diff --git a/.gitattributes b/.gitattributes index aa168a0ae..b1f760ebe 100644 --- a/.gitattributes +++ b/.gitattributes @@ -16,3 +16,14 @@ *.parquet binary *.zip binary *.duckdb binary + +# A captured PostgreSQL log is a FIXTURE of real bytes: the parser keys on the exact line endings +# and the tab-indented continuation, so line-ending normalisation would change what it is evidence of. +Darling/Darling.Tests/Fixtures/auto_explain_real_block.txt -text + +# The PostgreSQL verification rig runs INSIDE Linux containers. A Dockerfile whose RUN lines carry a +# trailing CR breaks the shell that executes them, and the seed is fed to psql in the container the same +# way - so these keep LF regardless of the repo-wide eol=crlf default. +tools/pg-verification-rig/Dockerfile text eol=lf +tools/pg-verification-rig/*.yml text eol=lf +tools/pg-verification-rig/*.sql text eol=lf diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index 11c263ce9..cfeea7ebc 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -1,876 +1,926 @@ -name: Build - -on: - push: - branches: [main, dev] - pull_request: - branches: [main, dev] - release: - types: [published] - # Merge-queue runs. Inert until a queue ruleset is enabled on a branch (a repo setting), - # but the required checks must handle the event BEFORE that click, or every queued PR - # stalls on checks that never report. dorny/paths-filter v4.0.1+ resolves merge_group - # diffs from the payload's base_sha/head_sha whenever the base input is empty — exactly - # what the filter steps pass for non-push events — so path classification works - # unchanged in a queue run. - merge_group: - -permissions: - contents: write - id-token: write - actions: read - -# A re-push to a PR cancels that PR's superseded in-flight run, and a push to dev/main -# cancels that BRANCH's superseded in-flight run — newest SHA wins. Finishing a build of -# code that is no longer the head helps nobody, and the shared Windows runner pool is what -# serializes everyone's CI (#1697 sat queued behind two dev builds; on 2026-07-26 a -# ~20-merge train left 13 of the day's 30 dev-push runs finishing SHAs a newer merge had -# already replaced — ~60 reclaimable runner-minutes in one evening). Cancelling a -# superseded PUSH run is safe because a push run produces nothing any other run consumes: -# every upload-artifact step in this workflow is gated to the release event (the SignPath -# signing path) or to failure() (darling-pg diagnostics), nothing in the repo downloads -# cross-run artifacts (no download-artifact, gh run download, or workflow_run consumer -# exists), nightly.yml builds its own tree from its own checkout, and a release compiles -# fresh on the release event. The accepted trade: push builds are diff-scoped, so a -# cancelled run's areas are not re-verified until the next change touches them — the -# nightly and the all-areas dev->main release PR are the backstops. Release and -# merge-queue runs deliberately keep a UNIQUE group per run (run_id) and are NEVER -# cancelled: a release build waits on SignPath's manual approval gate, and a queue -# validation is the last check before its result lands on dev. -concurrency: - group: ${{ github.event_name == 'pull_request' && format('build-pr-{0}', github.event.pull_request.number) || github.event_name == 'push' && format('build-push-{0}', github.ref) || format('build-run-{0}', github.run_id) }} - cancel-in-progress: ${{ github.event_name == 'pull_request' || github.event_name == 'push' }} - -jobs: - build: - runs-on: windows-latest - - steps: - - uses: actions/checkout@v7 - - - name: Detect changed paths - id: filter - if: github.event_name != 'release' - uses: dorny/paths-filter@v4 - with: - # On push events, compare against the previous commit on this branch - # (github.event.before). Without this, the action defaults to comparing - # against the default branch on non-default branch pushes, which would - # match every accumulated change and defeat the filter. - base: ${{ github.event_name == 'push' && github.event.before || '' }} - # Emit the matched file list so the fast-path step can NAME what it classified - # as documentation. A fast path that silently under-builds is the failure mode - # worth guarding against, so the reason is always printed, never inferred. - list-files: shell - filters: | - # A change to a root build file (the solution, restore config, or THIS workflow) can - # affect every product, so it forces a full build/test/publish. - root: - - 'PerformanceMonitor.sln' - - 'global.json' - - 'nuget.config' - - 'NuGet.config' - - '.github/workflows/build.yml' - # The shared PerformanceMonitor.* core libraries feed Lite, the Full Dashboard, AND - # Darling (verified via ProjectReference), so a change here fans out to all three. - # NOT the CLI Installer — it references only Installer.Core. - # - # Every area pattern says `dir/**/!(*.md)` — any non-markdown file under the - # area — instead of the old `dir/**` include plus a bare `!**/*.md` exclude. - # That is not style: dorny v4 evaluates each pattern as an INDEPENDENT - # predicate under the default predicate-quantifier 'some' (a filter is true - # when any changed file matches at least one rule), so a bare `!**/*.md` line - # is not a subtraction — it is its own rule meaning "any file that is not - # markdown", which silently made every area filter true for ANY non-markdown - # change anywhere in the repo. Measured proof: a single root .gitignore edit - # built and tested all four products and ran the full Darling PG suite - # (PR #1714, run 30219202642, filter log: "Filter darling = true, Matching - # files: .gitignore"). The extglob keeps the markdown carve-out INSIDE the - # include, where quantifier semantics cannot detach it. - core: - - 'PerformanceMonitor.Alerting/**/!(*.md)' - - 'PerformanceMonitor.Analysis/**/!(*.md)' - - 'PerformanceMonitor.Collectors/**/!(*.md)' - - 'PerformanceMonitor.Common/**/!(*.md)' - - 'PerformanceMonitor.Notifications/**/!(*.md)' - - 'PerformanceMonitor.PlanAnalysis/**/!(*.md)' - - 'PerformanceMonitor.Ui/**/!(*.md)' - # Installer.Core is shared by the CLI Installer AND the Full Dashboard's integrated - # installer — a change rebuilds both, and nothing else. - installer_core: - - 'deprecated/Installer.Core/**/!(*.md)' - dashboard: - - 'deprecated/Dashboard/**/!(*.md)' - - 'deprecated/Dashboard.Tests/**/!(*.md)' - lite: - - 'Lite/**/!(*.md)' - - 'Lite.Tests/**/!(*.md)' - # Same silently-stops-guarding reason as the darling filter's Lite entries below: - # Lite.Tests/ThemeParityLiteDarlingTests.cs READS the Darling viewer's theme - # dictionaries to assert the two apps' shared brush keys still resolve to the same - # colors. A Darling-theme-only edit is exactly the drift that guard exists to catch, - # so it has to reach the suite. - - 'Darling/PerformanceMonitor.Darling.Viewer/Themes/*.xaml' - installer: - - 'deprecated/Installer/**/!(*.md)' - - 'deprecated/Installer.Tests/**/!(*.md)' - - 'install/**/!(*.md)' - - 'upgrades/**/!(*.md)' - darling: - - 'Darling/**/!(*.md)' - # nightly.yml is not a build input, but Darling.Tests PARSES it: the #1888 - # guard reads both workflows' throwaway-cluster settings and compares them - # against the product's worker-sizing formula. Without this, a nightly-only - # edit would change a file the guard asserts on while never running the - # guard — a guard that silently stops guarding, which is the exact failure - # mode the source-parsing tests here exist to prevent. - - '.github/workflows/nightly.yml' - # Same reason, Lite side: the #1949 pin in Darling.Tests asserts every twinned - # query grid carries the SAME column sequence in both front ends, so it reads - # these six Lite files. A Lite-only XAML edit has to reach the suite or the - # parity half of that guard stops guarding. - - 'Lite/Controls/ServerTab.xaml' - - 'Lite/Controls/FinOpsTab.xaml' - - 'Lite/Windows/WaitDrillDownWindow.xaml' - - 'Lite/Windows/ProcedureHistoryWindow.xaml' - - 'Lite/Windows/QueryStatsHistoryWindow.xaml' - - 'Lite/Windows/QueryStoreHistoryWindow.xaml' - # #2114: XamlStaticResourceHygieneTests scans EVERY Lite XAML file — a StaticResource - # regression in one outside the six named above must still trigger the Darling job - # that runs the guard, or it slips to the nightly. - - 'Lite/**/*.xaml' - # The DOCUMENTATION allowlist: files that cannot affect a build under any - # job in this workflow. Deliberately an allowlist of non-executable content, - # not a "everything that isn't code" subtraction — a new file type defaults - # to being treated as code, which is the safe direction to be wrong in. - # - # NOT here, on purpose: *.sql (the installer and sql-validation compile it), - # *.yml (workflows), *.csproj / *.props / packages.lock.json (build inputs), - # and *.cs regardless of how comment-only the change looks — an XML doc - # comment still recompiles, and the compiler is what proves it still builds. - # - # The docs/ and Screenshots/ entries are extension-explicit rather than bare - # directory globs for the same reason: everything in them today is markdown, - # SVG, or a screenshot image, and a .sql or script dropped into either - # directory tomorrow should default to being code, not inherit a free pass - # from its parent directory. - docs: - - '**/*.md' - - 'LICENSE' - - 'CITATION.cff' - - '.gitignore' - - '.gitattributes' - - 'docs/**/*.{md,svg,png,jpg,jpeg,gif}' - - 'Screenshots/**/*.{md,svg,png,jpg,jpeg,gif}' - # Catch-all COUNTER, not a boolean gate: the classify step below decides - # "documentation-only" by comparing all_count to docs_count — they are equal - # exactly when every changed file sits on the docs allowlist. Stated as a - # count comparison because the previous shape ('**' plus '!' exclusions, - # a code: filter) could never be false under predicate-quantifier 'some' — - # every file matches '**', so the #1712 fast path shipped unable to engage - # (throwaway PR #1714: a .gitignore-only diff still paid setup + restore and, - # via the predicate bug above, a full build). - all: - - '**' - - # Decides the docs fast path ONCE, in one place, and says so out loud. Guards keep - # it off every path where a skipped restore would be a real loss: - # release — the filter step does not even run there, and a release must always - # compile and publish from a cold, fully restored tree. - # push — dev/main pushes are the integration signal for what just merged, so - # they restore unconditionally even for a docs-only commit. Cheap - # insurance: this only forces the restore back on, it does not force - # the per-product build/test steps, which stay path-gated as before. - # merge_group — a queue run is the LAST validation before its result lands on dev, - # so it takes the same always-restore path as a push. - # areas — belt and suspenders: even when the counts say docs-only, any lit - # area filter vetoes the fast path, because an area=true with restore - # skipped would run `dotnet build --no-restore` against nothing. The - # two classifications are built from the same allowlist so they cannot - # disagree today; this guard is for the day someone edits one and not - # the other. - # Everything else (pull_request) is eligible, and engages only when EVERY changed - # file is on the documentation allowlist (all_count == docs_count). - - name: Classify change for the docs fast path - id: fastpath - shell: bash - env: - ALL_COUNT: ${{ steps.filter.outputs.all_count }} - DOCS_COUNT: ${{ steps.filter.outputs.docs_count }} - DOCS_FILES: ${{ steps.filter.outputs.docs_files }} - AREAS: 'root=${{ steps.filter.outputs.root }} core=${{ steps.filter.outputs.core }} installer_core=${{ steps.filter.outputs.installer_core }} dashboard=${{ steps.filter.outputs.dashboard }} lite=${{ steps.filter.outputs.lite }} installer=${{ steps.filter.outputs.installer }} darling=${{ steps.filter.outputs.darling }}' - run: | - set -euo pipefail - - if [ "${{ github.event_name }}" = "release" ]; then - echo "engaged=false" >> "$GITHUB_OUTPUT" - echo "::notice title=Full build::Release event - the docs fast path never applies to a release." - exit 0 - fi - - if [ "${{ github.event_name }}" = "push" ] || [ "${{ github.event_name }}" = "merge_group" ]; then - echo "engaged=false" >> "$GITHUB_OUTPUT" - echo "::notice title=Full build::${{ github.event_name }} on '${{ github.ref_name }}' - integration runs always restore, even for a docs-only change." - exit 0 - fi - - echo "Changed files: ${ALL_COUNT:-0} total, ${DOCS_COUNT:-0} on the documentation allowlist. Areas: ${AREAS}" - - if [ "${ALL_COUNT:-0}" -gt 0 ] && [ "${ALL_COUNT:-0}" -eq "${DOCS_COUNT:-0}" ] && [[ "${AREAS}" != *"=true"* ]]; then - echo "engaged=true" >> "$GITHUB_OUTPUT" - echo "::notice title=DOCS FAST PATH ENGAGED::All ${ALL_COUNT} changed files are on the documentation allowlist, so .NET setup, restore and versioning are skipped. This job still reports its result." - echo "Documentation files classified in this change:" - for f in ${DOCS_FILES}; do echo " - ${f}"; done - else - echo "engaged=false" >> "$GITHUB_OUTPUT" - echo "::notice title=Full build::At least one changed file is off the documentation allowlist (${DOCS_COUNT:-0} of ${ALL_COUNT:-0} classified as documentation)." - fi - - - name: Setup .NET 10.0 - if: steps.fastpath.outputs.engaged != 'true' - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - - name: Restore dependencies - if: steps.fastpath.outputs.engaged != 'true' - run: | - dotnet restore Lite/PerformanceMonitorLite.csproj --locked-mode - dotnet restore Lite.Tests/Lite.Tests.csproj --locked-mode - dotnet restore deprecated/Installer.Tests/Installer.Tests.csproj --locked-mode - dotnet restore deprecated/Dashboard.Tests/Dashboard.Tests.csproj --locked-mode - dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode - dotnet restore Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj --locked-mode - - - name: Build Lite.Tests - if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet build Lite.Tests/Lite.Tests.csproj -c Release --no-restore - - - name: Build Installer.Tests - if: steps.filter.outputs.installer == 'true' || steps.filter.outputs.installer_core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet build deprecated/Installer.Tests/Installer.Tests.csproj -c Release --no-restore - - # The 'dashboard' path filter was defined when the Full Dashboard moved to deprecated/ (#1612) but - # never wired to a step, so its build and tests silently stopped running — which is how a batch of - # compiler warnings and three broken ThemeParityTests accumulated unnoticed (#1643). Deprecated means - # bug-fix-only, not unverified: it still compiles warning-free and its tests still guard cross-app - # parity (the theme palettes it checks are LITE's too). - - name: Build Dashboard.Tests - if: steps.filter.outputs.dashboard == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet build deprecated/Dashboard.Tests/Dashboard.Tests.csproj -c Release --no-restore - - - name: Build Darling - if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: | - dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore - dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release --no-restore - - # One step for the whole Lite suite. It was split into fast / analysis-heavy halves when the - # seven analysis classes rebuilt the full DuckDB schema inside every test and their subset - # alone cost ~9 minutes; after the shared class fixtures (#1693, #1698) and batched seeding - # (#1694) that subset runs in ~1 minute, so the split — and the narrower lite_analysis path - # gate that let non-analysis Lite changes skip it — stopped earning its second test-host - # spin-up and its filter-drift risk. - - name: Run Lite tests - if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet run --project Lite.Tests/Lite.Tests.csproj -c Release --no-build - - - name: Run Installer tests - if: steps.filter.outputs.installer == 'true' || steps.filter.outputs.installer_core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet run --project deprecated/Installer.Tests/Installer.Tests.csproj -c Release --no-build -- -class- "Installer.Tests.VersionDetectionTests" -class- "Installer.Tests.IdempotencyTests" -class- "Installer.Tests.AdversarialTests" - - - name: Run Dashboard tests - if: steps.filter.outputs.dashboard == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet run --project deprecated/Dashboard.Tests/Dashboard.Tests.csproj -c Release --no-build - - - name: Run Darling tests - if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build - - - name: Get version - if: steps.fastpath.outputs.engaged != 'true' - id: version - shell: pwsh - run: | - $version = ([xml](Get-Content Lite/PerformanceMonitorLite.csproj)).Project.PropertyGroup.Version | Where-Object { $_ } - echo "VERSION=$version" >> $env:GITHUB_OUTPUT - - - name: Publish Lite - if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -o publish/Lite - - - name: Publish Lite (self-contained for Velopack) - if: github.event_name == 'release' - run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -r win-x64 --self-contained -o publish/Lite-velopack - - - name: Publish Darling Service - if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o publish/DarlingService - - - name: Publish Darling Viewer - if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' - run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -o publish/DarlingViewer - - - name: Publish Darling Viewer (self-contained for Velopack) - if: github.event_name == 'release' - run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -r win-x64 --self-contained -o publish/DarlingViewer-velopack - - # Darling bundles a PostgreSQL 18 + TimescaleDB runtime (pg-runtime.zip) that ships beside - # the service exe; DarlingManagedPostgres extracts it on first run. The fetch script pulls - # ~340MB of pinned EDB/TimescaleDB archives, so this is release-only and cached. The key is - # the fetch script's own content hash (the SHA256 pins live inside it): a re-release with - # unchanged pins restores the assembled zip and skips both the download and the assembly, - # and any pin/version bump edits the script and invalidates the cache automatically. - - name: Cache Darling pg-runtime.zip - if: github.event_name == 'release' - id: cache-pg-runtime - uses: actions/cache@v6 - with: - path: Darling/artifacts/pg-runtime.zip - key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} - - - name: Build Darling pg-runtime.zip - if: github.event_name == 'release' && steps.cache-pg-runtime.outputs.cache-hit != 'true' - shell: pwsh - run: ./Darling/tools/fetch-pg-runtime.ps1 - - - name: Package release artifacts - if: github.event_name == 'release' - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - New-Item -ItemType Directory -Force -Path releases - - # Lite ZIP - portable artifact for advanced/air-gapped users. The README points end - # users at Setup.exe (Velopack); this ZIP is the explicit fallback. - Compress-Archive -Path 'publish/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force - - # upload-artifact is deliberately HELD at v6 (#1653): every signing step below consumes - # `steps.upload-*.outputs.artifact-id`, and v7 changes artifact archiving semantics (the - # `archive` parameter). The signing path only executes on `release: [published]`, so a broken - # bump surfaces at release time — bump only alongside a validated real signing run. - # Dependabot is configured to skip this major (see .github/dependabot.yml). - - name: Upload Lite for signing - if: github.event_name == 'release' - id: upload-lite - uses: actions/upload-artifact@v6 - with: - name: Lite-unsigned - path: publish/Lite/ - - - name: Stage Darling for signing - if: github.event_name == 'release' - shell: pwsh - run: | - $stage = 'publish/Darling-signing' - if (Test-Path $stage) { Remove-Item -Recurse -Force $stage } - New-Item -ItemType Directory -Force -Path "$stage/viewer" | Out-Null - Copy-Item 'publish/DarlingService/*' $stage -Recurse - Copy-Item 'publish/DarlingViewer/*' "$stage/viewer" -Recurse - - - name: Upload Darling for signing - if: github.event_name == 'release' - id: upload-darling - uses: actions/upload-artifact@v6 - with: - name: Darling-unsigned - path: publish/Darling-signing/ - - - name: Sign Lite - if: github.event_name == 'release' - uses: signpath/github-action-submit-signing-request@v2 - with: - api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' - organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' - project-slug: 'PerformanceMonitor' - signing-policy-slug: 'release-signing' - artifact-configuration-slug: 'Lite' - github-artifact-id: '${{ steps.upload-lite.outputs.artifact-id }}' - wait-for-completion: true - output-artifact-directory: 'signed/Lite' - - - name: Sign Darling - if: github.event_name == 'release' - uses: signpath/github-action-submit-signing-request@v2 - with: - api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' - organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' - project-slug: 'PerformanceMonitor' - signing-policy-slug: 'release-signing' - artifact-configuration-slug: 'Darling' - github-artifact-id: '${{ steps.upload-darling.outputs.artifact-id }}' - wait-for-completion: true - output-artifact-directory: 'signed/Darling' - - - name: Replace with signed artifacts - if: github.event_name == 'release' - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - # Re-zip signed files into release archives - Remove-Item "releases/PerformanceMonitorLite-$version.zip" -ErrorAction SilentlyContinue - Compress-Archive -Path 'signed/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force - - - name: Package Darling (signed) - if: github.event_name == 'release' - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - # One product, one zip (mirrors the one-zip-per-product convention above). signed/Darling - # already holds the signed tree in its final layout — the service at the archive root (its - # darling.sample.json alongside), the viewer in a viewer\ subfolder. Drop pg-runtime.zip - # beside the service exe, exactly where DarlingManagedPostgres looks (AppContext.BaseDirectory) - # and extracts it on first run. The EDB PostgreSQL binaries inside pg-runtime.zip are shipped - # as opaque data and were never signed. - Copy-Item 'Darling/artifacts/pg-runtime.zip' 'signed/Darling' - - Remove-Item "releases/PerformanceMonitorDarling-$version.zip" -ErrorAction SilentlyContinue - Compress-Archive -Path 'signed/Darling/*' -DestinationPath "releases/PerformanceMonitorDarling-$version.zip" -Force - - # The Velopack (Setup.exe) path publishes a SEPARATE self-contained build - # (publish/Dashboard-velopack, publish/Lite-velopack -- see the "self-contained - # for Velopack" steps above) that previously went straight into `vpk pack` - # without ever being uploaded to SignPath. Only the framework-dependent trees - # used for the legacy ZIPs were signed, so every Setup.exe shipped unsigned - # since Velopack packaging was introduced. These steps close that gap by - # mirroring the exact upload/sign pattern used for Dashboard/Lite/Installer - # above, and vpk pack below now reads from the signed output. - - name: Upload Lite (Velopack) for signing - if: github.event_name == 'release' - id: upload-lite-velopack - uses: actions/upload-artifact@v6 - with: - name: Lite-Velopack-unsigned - path: publish/Lite-velopack/ - - - name: Upload Darling Viewer (Velopack) for signing - if: github.event_name == 'release' - id: upload-darlingviewer-velopack - uses: actions/upload-artifact@v6 - with: - name: DarlingViewer-Velopack-unsigned - path: publish/DarlingViewer-velopack/ - - - name: Sign Lite (Velopack) - if: github.event_name == 'release' - uses: signpath/github-action-submit-signing-request@v2 - with: - api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' - organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' - project-slug: 'PerformanceMonitor' - signing-policy-slug: 'release-signing' - artifact-configuration-slug: 'Lite' - github-artifact-id: '${{ steps.upload-lite-velopack.outputs.artifact-id }}' - wait-for-completion: true - output-artifact-directory: 'signed/Lite-Velopack' - - # The remote-seat viewer Setup.exe (#1555) is signed with its OWN 'DarlingViewer' artifact - # configuration — the co-located-zip 'Darling' slug signs a service+viewer\ tree layout, which - # does not match this self-contained viewer-at-root publish. Like the 'Darling' slug, the - # 'DarlingViewer' config (which files get signed) lives outside this repo on signpath.io; until - # Erik creates the slug this step fails the release — the standing #1340 SignPath prerequisite. - - name: Sign Darling Viewer (Velopack) - if: github.event_name == 'release' - uses: signpath/github-action-submit-signing-request@v2 - with: - api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' - organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' - project-slug: 'PerformanceMonitor' - signing-policy-slug: 'release-signing' - artifact-configuration-slug: 'DarlingViewer' - github-artifact-id: '${{ steps.upload-darlingviewer-velopack.outputs.artifact-id }}' - wait-for-completion: true - output-artifact-directory: 'signed/DarlingViewer-Velopack' - - - name: Create Velopack releases (Lite + Darling Viewer) - if: github.event_name == 'release' - shell: pwsh - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - VERSION: ${{ steps.version.outputs.VERSION }} - run: | - # Pin vpk to the Velopack library version (keep in sync with the Velopack - # PackageReference in PerformanceMonitorLite.csproj). - dotnet tool install -g vpk --version 1.2.0 - New-Item -ItemType Directory -Force -Path releases/velopack-lite - New-Item -ItemType Directory -Force -Path releases/velopack-darlingviewer - - # Lite: download previous + pack (from the SIGNED velopack output, not the raw publish dir) - vpk download github --repoUrl https://github.com/${{ github.repository }} --channel lite -o releases/velopack-lite --token $env:GH_TOKEN - vpk pack -u PerformanceMonitorLite -v $env:VERSION -p signed/Lite-Velopack -e PerformanceMonitorLite.exe -o releases/velopack-lite --channel lite - - # Darling Viewer remote-seat installer (#1555): download previous + pack (from the SIGNED - # velopack output). Its own 'darlingviewer' channel/delta feed, separate from the co-located - # viewer inside PerformanceMonitorDarling-*.zip (which stays plain-zip only). - vpk download github --repoUrl https://github.com/${{ github.repository }} --channel darlingviewer -o releases/velopack-darlingviewer --token $env:GH_TOKEN - vpk pack -u PerformanceMonitorDarlingViewer -v $env:VERSION -p signed/DarlingViewer-Velopack -e PerformanceMonitor.Darling.Viewer.exe -o releases/velopack-darlingviewer --channel darlingviewer - - - name: Generate checksums - if: github.event_name == 'release' - shell: pwsh - run: | - $checksums = Get-ChildItem releases/*.zip | ForEach-Object { - $hash = (Get-FileHash $_.FullName -Algorithm SHA256).Hash.ToLower() - "$hash $($_.Name)" - } - $checksums | Out-File -FilePath releases/SHA256SUMS.txt -Encoding utf8 - Write-Host "Checksums:" - $checksums | ForEach-Object { Write-Host $_ } - - - name: Upload release assets - if: github.event_name == 'release' - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - run: | - gh release upload ${{ github.event.release.tag_name }} releases/*.zip releases/SHA256SUMS.txt --clobber - - - name: Upload Lite Velopack artifacts - if: github.event_name == 'release' - shell: pwsh - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - VERSION: ${{ steps.version.outputs.VERSION }} - run: | - vpk upload github --repoUrl https://github.com/${{ github.repository }} --channel lite -o releases/velopack-lite --releaseName "v$env:VERSION" --tag "v$env:VERSION" --merge --token $env:GH_TOKEN - - - name: Upload Darling Viewer Velopack artifacts - if: github.event_name == 'release' - shell: pwsh - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - VERSION: ${{ steps.version.outputs.VERSION }} - run: | - vpk upload github --repoUrl https://github.com/${{ github.repository }} --channel darlingviewer -o releases/velopack-darlingviewer --releaseName "v$env:VERSION" --tag "v$env:VERSION" --merge --token $env:GH_TOKEN - - # #1587: gated-live Darling coverage BEFORE merge, not only in the nightly. Darling.Tests has - # live-PostgreSQL tests (the *_AgainstDevPostgres classes) gated on DARLING_TEST_PG, which the - # build job above never sets — so those tests only ever ran post-merge in nightly.yml. That is - # exactly how #1586's alter_job(bigint) bug merged AND deployed clean: the test that catches it is - # gated-live, and PR CI did not run it. This job stands up a throwaway PostgreSQL + TimescaleDB - # from the bundled pg-runtime and runs the FULL Darling suite against it — but ONLY when Darling - # code (or this workflow) changed, so a Lite/Dashboard-only PR pays nothing. On a non-Darling - # change every step below no-ops via the path filter and the job still reports SUCCESS, so it can - # be made a required check without blocking unrelated PRs (the same always-runs-reports-a-result - # shape the build job uses for doc-only changes). Mirrors the nightly darling-pg job step-for-step; - # the one live-SQL-Server E2E stays skipped (DARLING_TEST_SQL unset — no SQL Server on the runner). - darling-pg: - name: Darling PostgreSQL tests - runs-on: windows-latest - # Max observed on a warm cache is ~3m40s; a cold pg-runtime cache adds a ~340MB fetch. - # 30 minutes is 3x headroom over the cold path — past that, something is hung (pg_ctl -w - # waiting on a cluster that will never come up), and the default 6h timeout would hold a - # shared-pool Windows runner hostage for the duration. The build job above deliberately - # has NO timeout: on release it waits on SignPath's manual approval gate, which can - # legitimately take hours. - timeout-minutes: 30 - permissions: - contents: read - - steps: - - uses: actions/checkout@v7 - - # Only do the expensive TimescaleDB work when Darling code changed — or when THIS workflow - # changed, so a change to the gate itself is exercised by the gate (this is what makes the PR - # that introduces this job validate itself end-to-end). Doc-only Darling edits don't trigger - # it. Skipped entirely on release: the dev push that produced the release commit already ran - # it, so the filter step doesn't run and every step below no-ops. - - name: Detect changed paths - id: filter - if: github.event_name != 'release' - uses: dorny/paths-filter@v4 - with: - # On push, compare against the previous commit on this branch (mirrors the build job); - # on pull_request, an empty base makes the action diff against the PR base branch. - base: ${{ github.event_name == 'push' && github.event.before || '' }} - # `Darling/**/!(*.md)` instead of a `Darling/**` include plus a `!Darling/**/*.md` - # exclude: dorny v4 treats each pattern as an independent predicate under the - # default quantifier, so the old bare negation was itself a match-all-non-Darling-md - # rule — this job ran the full TimescaleDB suite on every PR, including md-only - # ones (run 30218459544: "Filter darling = true, Matching files: CHANGELOG.md"). - filters: | - darling: - - 'Darling/**/!(*.md)' - - '.github/workflows/build.yml' - # Same reason as the build job's darling filter: the #1888 cluster-sizing - # guard parses nightly.yml, so an edit to it has to reach the suite. - - '.github/workflows/nightly.yml' - - # This job's gate was already correct for documentation — a docs-only change leaves - # 'darling' false and every step below no-ops. What it lacked was SAYING so: a job - # that reports success having quietly run nothing looks identical to one that tested - # everything. Costs one step; buys a log you can point at when asking "did this - # actually get tested?". - - name: Report the Darling PG gate decision - shell: bash - run: | - set -euo pipefail - - if [ "${{ github.event_name }}" = "release" ]; then - echo "::notice title=Darling PG tests skipped::Release event - the dev push that produced this commit already ran them." - elif [ "${{ steps.filter.outputs.darling }}" = "true" ]; then - echo "::notice title=Darling PG tests running::Darling code (or this workflow) changed." - else - echo "::notice title=Darling PG tests skipped::No Darling code changed - documentation-only Darling edits do not trigger the TimescaleDB suite." - fi - - - name: Setup .NET 10.0 - if: steps.filter.outputs.darling == 'true' - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - # Same cache key the release job and nightly.yml use (the fetch script's own content hash), so - # a warm cache from any of the three means no ~340MB EDB/TimescaleDB download here. - - name: Cache Darling pg-runtime.zip - if: steps.filter.outputs.darling == 'true' - id: cache-pg-runtime - uses: actions/cache@v6 - with: - path: Darling/artifacts/pg-runtime.zip - key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} - - - name: Build Darling pg-runtime.zip (cache miss only) - if: steps.filter.outputs.darling == 'true' && steps.cache-pg-runtime.outputs.cache-hit != 'true' - shell: pwsh - run: ./Darling/tools/fetch-pg-runtime.ps1 - - - name: Extract pg-runtime - if: steps.filter.outputs.darling == 'true' - shell: pwsh - run: | - $zip = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime.zip" - $dest = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime" - if (Test-Path $dest) { Remove-Item -Recurse -Force $dest } - Add-Type -AssemblyName System.IO.Compression.FileSystem - [System.IO.Compression.ZipFile]::ExtractToDirectory($zip, $dest) - if (-not (Test-Path "$dest\pgsql\bin\pg_ctl.exe")) { throw "pg-runtime missing pgsql\bin\pg_ctl.exe" } - - - name: Restore Darling.Tests - if: steps.filter.outputs.darling == 'true' - run: dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode - - - name: Build Darling.Tests - if: steps.filter.outputs.darling == 'true' - run: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore - - # Throwaway cluster from the bundled runtime. initdb TRUST auth is acceptable ONLY here: an - # ephemeral CI runner, loopback-only, throwaway data (the product default is scram-sha-256 + - # a generated credential). Superuser is "darling" (not "postgres") so CREATE SCHEMA ... - # AUTHORIZATION darling in the V8 schema-split migration resolves. Settings mirror - # DarlingManagedPostgres.BuildConfAppend (timescaledb preload, port 5541, loopback bind) - # and BuildWorkerSizingConfAppend (the two worker settings below). - # - # #1888: the worker settings are NOT optional garnish. PostgreSQL's default - # max_worker_processes = 8 cannot launch TimescaleDB's per-hypertable compression, - # retention and continuous-aggregate policy jobs — the postmaster logs "failed to start a - # background worker" storms and most policy runs never happen. Without them this job tested a - # configuration NO customer runs (the product refuses to: BuildWorkerSizingConfAppend writes - # these on every managed start), and worse, made failures luck-of-the-slot rather than - # reproducible — which is exactly how #1862's compression flake failed twice on CI in two - # unrecognizably different ways and could not be reproduced locally at all. - # - # The values are the product's own derivation from the live hypertable count - # (TimescaleSupport.HypertableCount = the 50-collector catalog + collection_log = 51): - # timescaledb.max_background_workers = HypertableCount + 2 = 53 - # max_worker_processes = 3 + (HypertableCount + 2) + 8 = 64 - # Hard-coded here because a workflow cannot call into the product — so - # CiClusterWorkerSizingTests parses THIS FILE and fails the build if either number stops - # matching the formula as collectors are added, and CiClusterWorkerSizingLiveTests asserts - # the running cluster actually serves them (a conf line that never took effect is - # indistinguishable from one that did, by inspection). - - name: Initialize and start throwaway PostgreSQL - if: steps.filter.outputs.darling == 'true' - shell: pwsh - run: | - $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" - $dataDir = "$env:RUNNER_TEMP\darling-pgdata" - $logFile = "$env:RUNNER_TEMP\darling-pg.log" - & "$bin\initdb.exe" -D $dataDir -U darling -A trust --encoding=UTF8 - if ($LASTEXITCODE -ne 0) { throw "initdb failed ($LASTEXITCODE)" } - Add-Content -Path "$dataDir\postgresql.conf" -Value "shared_preload_libraries = 'timescaledb'" - Add-Content -Path "$dataDir\postgresql.conf" -Value "port = 5541" - Add-Content -Path "$dataDir\postgresql.conf" -Value "listen_addresses = '127.0.0.1'" - Add-Content -Path "$dataDir\postgresql.conf" -Value "timescaledb.max_background_workers = 53" - Add-Content -Path "$dataDir\postgresql.conf" -Value "max_worker_processes = 64" - & "$bin\pg_ctl.exe" -D $dataDir -l $logFile -w start - if ($LASTEXITCODE -ne 0) { if (Test-Path $logFile) { Get-Content $logFile -Tail 50 }; throw "pg_ctl start failed ($LASTEXITCODE)" } - & "$bin\createdb.exe" -h 127.0.0.1 -p 5541 -U darling darling - if ($LASTEXITCODE -ne 0) { throw "createdb failed ($LASTEXITCODE)" } - - # DARLING_TEST_PG lights up the [Collection("live-postgres")] classes; DARLING_TEST_PGRUNTIME - # lights up the managed-bootstrap E2E. DARLING_TEST_SQL is intentionally unset — the one - # live-SQL-Server E2E stays skipped (no SQL Server on the runner). Full suite (not a filtered - # subset): the ungated ~3s overlap with the build job is negligible and a filter could hide a - # test the way narrow filters have bitten before. - - name: Run Darling PG tests - if: steps.filter.outputs.darling == 'true' - shell: pwsh - env: - DARLING_TEST_PG: "Host=127.0.0.1;Port=5541;Username=darling;Database=darling" - DARLING_TEST_PGRUNTIME: ${{ github.workspace }}\Darling\artifacts\pg-runtime - run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build -- -trx TestResults/darling-pr.trx - - - name: Stop PostgreSQL - if: always() && steps.filter.outputs.darling == 'true' - shell: pwsh - run: | - $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" - $dataDir = "$env:RUNNER_TEMP\darling-pgdata" - if (Test-Path "$bin\pg_ctl.exe") { & "$bin\pg_ctl.exe" -D $dataDir -m fast -w stop } - exit 0 - - - name: Upload PG log and test results on failure - if: failure() && steps.filter.outputs.darling == 'true' - uses: actions/upload-artifact@v6 - with: - name: darling-pg-failure - path: | - ${{ runner.temp }}/darling-pg.log - TestResults/ - if-no-files-found: ignore - - # ── Linux service build + container image (#1804) ──────────────────────────────────────────────── - # The Darling service is cross-platform .NET on purpose, but until this job nothing PROVED it on - # every PR — the linux-x64 publish and the container image both built for the first time at release - # time or never. Same path-filter shape as darling-pg above: only runs the expensive work when - # Darling/service code (or this workflow, or the Dockerfile) changed, always reports a result so it - # can be a required check. No tests run here — the test projects are net10.0-windows (they reference - # the WPF apps); the cross-platform behavior they pin is exercised by the Windows jobs, and the - # container smoke lives in the compose quickstart. On PRs/pushes this job answers exactly two - # questions: does the service still publish for linux-x64, and does the image still build. - # - # On the RELEASE event this job is the Linux PUBLISHER: before it, a stable release shipped Windows - # zips and Setup.exes while the linux tar.gz and the ghcr image existed only at nightly quality - # (nightly.yml, tag :nightly) — the compose quickstart pointed released users at a nightly image. - # Now the release uploads PerformanceMonitorDarling-linux-x64-.tar.gz + SHA256SUMS-linux.txt to - # the release and pushes ghcr : and :latest. Linux binaries are NOT SignPath-signed - # (SignPath signs Windows PEs); the checksums file is the integrity story, same as the nightly. - darling-linux: - name: Darling Linux build - runs-on: ubuntu-latest - timeout-minutes: 30 - permissions: - contents: write - packages: write - # Sigstore keyless signing + GitHub provenance (proven on the nightly first): id-token yields - # the OIDC identity Fulcio certifies against; attestations stores the tarball's provenance. - id-token: write - attestations: write - - steps: - - uses: actions/checkout@v7 - - - name: Detect changed paths - id: filter - if: github.event_name != 'release' - uses: dorny/paths-filter@v4 - with: - base: ${{ github.event_name == 'push' && github.event.before || '' }} - filters: | - darling: - - 'Darling/**/!(*.md)' - - 'PerformanceMonitor.Common/**' - - 'PerformanceMonitor.Collectors/**' - - 'PerformanceMonitor.Analysis/**' - - '.github/workflows/build.yml' - - - name: Report the Linux gate decision - shell: bash - run: | - set -euo pipefail - if [ "${{ github.event_name }}" = "release" ]; then - echo "::notice title=Darling Linux publishing::Release event - packaging the versioned linux tar.gz and pushing the ghcr image." - elif [ "${{ steps.filter.outputs.darling }}" = "true" ]; then - echo "::notice title=Darling Linux build running::Darling/service code (or this workflow) changed." - else - echo "::notice title=Darling Linux build skipped::No Darling/service code changed." - fi - - - name: Setup .NET 10.0 - if: github.event_name == 'release' || steps.filter.outputs.darling == 'true' - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - - name: Publish service (linux-x64) - if: github.event_name == 'release' || steps.filter.outputs.darling == 'true' - run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -r linux-x64 --self-contained false -o publish/DarlingService-linux - - - name: Build container image - if: github.event_name != 'release' && steps.filter.outputs.darling == 'true' - run: docker build -f Darling/Dockerfile -t performancemonitor-darling:pr . - - # ── Release-only publishing (mirrors nightly.yml's linux job, versioned instead of :nightly) ── - - - name: Get version - if: github.event_name == 'release' - id: version - shell: bash - run: | - set -euo pipefail - version="$(grep -oPm1 '(?<=)[^<]+' Lite/PerformanceMonitorLite.csproj)" - echo "VERSION=${version}" >> "$GITHUB_OUTPUT" - - - name: Package linux artifact + checksum - if: github.event_name == 'release' - shell: bash - run: | - set -euo pipefail - version="${{ steps.version.outputs.VERSION }}" - mkdir -p releases - tar -C publish/DarlingService-linux -czf "releases/PerformanceMonitorDarling-linux-x64-${version}.tar.gz" . - (cd releases && sha256sum "PerformanceMonitorDarling-linux-x64-${version}.tar.gz" > SHA256SUMS-linux.txt && cat SHA256SUMS-linux.txt) - - - name: Upload linux artifact to the release - if: github.event_name == 'release' - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - run: gh release upload ${{ github.event.release.tag_name }} releases/PerformanceMonitorDarling-linux-x64-*.tar.gz releases/SHA256SUMS-linux.txt --clobber - - - name: Build and push container image (ghcr, versioned + latest) - if: github.event_name == 'release' - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - shell: bash - run: | - set -euo pipefail - version="${{ steps.version.outputs.VERSION }}" - image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" - echo "$GH_TOKEN" | docker login ghcr.io -u "${{ github.actor }}" --password-stdin - docker build -f Darling/Dockerfile -t "${image}:${version}" -t "${image}:latest" . - docker push "${image}:${version}" - docker push "${image}:latest" - - # Keyless Sigstore signature on the released image — the same steps the nightly runs (verified - # end-to-end from a client 2026-08-06: cosign validated the claims, the Rekor log entry, and the - # workflow identity). SignPath's cosign support is edition-gated, and this path needs no keys or - # subscription at all. Verify: - # cosign verify ghcr.io/erikdarlingdata/performancemonitor-darling: \ - # --certificate-identity-regexp 'github.com/erikdarlingdata/PerformanceMonitor' \ - # --certificate-oidc-issuer https://token.actions.githubusercontent.com - - name: Install cosign - if: github.event_name == 'release' - uses: sigstore/cosign-installer@v3 - - - name: Sign container image (keyless, Sigstore) - if: github.event_name == 'release' - shell: bash - run: | - set -euo pipefail - version="${{ steps.version.outputs.VERSION }}" - image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" - digest="$(docker inspect --format='{{index .RepoDigests 0}}' "${image}:${version}")" - cosign sign --yes "${digest}" - - # GitHub-native SLSA provenance for the released tarball: gh attestation verify -R . - - name: Attest the linux tarball (GitHub provenance) - if: github.event_name == 'release' - uses: actions/attest-build-provenance@v4 - with: - subject-path: releases/PerformanceMonitorDarling-linux-x64-*.tar.gz +name: Build + +on: + push: + branches: [main, dev] + pull_request: + branches: [main, dev] + release: + types: [published] + # Merge-queue runs. Inert until a queue ruleset is enabled on a branch (a repo setting), + # but the required checks must handle the event BEFORE that click, or every queued PR + # stalls on checks that never report. dorny/paths-filter v4.0.1+ resolves merge_group + # diffs from the payload's base_sha/head_sha whenever the base input is empty — exactly + # what the filter steps pass for non-push events — so path classification works + # unchanged in a queue run. + merge_group: + +permissions: + contents: write + id-token: write + actions: read + +# A re-push to a PR cancels that PR's superseded in-flight run, and a push to dev/main +# cancels that BRANCH's superseded in-flight run — newest SHA wins. Finishing a build of +# code that is no longer the head helps nobody, and the shared Windows runner pool is what +# serializes everyone's CI (#1697 sat queued behind two dev builds; on 2026-07-26 a +# ~20-merge train left 13 of the day's 30 dev-push runs finishing SHAs a newer merge had +# already replaced — ~60 reclaimable runner-minutes in one evening). Cancelling a +# superseded PUSH run is safe because a push run produces nothing any other run consumes: +# every upload-artifact step in this workflow is gated to the release event (the SignPath +# signing path) or to failure() (darling-pg diagnostics), nothing in the repo downloads +# cross-run artifacts (no download-artifact, gh run download, or workflow_run consumer +# exists), nightly.yml builds its own tree from its own checkout, and a release compiles +# fresh on the release event. The accepted trade: push builds are diff-scoped, so a +# cancelled run's areas are not re-verified until the next change touches them — the +# nightly and the all-areas dev->main release PR are the backstops. Release and +# merge-queue runs deliberately keep a UNIQUE group per run (run_id) and are NEVER +# cancelled: a release build waits on SignPath's manual approval gate, and a queue +# validation is the last check before its result lands on dev. +concurrency: + group: ${{ github.event_name == 'pull_request' && format('build-pr-{0}', github.event.pull_request.number) || github.event_name == 'push' && format('build-push-{0}', github.ref) || format('build-run-{0}', github.run_id) }} + cancel-in-progress: ${{ github.event_name == 'pull_request' || github.event_name == 'push' }} + +jobs: + build: + runs-on: windows-latest + + steps: + - uses: actions/checkout@v7 + + - name: Detect changed paths + id: filter + if: github.event_name != 'release' + uses: dorny/paths-filter@v4 + with: + # On push events, compare against the previous commit on this branch + # (github.event.before). Without this, the action defaults to comparing + # against the default branch on non-default branch pushes, which would + # match every accumulated change and defeat the filter. + base: ${{ github.event_name == 'push' && github.event.before || '' }} + # Emit the matched file list so the fast-path step can NAME what it classified + # as documentation. A fast path that silently under-builds is the failure mode + # worth guarding against, so the reason is always printed, never inferred. + list-files: shell + filters: | + # A change to a root build file (the solution, restore config, or THIS workflow) can + # affect every product, so it forces a full build/test/publish. + root: + - 'PerformanceMonitor.sln' + - 'global.json' + - 'nuget.config' + - 'NuGet.config' + - '.github/workflows/build.yml' + # The shared PerformanceMonitor.* core libraries feed Lite, the Full Dashboard, AND + # Darling (verified via ProjectReference), so a change here fans out to all three. + # NOT the CLI Installer — it references only Installer.Core. + # + # Every area pattern says `dir/**/!(*.md)` — any non-markdown file under the + # area — instead of the old `dir/**` include plus a bare `!**/*.md` exclude. + # That is not style: dorny v4 evaluates each pattern as an INDEPENDENT + # predicate under the default predicate-quantifier 'some' (a filter is true + # when any changed file matches at least one rule), so a bare `!**/*.md` line + # is not a subtraction — it is its own rule meaning "any file that is not + # markdown", which silently made every area filter true for ANY non-markdown + # change anywhere in the repo. Measured proof: a single root .gitignore edit + # built and tested all four products and ran the full Darling PG suite + # (PR #1714, run 30219202642, filter log: "Filter darling = true, Matching + # files: .gitignore"). The extglob keeps the markdown carve-out INSIDE the + # include, where quantifier semantics cannot detach it. + core: + - 'PerformanceMonitor.Alerting/**/!(*.md)' + - 'PerformanceMonitor.Analysis/**/!(*.md)' + - 'PerformanceMonitor.Collectors/**/!(*.md)' + - 'PerformanceMonitor.Common/**/!(*.md)' + - 'PerformanceMonitor.Notifications/**/!(*.md)' + - 'PerformanceMonitor.PlanAnalysis/**/!(*.md)' + - 'PerformanceMonitor.Ui/**/!(*.md)' + # Installer.Core is shared by the CLI Installer AND the Full Dashboard's integrated + # installer — a change rebuilds both, and nothing else. + installer_core: + - 'deprecated/Installer.Core/**/!(*.md)' + dashboard: + - 'deprecated/Dashboard/**/!(*.md)' + - 'deprecated/Dashboard.Tests/**/!(*.md)' + lite: + - 'Lite/**/!(*.md)' + - 'Lite.Tests/**/!(*.md)' + # #2489: LiteRuntimePrerequisiteDocsTests derives Lite's .NET runtime prerequisites + # from the BUILT runtimeconfig and from the publish shapes in this very file, then + # asserts these two READMEs say so - which artifact is self-contained, which needs two + # runtimes. Naming them here is not optional. Every area filter carves markdown out + # (dir/**/!(*.md)), and a docs-only PR additionally engages the fast path above, which + # skips .NET setup and restore outright - so without these two entries the guard could + # not run on a README-only change, i.e. on exactly the edit that would undo the fix. + # The cost is that a root-README edit now runs the Lite suite; that is the same trade + # the Darling entries below already take, and the alternative is a guard that quietly + # stops guarding. + - 'README.md' + - 'Lite/README.md' + # Same silently-stops-guarding reason as the darling filter's Lite entries below: + # Lite.Tests/ThemeParityLiteDarlingTests.cs READS the Darling viewer's theme + # dictionaries to assert the two apps' shared brush keys still resolve to the same + # colors. A Darling-theme-only edit is exactly the drift that guard exists to catch, + # so it has to reach the suite. + - 'Darling/PerformanceMonitor.Darling.Viewer/Themes/*.xaml' + # Same reason again, and a whole tree this time: WatermarkPolicyTests' + # TheCatchUpHorizon_IsWrittenDownInExactlyOnePlace (#2468) scans the Darling SERVICE for + # comments that restate the catch-up horizon's number instead of pointing at + # WatermarkPolicy.MaxCatchup. Three of the twelve sites it was written for live here, so a + # Darling-only PR reintroducing one would fire `darling` but not `lite` and skip the step + # that runs the guard — caught a day later by the nightly, which is exactly the silent + # drift the pin exists to close. The cost is that Darling service PRs now also run the Lite + # suite; that is the same trade the `darling` filter already takes for Lite/**/*.xaml. + - 'Darling/PerformanceMonitor.Darling.Service/**/!(*.md)' + installer: + - 'deprecated/Installer/**/!(*.md)' + - 'deprecated/Installer.Tests/**/!(*.md)' + - 'install/**/!(*.md)' + - 'upgrades/**/!(*.md)' + darling: + - 'Darling/**/!(*.md)' + # nightly.yml is not a build input, but Darling.Tests PARSES it: the #1888 + # guard reads both workflows' throwaway-cluster settings and compares them + # against the product's worker-sizing formula. Without this, a nightly-only + # edit would change a file the guard asserts on while never running the + # guard — a guard that silently stops guarding, which is the exact failure + # mode the source-parsing tests here exist to prevent. + - '.github/workflows/nightly.yml' + # Same reason, Lite side: the #1949 pin in Darling.Tests asserts every twinned + # query grid carries the SAME column sequence in both front ends, so it reads + # these six Lite files. A Lite-only XAML edit has to reach the suite or the + # parity half of that guard stops guarding. + - 'Lite/Controls/ServerTab.xaml' + - 'Lite/Controls/FinOpsTab.xaml' + - 'Lite/Windows/WaitDrillDownWindow.xaml' + - 'Lite/Windows/ProcedureHistoryWindow.xaml' + - 'Lite/Windows/QueryStatsHistoryWindow.xaml' + - 'Lite/Windows/QueryStoreHistoryWindow.xaml' + # #2114: XamlStaticResourceHygieneTests scans EVERY Lite XAML file — a StaticResource + # regression in one outside the six named above must still trigger the Darling job + # that runs the guard, or it slips to the nightly. + - 'Lite/**/*.xaml' + # The DOCUMENTATION allowlist: files that cannot affect a build under any + # job in this workflow. Deliberately an allowlist of non-executable content, + # not a "everything that isn't code" subtraction — a new file type defaults + # to being treated as code, which is the safe direction to be wrong in. + # + # NOT here, on purpose: *.sql (the installer and sql-validation compile it), + # *.yml (workflows), *.csproj / *.props / packages.lock.json (build inputs), + # and *.cs regardless of how comment-only the change looks — an XML doc + # comment still recompiles, and the compiler is what proves it still builds. + # + # The docs/ and Screenshots/ entries are extension-explicit rather than bare + # directory globs for the same reason: everything in them today is markdown, + # SVG, or a screenshot image, and a .sql or script dropped into either + # directory tomorrow should default to being code, not inherit a free pass + # from its parent directory. + docs: + - '**/*.md' + - 'LICENSE' + - 'CITATION.cff' + - '.gitignore' + - '.gitattributes' + - 'docs/**/*.{md,svg,png,jpg,jpeg,gif}' + - 'Screenshots/**/*.{md,svg,png,jpg,jpeg,gif}' + # Catch-all COUNTER, not a boolean gate: the classify step below decides + # "documentation-only" by comparing all_count to docs_count — they are equal + # exactly when every changed file sits on the docs allowlist. Stated as a + # count comparison because the previous shape ('**' plus '!' exclusions, + # a code: filter) could never be false under predicate-quantifier 'some' — + # every file matches '**', so the #1712 fast path shipped unable to engage + # (throwaway PR #1714: a .gitignore-only diff still paid setup + restore and, + # via the predicate bug above, a full build). + all: + - '**' + + # Decides the docs fast path ONCE, in one place, and says so out loud. Guards keep + # it off every path where a skipped restore would be a real loss: + # release — the filter step does not even run there, and a release must always + # compile and publish from a cold, fully restored tree. + # push — dev/main pushes are the integration signal for what just merged, so + # they restore unconditionally even for a docs-only commit. Cheap + # insurance: this only forces the restore back on, it does not force + # the per-product build/test steps, which stay path-gated as before. + # merge_group — a queue run is the LAST validation before its result lands on dev, + # so it takes the same always-restore path as a push. + # areas — belt and suspenders: even when the counts say docs-only, any lit + # area filter vetoes the fast path, because an area=true with restore + # skipped would run `dotnet build --no-restore` against nothing. The + # two classifications are built from the same allowlist so they cannot + # disagree today; this guard is for the day someone edits one and not + # the other. + # Everything else (pull_request) is eligible, and engages only when EVERY changed + # file is on the documentation allowlist (all_count == docs_count). + - name: Classify change for the docs fast path + id: fastpath + shell: bash + env: + ALL_COUNT: ${{ steps.filter.outputs.all_count }} + DOCS_COUNT: ${{ steps.filter.outputs.docs_count }} + DOCS_FILES: ${{ steps.filter.outputs.docs_files }} + AREAS: 'root=${{ steps.filter.outputs.root }} core=${{ steps.filter.outputs.core }} installer_core=${{ steps.filter.outputs.installer_core }} dashboard=${{ steps.filter.outputs.dashboard }} lite=${{ steps.filter.outputs.lite }} installer=${{ steps.filter.outputs.installer }} darling=${{ steps.filter.outputs.darling }}' + run: | + set -euo pipefail + + if [ "${{ github.event_name }}" = "release" ]; then + echo "engaged=false" >> "$GITHUB_OUTPUT" + echo "::notice title=Full build::Release event - the docs fast path never applies to a release." + exit 0 + fi + + if [ "${{ github.event_name }}" = "push" ] || [ "${{ github.event_name }}" = "merge_group" ]; then + echo "engaged=false" >> "$GITHUB_OUTPUT" + echo "::notice title=Full build::${{ github.event_name }} on '${{ github.ref_name }}' - integration runs always restore, even for a docs-only change." + exit 0 + fi + + echo "Changed files: ${ALL_COUNT:-0} total, ${DOCS_COUNT:-0} on the documentation allowlist. Areas: ${AREAS}" + + if [ "${ALL_COUNT:-0}" -gt 0 ] && [ "${ALL_COUNT:-0}" -eq "${DOCS_COUNT:-0}" ] && [[ "${AREAS}" != *"=true"* ]]; then + echo "engaged=true" >> "$GITHUB_OUTPUT" + echo "::notice title=DOCS FAST PATH ENGAGED::All ${ALL_COUNT} changed files are on the documentation allowlist, so .NET setup, restore and versioning are skipped. This job still reports its result." + echo "Documentation files classified in this change:" + for f in ${DOCS_FILES}; do echo " - ${f}"; done + else + echo "engaged=false" >> "$GITHUB_OUTPUT" + echo "::notice title=Full build::At least one changed file is off the documentation allowlist (${DOCS_COUNT:-0} of ${ALL_COUNT:-0} classified as documentation)." + fi + + - name: Setup .NET 10.0 + if: steps.fastpath.outputs.engaged != 'true' + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + - name: Restore dependencies + if: steps.fastpath.outputs.engaged != 'true' + run: | + dotnet restore Lite/PerformanceMonitorLite.csproj --locked-mode + dotnet restore Lite.Tests/Lite.Tests.csproj --locked-mode + dotnet restore deprecated/Installer.Tests/Installer.Tests.csproj --locked-mode + dotnet restore deprecated/Dashboard.Tests/Dashboard.Tests.csproj --locked-mode + dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode + dotnet restore Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj --locked-mode + + - name: Build Lite.Tests + if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet build Lite.Tests/Lite.Tests.csproj -c Release --no-restore + + - name: Build Installer.Tests + if: steps.filter.outputs.installer == 'true' || steps.filter.outputs.installer_core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet build deprecated/Installer.Tests/Installer.Tests.csproj -c Release --no-restore + + # The 'dashboard' path filter was defined when the Full Dashboard moved to deprecated/ (#1612) but + # never wired to a step, so its build and tests silently stopped running — which is how a batch of + # compiler warnings and three broken ThemeParityTests accumulated unnoticed (#1643). Deprecated means + # bug-fix-only, not unverified: it still compiles warning-free and its tests still guard cross-app + # parity (the theme palettes it checks are LITE's too). + - name: Build Dashboard.Tests + if: steps.filter.outputs.dashboard == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet build deprecated/Dashboard.Tests/Dashboard.Tests.csproj -c Release --no-restore + + - name: Build Darling + if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: | + dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore + dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release --no-restore + + # One step for the whole Lite suite. It was split into fast / analysis-heavy halves when the + # seven analysis classes rebuilt the full DuckDB schema inside every test and their subset + # alone cost ~9 minutes; after the shared class fixtures (#1693, #1698) and batched seeding + # (#1694) that subset runs in ~1 minute, so the split — and the narrower lite_analysis path + # gate that let non-analysis Lite changes skip it — stopped earning its second test-host + # spin-up and its filter-drift risk. + - name: Run Lite tests + if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet run --project Lite.Tests/Lite.Tests.csproj -c Release --no-build + + - name: Run Installer tests + if: steps.filter.outputs.installer == 'true' || steps.filter.outputs.installer_core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet run --project deprecated/Installer.Tests/Installer.Tests.csproj -c Release --no-build -- -class- "Installer.Tests.VersionDetectionTests" -class- "Installer.Tests.IdempotencyTests" -class- "Installer.Tests.AdversarialTests" + + - name: Run Dashboard tests + if: steps.filter.outputs.dashboard == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet run --project deprecated/Dashboard.Tests/Dashboard.Tests.csproj -c Release --no-build + + - name: Run Darling tests + if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build + + - name: Get version + if: steps.fastpath.outputs.engaged != 'true' + id: version + shell: pwsh + run: | + $version = ([xml](Get-Content Lite/PerformanceMonitorLite.csproj)).Project.PropertyGroup.Version | Where-Object { $_ } + echo "VERSION=$version" >> $env:GITHUB_OUTPUT + + # #2501: -r win-x64 --self-contained, and it makes the portable ZIP SMALLER, not bigger. + # The RID-agnostic framework-dependent publish this replaced copied every platform its + # packages ship - 537 MB of runtimes\ on a 565 MB tree, of which only the 52 MB win-x64 + # folder can ever load on Windows (DuckDB.NET.Bindings.Full is most of it, SkiaSharp and + # SqlClient behind it). Pinning the RID drops ~485 MB of unloadable native payload, which + # is far more than the bundled .NET/WPF/ASP.NET runtime adds back. Measured on the same + # commit and SDK: 565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB. + # + # The size was the secondary reason. The primary one is #2489: a framework-dependent ZIP + # unzipped onto a stock Windows Server dies on the .NET host's own "You must install .NET + # to run this application" before a line of our code runs, so Lite cannot report it, and + # #2499 could only ship a READ-ME-FIRST.txt next to the exe. Self-contained there is + # nothing to install and nothing to fail. That matters most for the nightly ZIP, which is + # the UAT download and is not offered as a Setup.exe at all. + # + # Lite/PerformanceMonitorLite.csproj declares RuntimeIdentifiers so the committed lock + # file covers this RID; see the comment there before removing either half. + - name: Publish Lite + if: steps.filter.outputs.lite == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -r win-x64 --self-contained -o publish/Lite + + - name: Publish Lite (self-contained for Velopack) + if: github.event_name == 'release' + run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -r win-x64 --self-contained -o publish/Lite-velopack + + - name: Publish Darling Service + if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o publish/DarlingService + + - name: Publish Darling Viewer + if: steps.filter.outputs.darling == 'true' || steps.filter.outputs.core == 'true' || steps.filter.outputs.root == 'true' || github.event_name == 'release' + run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -o publish/DarlingViewer + + - name: Publish Darling Viewer (self-contained for Velopack) + if: github.event_name == 'release' + run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -r win-x64 --self-contained -o publish/DarlingViewer-velopack + + # Darling bundles a PostgreSQL 18 + TimescaleDB runtime (pg-runtime.zip) that ships beside + # the service exe; DarlingManagedPostgres extracts it on first run. The fetch script pulls + # ~340MB of pinned EDB/TimescaleDB archives, so this is release-only and cached. The key is + # the fetch script's own content hash (the SHA256 pins live inside it): a re-release with + # unchanged pins restores the assembled zip and skips both the download and the assembly, + # and any pin/version bump edits the script and invalidates the cache automatically. + - name: Cache Darling pg-runtime.zip + if: github.event_name == 'release' + id: cache-pg-runtime + uses: actions/cache@v6 + with: + path: Darling/artifacts/pg-runtime.zip + key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} + + - name: Build Darling pg-runtime.zip + if: github.event_name == 'release' && steps.cache-pg-runtime.outputs.cache-hit != 'true' + shell: pwsh + run: ./Darling/tools/fetch-pg-runtime.ps1 + + - name: Package release artifacts + if: github.event_name == 'release' + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + New-Item -ItemType Directory -Force -Path releases + + # Lite ZIP - portable artifact for advanced/air-gapped users. The README points end + # users at Setup.exe (Velopack); this ZIP is the explicit fallback. + Compress-Archive -Path 'publish/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force + + # upload-artifact is deliberately HELD at v6 (#1653): every signing step below consumes + # `steps.upload-*.outputs.artifact-id`, and v7 changes artifact archiving semantics (the + # `archive` parameter). The signing path only executes on `release: [published]`, so a broken + # bump surfaces at release time — bump only alongside a validated real signing run. + # Dependabot is configured to skip this major (see .github/dependabot.yml). + - name: Upload Lite for signing + if: github.event_name == 'release' + id: upload-lite + uses: actions/upload-artifact@v6 + with: + name: Lite-unsigned + path: publish/Lite/ + + - name: Stage Darling for signing + if: github.event_name == 'release' + shell: pwsh + run: | + $stage = 'publish/Darling-signing' + if (Test-Path $stage) { Remove-Item -Recurse -Force $stage } + New-Item -ItemType Directory -Force -Path "$stage/viewer" | Out-Null + Copy-Item 'publish/DarlingService/*' $stage -Recurse + Copy-Item 'publish/DarlingViewer/*' "$stage/viewer" -Recurse + + - name: Upload Darling for signing + if: github.event_name == 'release' + id: upload-darling + uses: actions/upload-artifact@v6 + with: + name: Darling-unsigned + path: publish/Darling-signing/ + + - name: Sign Lite + if: github.event_name == 'release' + uses: signpath/github-action-submit-signing-request@v2 + with: + api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' + organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' + project-slug: 'PerformanceMonitor' + signing-policy-slug: 'release-signing' + artifact-configuration-slug: 'Lite' + github-artifact-id: '${{ steps.upload-lite.outputs.artifact-id }}' + wait-for-completion: true + output-artifact-directory: 'signed/Lite' + + - name: Sign Darling + if: github.event_name == 'release' + uses: signpath/github-action-submit-signing-request@v2 + with: + api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' + organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' + project-slug: 'PerformanceMonitor' + signing-policy-slug: 'release-signing' + artifact-configuration-slug: 'Darling' + github-artifact-id: '${{ steps.upload-darling.outputs.artifact-id }}' + wait-for-completion: true + output-artifact-directory: 'signed/Darling' + + - name: Replace with signed artifacts + if: github.event_name == 'release' + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + # Re-zip signed files into release archives + Remove-Item "releases/PerformanceMonitorLite-$version.zip" -ErrorAction SilentlyContinue + Compress-Archive -Path 'signed/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force + + - name: Package Darling (signed) + if: github.event_name == 'release' + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + # One product, one zip (mirrors the one-zip-per-product convention above). signed/Darling + # already holds the signed tree in its final layout — the service at the archive root (its + # darling.sample.json alongside), the viewer in a viewer\ subfolder. Drop pg-runtime.zip + # beside the service exe, exactly where DarlingManagedPostgres looks (AppContext.BaseDirectory) + # and extracts it on first run. The EDB PostgreSQL binaries inside pg-runtime.zip are shipped + # as opaque data and were never signed. + Copy-Item 'Darling/artifacts/pg-runtime.zip' 'signed/Darling' + + Remove-Item "releases/PerformanceMonitorDarling-$version.zip" -ErrorAction SilentlyContinue + Compress-Archive -Path 'signed/Darling/*' -DestinationPath "releases/PerformanceMonitorDarling-$version.zip" -Force + + # The Velopack (Setup.exe) path publishes a SEPARATE self-contained build + # (publish/Dashboard-velopack, publish/Lite-velopack -- see the "self-contained + # for Velopack" steps above) that previously went straight into `vpk pack` + # without ever being uploaded to SignPath. Only the framework-dependent trees + # used for the legacy ZIPs were signed, so every Setup.exe shipped unsigned + # since Velopack packaging was introduced. These steps close that gap by + # mirroring the exact upload/sign pattern used for Dashboard/Lite/Installer + # above, and vpk pack below now reads from the signed output. + - name: Upload Lite (Velopack) for signing + if: github.event_name == 'release' + id: upload-lite-velopack + uses: actions/upload-artifact@v6 + with: + name: Lite-Velopack-unsigned + path: publish/Lite-velopack/ + + - name: Upload Darling Viewer (Velopack) for signing + if: github.event_name == 'release' + id: upload-darlingviewer-velopack + uses: actions/upload-artifact@v6 + with: + name: DarlingViewer-Velopack-unsigned + path: publish/DarlingViewer-velopack/ + + - name: Sign Lite (Velopack) + if: github.event_name == 'release' + uses: signpath/github-action-submit-signing-request@v2 + with: + api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' + organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' + project-slug: 'PerformanceMonitor' + signing-policy-slug: 'release-signing' + artifact-configuration-slug: 'Lite' + github-artifact-id: '${{ steps.upload-lite-velopack.outputs.artifact-id }}' + wait-for-completion: true + output-artifact-directory: 'signed/Lite-Velopack' + + # The remote-seat viewer Setup.exe (#1555) is signed with its OWN 'DarlingViewer' artifact + # configuration — the co-located-zip 'Darling' slug signs a service+viewer\ tree layout, which + # does not match this self-contained viewer-at-root publish. Like the 'Darling' slug, the + # 'DarlingViewer' config (which files get signed) lives outside this repo on signpath.io; until + # Erik creates the slug this step fails the release — the standing #1340 SignPath prerequisite. + - name: Sign Darling Viewer (Velopack) + if: github.event_name == 'release' + uses: signpath/github-action-submit-signing-request@v2 + with: + api-token: '${{ secrets.SIGNPATH_API_TOKEN }}' + organization-id: '7969f8b6-d946-4a74-9bac-a55856d8b8e0' + project-slug: 'PerformanceMonitor' + signing-policy-slug: 'release-signing' + artifact-configuration-slug: 'DarlingViewer' + github-artifact-id: '${{ steps.upload-darlingviewer-velopack.outputs.artifact-id }}' + wait-for-completion: true + output-artifact-directory: 'signed/DarlingViewer-Velopack' + + - name: Create Velopack releases (Lite + Darling Viewer) + if: github.event_name == 'release' + shell: pwsh + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + VERSION: ${{ steps.version.outputs.VERSION }} + run: | + # Pin vpk to the Velopack library version (keep in sync with the Velopack + # PackageReference in PerformanceMonitorLite.csproj). + dotnet tool install -g vpk --version 1.2.0 + New-Item -ItemType Directory -Force -Path releases/velopack-lite + New-Item -ItemType Directory -Force -Path releases/velopack-darlingviewer + + # Lite: download previous + pack (from the SIGNED velopack output, not the raw publish dir) + vpk download github --repoUrl https://github.com/${{ github.repository }} --channel lite -o releases/velopack-lite --token $env:GH_TOKEN + vpk pack -u PerformanceMonitorLite -v $env:VERSION -p signed/Lite-Velopack -e PerformanceMonitorLite.exe -o releases/velopack-lite --channel lite + + # Darling Viewer remote-seat installer (#1555): download previous + pack (from the SIGNED + # velopack output). Its own 'darlingviewer' channel/delta feed, separate from the co-located + # viewer inside PerformanceMonitorDarling-*.zip (which stays plain-zip only). + vpk download github --repoUrl https://github.com/${{ github.repository }} --channel darlingviewer -o releases/velopack-darlingviewer --token $env:GH_TOKEN + vpk pack -u PerformanceMonitorDarlingViewer -v $env:VERSION -p signed/DarlingViewer-Velopack -e PerformanceMonitor.Darling.Viewer.exe -o releases/velopack-darlingviewer --channel darlingviewer + + - name: Generate checksums + if: github.event_name == 'release' + shell: pwsh + run: | + $checksums = Get-ChildItem releases/*.zip | ForEach-Object { + $hash = (Get-FileHash $_.FullName -Algorithm SHA256).Hash.ToLower() + "$hash $($_.Name)" + } + # LF and no BOM, written explicitly (#2383). Out-File on Windows emits CRLF whatever + # the encoding, and `shasum -c` on macOS/Linux then takes the trailing \r as part of + # the FILENAME - every entry reports "No such file or directory / FAILED open or + # read", which is precisely what a tampered download looks like. The hashes were + # always right; the file just could not be used by the tool it exists for, on the two + # platforms where that tool is the default. A BOM breaks the first line the same way. + # + # There is a SECOND copy of this step in nightly.yml. #2384 fixed only this one, and the + # nightly went on publishing CRLF for another day; keep them in step. + [IO.File]::WriteAllText( + "$PWD/releases/SHA256SUMS.txt", + (($checksums -join "`n") + "`n"), + [Text.UTF8Encoding]::new($false)) + Write-Host "Checksums:" + $checksums | ForEach-Object { Write-Host $_ } + + - name: Upload release assets + if: github.event_name == 'release' + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: | + gh release upload ${{ github.event.release.tag_name }} releases/*.zip releases/SHA256SUMS.txt --clobber + + - name: Upload Lite Velopack artifacts + if: github.event_name == 'release' + shell: pwsh + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + VERSION: ${{ steps.version.outputs.VERSION }} + run: | + vpk upload github --repoUrl https://github.com/${{ github.repository }} --channel lite -o releases/velopack-lite --releaseName "v$env:VERSION" --tag "v$env:VERSION" --merge --token $env:GH_TOKEN + + - name: Upload Darling Viewer Velopack artifacts + if: github.event_name == 'release' + shell: pwsh + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + VERSION: ${{ steps.version.outputs.VERSION }} + run: | + vpk upload github --repoUrl https://github.com/${{ github.repository }} --channel darlingviewer -o releases/velopack-darlingviewer --releaseName "v$env:VERSION" --tag "v$env:VERSION" --merge --token $env:GH_TOKEN + + # #1587: gated-live Darling coverage BEFORE merge, not only in the nightly. Darling.Tests has + # live-PostgreSQL tests (the *_AgainstDevPostgres classes) gated on DARLING_TEST_PG, which the + # build job above never sets — so those tests only ever ran post-merge in nightly.yml. That is + # exactly how #1586's alter_job(bigint) bug merged AND deployed clean: the test that catches it is + # gated-live, and PR CI did not run it. This job stands up a throwaway PostgreSQL + TimescaleDB + # from the bundled pg-runtime and runs the FULL Darling suite against it — but ONLY when Darling + # code (or this workflow) changed, so a Lite/Dashboard-only PR pays nothing. On a non-Darling + # change every step below no-ops via the path filter and the job still reports SUCCESS, so it can + # be made a required check without blocking unrelated PRs (the same always-runs-reports-a-result + # shape the build job uses for doc-only changes). Mirrors the nightly darling-pg job step-for-step; + # the one live-SQL-Server E2E stays skipped (DARLING_TEST_SQL unset — no SQL Server on the runner). + darling-pg: + name: Darling PostgreSQL tests + runs-on: windows-latest + # Max observed on a warm cache is ~3m40s; a cold pg-runtime cache adds a ~340MB fetch. + # 30 minutes is 3x headroom over the cold path — past that, something is hung (pg_ctl -w + # waiting on a cluster that will never come up), and the default 6h timeout would hold a + # shared-pool Windows runner hostage for the duration. The build job above deliberately + # has NO timeout: on release it waits on SignPath's manual approval gate, which can + # legitimately take hours. + timeout-minutes: 30 + permissions: + contents: read + + steps: + - uses: actions/checkout@v7 + + # Only do the expensive TimescaleDB work when Darling code changed — or when THIS workflow + # changed, so a change to the gate itself is exercised by the gate (this is what makes the PR + # that introduces this job validate itself end-to-end). Doc-only Darling edits don't trigger + # it. Skipped entirely on release: the dev push that produced the release commit already ran + # it, so the filter step doesn't run and every step below no-ops. + - name: Detect changed paths + id: filter + if: github.event_name != 'release' + uses: dorny/paths-filter@v4 + with: + # On push, compare against the previous commit on this branch (mirrors the build job); + # on pull_request, an empty base makes the action diff against the PR base branch. + base: ${{ github.event_name == 'push' && github.event.before || '' }} + # `Darling/**/!(*.md)` instead of a `Darling/**` include plus a `!Darling/**/*.md` + # exclude: dorny v4 treats each pattern as an independent predicate under the + # default quantifier, so the old bare negation was itself a match-all-non-Darling-md + # rule — this job ran the full TimescaleDB suite on every PR, including md-only + # ones (run 30218459544: "Filter darling = true, Matching files: CHANGELOG.md"). + filters: | + darling: + - 'Darling/**/!(*.md)' + - '.github/workflows/build.yml' + # Same reason as the build job's darling filter: the #1888 cluster-sizing + # guard parses nightly.yml, so an edit to it has to reach the suite. + - '.github/workflows/nightly.yml' + + # This job's gate was already correct for documentation — a docs-only change leaves + # 'darling' false and every step below no-ops. What it lacked was SAYING so: a job + # that reports success having quietly run nothing looks identical to one that tested + # everything. Costs one step; buys a log you can point at when asking "did this + # actually get tested?". + - name: Report the Darling PG gate decision + shell: bash + run: | + set -euo pipefail + + if [ "${{ github.event_name }}" = "release" ]; then + echo "::notice title=Darling PG tests skipped::Release event - the dev push that produced this commit already ran them." + elif [ "${{ steps.filter.outputs.darling }}" = "true" ]; then + echo "::notice title=Darling PG tests running::Darling code (or this workflow) changed." + else + echo "::notice title=Darling PG tests skipped::No Darling code changed - documentation-only Darling edits do not trigger the TimescaleDB suite." + fi + + - name: Setup .NET 10.0 + if: steps.filter.outputs.darling == 'true' + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + # Same cache key the release job and nightly.yml use (the fetch script's own content hash), so + # a warm cache from any of the three means no ~340MB EDB/TimescaleDB download here. + - name: Cache Darling pg-runtime.zip + if: steps.filter.outputs.darling == 'true' + id: cache-pg-runtime + uses: actions/cache@v6 + with: + path: Darling/artifacts/pg-runtime.zip + key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} + + - name: Build Darling pg-runtime.zip (cache miss only) + if: steps.filter.outputs.darling == 'true' && steps.cache-pg-runtime.outputs.cache-hit != 'true' + shell: pwsh + run: ./Darling/tools/fetch-pg-runtime.ps1 + + - name: Extract pg-runtime + if: steps.filter.outputs.darling == 'true' + shell: pwsh + run: | + $zip = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime.zip" + $dest = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime" + if (Test-Path $dest) { Remove-Item -Recurse -Force $dest } + Add-Type -AssemblyName System.IO.Compression.FileSystem + [System.IO.Compression.ZipFile]::ExtractToDirectory($zip, $dest) + if (-not (Test-Path "$dest\pgsql\bin\pg_ctl.exe")) { throw "pg-runtime missing pgsql\bin\pg_ctl.exe" } + + - name: Restore Darling.Tests + if: steps.filter.outputs.darling == 'true' + run: dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode + + - name: Build Darling.Tests + if: steps.filter.outputs.darling == 'true' + run: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore + + # Throwaway cluster from the bundled runtime. initdb TRUST auth is acceptable ONLY here: an + # ephemeral CI runner, loopback-only, throwaway data (the product default is scram-sha-256 + + # a generated credential). Superuser is "darling" (not "postgres") so CREATE SCHEMA ... + # AUTHORIZATION darling in the V8 schema-split migration resolves. Settings mirror + # DarlingManagedPostgres.BuildConfAppend (timescaledb preload, port 5541, loopback bind) + # and BuildWorkerSizingConfAppend (the two worker settings below). + # + # #1888: the worker settings are NOT optional garnish. PostgreSQL's default + # max_worker_processes = 8 cannot launch TimescaleDB's per-hypertable compression, + # retention and continuous-aggregate policy jobs — the postmaster logs "failed to start a + # background worker" storms and most policy runs never happen. Without them this job tested a + # configuration NO customer runs (the product refuses to: BuildWorkerSizingConfAppend writes + # these on every managed start), and worse, made failures luck-of-the-slot rather than + # reproducible — which is exactly how #1862's compression flake failed twice on CI in two + # unrecognizably different ways and could not be reproduced locally at all. + # + # The values are the product's own derivation from the live hypertable count + # (TimescaleSupport.HypertableCount = the 62-collector catalog + collection_log = 63): + # timescaledb.max_background_workers = HypertableCount + 2 = 71 + # max_worker_processes = 3 + (HypertableCount + 2) + 8 = 82 + # Hard-coded here because a workflow cannot call into the product — so + # CiClusterWorkerSizingTests parses THIS FILE and fails the build if either number stops + # matching the formula as collectors are added, and CiClusterWorkerSizingLiveTests asserts + # the running cluster actually serves them (a conf line that never took effect is + # indistinguishable from one that did, by inspection). + - name: Initialize and start throwaway PostgreSQL + if: steps.filter.outputs.darling == 'true' + shell: pwsh + run: | + $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" + $dataDir = "$env:RUNNER_TEMP\darling-pgdata" + $logFile = "$env:RUNNER_TEMP\darling-pg.log" + & "$bin\initdb.exe" -D $dataDir -U darling -A trust --encoding=UTF8 + if ($LASTEXITCODE -ne 0) { throw "initdb failed ($LASTEXITCODE)" } + Add-Content -Path "$dataDir\postgresql.conf" -Value "shared_preload_libraries = 'timescaledb'" + Add-Content -Path "$dataDir\postgresql.conf" -Value "port = 5541" + Add-Content -Path "$dataDir\postgresql.conf" -Value "listen_addresses = '127.0.0.1'" + Add-Content -Path "$dataDir\postgresql.conf" -Value "timescaledb.max_background_workers = 71" + Add-Content -Path "$dataDir\postgresql.conf" -Value "max_worker_processes = 82" + & "$bin\pg_ctl.exe" -D $dataDir -l $logFile -w start + if ($LASTEXITCODE -ne 0) { if (Test-Path $logFile) { Get-Content $logFile -Tail 50 }; throw "pg_ctl start failed ($LASTEXITCODE)" } + & "$bin\createdb.exe" -h 127.0.0.1 -p 5541 -U darling darling + if ($LASTEXITCODE -ne 0) { throw "createdb failed ($LASTEXITCODE)" } + + # DARLING_TEST_PG lights up the [Collection("live-postgres")] classes; DARLING_TEST_PGRUNTIME + # lights up the managed-bootstrap E2E. DARLING_TEST_SQL is intentionally unset — the one + # live-SQL-Server E2E stays skipped (no SQL Server on the runner). Full suite (not a filtered + # subset): the ungated ~3s overlap with the build job is negligible and a filter could hide a + # test the way narrow filters have bitten before. + - name: Run Darling PG tests + if: steps.filter.outputs.darling == 'true' + shell: pwsh + env: + DARLING_TEST_PG: "Host=127.0.0.1;Port=5541;Username=darling;Database=darling" + DARLING_TEST_PGRUNTIME: ${{ github.workspace }}\Darling\artifacts\pg-runtime + run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build -- -trx TestResults/darling-pr.trx + + - name: Stop PostgreSQL + if: always() && steps.filter.outputs.darling == 'true' + shell: pwsh + run: | + $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" + $dataDir = "$env:RUNNER_TEMP\darling-pgdata" + if (Test-Path "$bin\pg_ctl.exe") { & "$bin\pg_ctl.exe" -D $dataDir -m fast -w stop } + exit 0 + + - name: Upload PG log and test results on failure + if: failure() && steps.filter.outputs.darling == 'true' + uses: actions/upload-artifact@v6 + with: + name: darling-pg-failure + path: | + ${{ runner.temp }}/darling-pg.log + TestResults/ + if-no-files-found: ignore + + # ── Linux service build + container image (#1804) ──────────────────────────────────────────────── + # The Darling service is cross-platform .NET on purpose, but until this job nothing PROVED it on + # every PR — the linux-x64 publish and the container image both built for the first time at release + # time or never. Same path-filter shape as darling-pg above: only runs the expensive work when + # Darling/service code (or this workflow, or the Dockerfile) changed, always reports a result so it + # can be a required check. No tests run here — the test projects are net10.0-windows (they reference + # the WPF apps); the cross-platform behavior they pin is exercised by the Windows jobs, and the + # container smoke lives in the compose quickstart. On PRs/pushes this job answers exactly two + # questions: does the service still publish for linux-x64, and does the image still build. + # + # On the RELEASE event this job is the Linux PUBLISHER: before it, a stable release shipped Windows + # zips and Setup.exes while the linux tar.gz and the ghcr image existed only at nightly quality + # (nightly.yml, tag :nightly) — the compose quickstart pointed released users at a nightly image. + # Now the release uploads PerformanceMonitorDarling-linux-x64-.tar.gz + SHA256SUMS-linux.txt to + # the release and pushes ghcr : and :latest. Linux binaries are NOT SignPath-signed + # (SignPath signs Windows PEs); the checksums file is the integrity story, same as the nightly. + darling-linux: + name: Darling Linux build + runs-on: ubuntu-latest + timeout-minutes: 30 + permissions: + contents: write + packages: write + # Sigstore keyless signing + GitHub provenance (proven on the nightly first): id-token yields + # the OIDC identity Fulcio certifies against; attestations stores the tarball's provenance. + id-token: write + attestations: write + + steps: + - uses: actions/checkout@v7 + + - name: Detect changed paths + id: filter + if: github.event_name != 'release' + uses: dorny/paths-filter@v4 + with: + base: ${{ github.event_name == 'push' && github.event.before || '' }} + filters: | + darling: + - 'Darling/**/!(*.md)' + - 'PerformanceMonitor.Common/**' + - 'PerformanceMonitor.Collectors/**' + - 'PerformanceMonitor.Analysis/**' + - '.github/workflows/build.yml' + + - name: Report the Linux gate decision + shell: bash + run: | + set -euo pipefail + if [ "${{ github.event_name }}" = "release" ]; then + echo "::notice title=Darling Linux publishing::Release event - packaging the versioned linux tar.gz and pushing the ghcr image." + elif [ "${{ steps.filter.outputs.darling }}" = "true" ]; then + echo "::notice title=Darling Linux build running::Darling/service code (or this workflow) changed." + else + echo "::notice title=Darling Linux build skipped::No Darling/service code changed." + fi + + - name: Setup .NET 10.0 + if: github.event_name == 'release' || steps.filter.outputs.darling == 'true' + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + - name: Publish service (linux-x64) + if: github.event_name == 'release' || steps.filter.outputs.darling == 'true' + run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -r linux-x64 --self-contained false -o publish/DarlingService-linux + + - name: Build container image + if: github.event_name != 'release' && steps.filter.outputs.darling == 'true' + run: docker build -f Darling/Dockerfile -t performancemonitor-darling:pr . + + # ── Release-only publishing (mirrors nightly.yml's linux job, versioned instead of :nightly) ── + + - name: Get version + if: github.event_name == 'release' + id: version + shell: bash + run: | + set -euo pipefail + version="$(grep -oPm1 '(?<=)[^<]+' Lite/PerformanceMonitorLite.csproj)" + echo "VERSION=${version}" >> "$GITHUB_OUTPUT" + + - name: Package linux artifact + checksum + if: github.event_name == 'release' + shell: bash + run: | + set -euo pipefail + version="${{ steps.version.outputs.VERSION }}" + mkdir -p releases + tar -C publish/DarlingService-linux -czf "releases/PerformanceMonitorDarling-linux-x64-${version}.tar.gz" . + (cd releases && sha256sum "PerformanceMonitorDarling-linux-x64-${version}.tar.gz" > SHA256SUMS-linux.txt && cat SHA256SUMS-linux.txt) + + - name: Upload linux artifact to the release + if: github.event_name == 'release' + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: gh release upload ${{ github.event.release.tag_name }} releases/PerformanceMonitorDarling-linux-x64-*.tar.gz releases/SHA256SUMS-linux.txt --clobber + + - name: Build and push container image (ghcr, versioned + latest) + if: github.event_name == 'release' + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + shell: bash + run: | + set -euo pipefail + version="${{ steps.version.outputs.VERSION }}" + image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" + echo "$GH_TOKEN" | docker login ghcr.io -u "${{ github.actor }}" --password-stdin + docker build -f Darling/Dockerfile -t "${image}:${version}" -t "${image}:latest" . + docker push "${image}:${version}" + docker push "${image}:latest" + + # Keyless Sigstore signature on the released image — the same steps the nightly runs (verified + # end-to-end from a client 2026-08-06: cosign validated the claims, the Rekor log entry, and the + # workflow identity). SignPath's cosign support is edition-gated, and this path needs no keys or + # subscription at all. Verify: + # cosign verify ghcr.io/erikdarlingdata/performancemonitor-darling: \ + # --certificate-identity-regexp 'github.com/erikdarlingdata/PerformanceMonitor' \ + # --certificate-oidc-issuer https://token.actions.githubusercontent.com + - name: Install cosign + if: github.event_name == 'release' + uses: sigstore/cosign-installer@v3 + + - name: Sign container image (keyless, Sigstore) + if: github.event_name == 'release' + shell: bash + run: | + set -euo pipefail + version="${{ steps.version.outputs.VERSION }}" + image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" + digest="$(docker inspect --format='{{index .RepoDigests 0}}' "${image}:${version}")" + cosign sign --yes "${digest}" + + # GitHub-native SLSA provenance for the released tarball: gh attestation verify -R . + - name: Attest the linux tarball (GitHub provenance) + if: github.event_name == 'release' + uses: actions/attest-build-provenance@v4 + with: + subject-path: releases/PerformanceMonitorDarling-linux-x64-*.tar.gz diff --git a/.github/workflows/nightly.yml b/.github/workflows/nightly.yml index 2ae97572a..f0855151a 100644 --- a/.github/workflows/nightly.yml +++ b/.github/workflows/nightly.yml @@ -1,487 +1,524 @@ -name: Nightly Build - -on: - schedule: - # 6:00 AM UTC (1:00 AM EST / 2:00 AM EDT) - - cron: '0 6 * * *' - workflow_dispatch: # manual trigger — and the vehicle the scheduled re-dispatch below rides - inputs: - from_schedule: - description: 'Set true by the scheduled re-dispatch so the 24h new-commit check applies. Leave false for manual runs, which always build.' - type: boolean - required: false - default: false - -permissions: - contents: write - # #1804: the linux job pushes the nightly container image to ghcr. - packages: write - # Sigstore keyless signing + GitHub provenance for the linux artifacts: id-token lets the job - # obtain its OIDC identity (Fulcio issues the short-lived signing cert against it), attestations - # lets attest-build-provenance store the tarball's provenance. Neither grants anything else. - id-token: write - attestations: write - -jobs: - # Scheduled workflows always execute the DEFAULT branch's copy of this file, while nightly - # artifacts deliberately build from dev's tree. That skew is how the 2026-07-26 nightly - # failed (run 30194606068): main's stale copy still read Dashboard/Dashboard.csproj, a path - # #1612 moved to deprecated/ on dev, so 'Set nightly version' died on a file missing from - # the tree it had just checked out — and the same trap bit before (#1550/#1551). The cure - # is structural, not another sync: on schedule this workflow does NOTHING but re-dispatch - # itself onto the dev REF, because a workflow_dispatch run executes the dispatched ref's - # copy of this file — dev's, current by definition. Once main carries this shape, its copy - # has exactly one job that must keep working, and that job references no tree paths at - # all; every future change to the real nightly logic lands on dev and takes effect the - # night it merges, no promotion to main needed. GITHUB_TOKEN can create workflow_dispatch - # runs (the Actions recursion guard exempts workflow_dispatch and repository_dispatch), - # and the dispatched run cannot loop back here because it arrives as workflow_dispatch, - # not schedule. Until main is synced once, the scheduled run still executes main's OLD - # copy and keeps failing nightly — the one-time sync is in the PR that introduced this. - redispatch: - if: github.event_name == 'schedule' - runs-on: ubuntu-latest - timeout-minutes: 5 - permissions: - actions: write - steps: - - name: Re-dispatch this workflow onto the dev ref - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - run: gh workflow run nightly.yml --repo ${{ github.repository }} --ref dev -f from_schedule=true - - # Everything below runs only in a workflow_dispatch run — the scheduled re-dispatch or a - # manual one — which executes from the dispatched ref (dev for the scheduled path). - check: - if: github.event_name == 'workflow_dispatch' - runs-on: ubuntu-latest - timeout-minutes: 10 - outputs: - has_changes: ${{ steps.check.outputs.has_changes }} - steps: - - uses: actions/checkout@v7 - with: - ref: dev - fetch-depth: 0 - - - name: Check for new commits in last 24 hours - id: check - run: | - RECENT=$(git log --since="24 hours ago" --oneline | head -1) - if [ -n "$RECENT" ]; then - echo "has_changes=true" >> $GITHUB_OUTPUT - echo "New commits found — building nightly" - else - echo "has_changes=false" >> $GITHUB_OUTPUT - echo "No new commits — skipping nightly build" - fi - - build: - needs: check - # Manual dispatches always build (from_schedule defaults false); the scheduled - # re-dispatch sets from_schedule=true and builds only when dev changed in the last - # 24h — the same policy the schedule applied when it ran these jobs directly. - if: needs.check.outputs.has_changes == 'true' || inputs.from_schedule != true - runs-on: windows-latest - # Full pipeline (restore, tests, four publishes, cold pg-runtime fetch, vpk pack, - # release upload) is well under an hour; 90 minutes means hung-not-slow. Nightly ships - # unsigned, so unlike build.yml's release path there is no manual signing gate to wait on. - timeout-minutes: 90 - - steps: - - uses: actions/checkout@v7 - with: - ref: dev - - - name: Setup .NET 10.0 - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - - name: Set nightly version - id: version - shell: pwsh - run: | - $base = ([xml](Get-Content Lite/PerformanceMonitorLite.csproj)).Project.PropertyGroup.Version | Where-Object { $_ } - $date = Get-Date -Format "yyyyMMdd" - $nightly = "$base-nightly.$date" - echo "VERSION=$nightly" >> $env:GITHUB_OUTPUT - echo "Nightly version: $nightly" - - - name: Restore dependencies - run: | - dotnet restore Lite/PerformanceMonitorLite.csproj --locked-mode - dotnet restore Lite.Tests/Lite.Tests.csproj --locked-mode - dotnet restore Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj --locked-mode - - - name: Run tests - run: dotnet run --project Lite.Tests/Lite.Tests.csproj -c Release - - - name: Publish Lite - run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -o publish/Lite - - - name: Publish Darling Service - run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o publish/DarlingService - - - name: Publish Darling Viewer - run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -o publish/DarlingViewer - - # Self-contained viewer publish that feeds the remote-seat Velopack Setup.exe (#1555), the same - # publish shape build.yml uses for the Dashboard/Lite Velopack packs. This is IN ADDITION to the - # framework-dependent "Publish Darling Viewer" above, which still feeds the co-located viewer\ - # folder inside PerformanceMonitorDarling-*.zip — that zip is unchanged. - - name: Publish Darling Viewer (self-contained for Velopack) - run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -r win-x64 --self-contained -o publish/DarlingViewer-velopack - - # Same cache key as build.yml's release path: the fetch script's content hash (the SHA256 - # pins live inside it). The ~340MB EDB/TimescaleDB fetch runs at most once per pin-set per - # branch; nightly and release runs share the assembled zip whenever the cache is visible. - - name: Cache Darling pg-runtime.zip - id: cache-pg-runtime - uses: actions/cache@v6 - with: - path: Darling/artifacts/pg-runtime.zip - key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} - - - name: Build Darling pg-runtime.zip - if: steps.cache-pg-runtime.outputs.cache-hit != 'true' - shell: pwsh - run: ./Darling/tools/fetch-pg-runtime.ps1 - - - name: Package artifacts - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - New-Item -ItemType Directory -Force -Path releases - - Compress-Archive -Path 'publish/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force - - - # Same layout as the release zip (build.yml "Package Darling (signed)"): service at the - # archive root with darling.sample.json alongside, viewer under viewer\, pg-runtime.zip - # beside the service exe where DarlingManagedPostgres extracts it on first run. Nightly - # zips are unsigned across the board, so this stages from publish/ instead of signed/. - $darlingDir = 'publish/Darling' - New-Item -ItemType Directory -Force -Path "$darlingDir/viewer" | Out-Null - Copy-Item 'publish/DarlingService/*' $darlingDir -Recurse - Copy-Item 'publish/DarlingViewer/*' "$darlingDir/viewer" -Recurse - Copy-Item 'Darling/artifacts/pg-runtime.zip' $darlingDir - - Compress-Archive -Path 'publish/Darling/*' -DestinationPath "releases/PerformanceMonitorDarling-$version.zip" -Force - - # Darling viewer remote-seat installer (#1555). Nightly ships it UNSIGNED like every other nightly - # artifact. Mirrors build.yml's release vpk pack (same pack id / exe / channel) but packs from the - # raw self-contained publish (no SignPath), and deliberately does NOT touch the Velopack update - # feed: the nightly GitHub release is deleted + recreated each night, so there is no persistent - # delta chain to `vpk download`/`vpk upload` from — we ship a standalone full Setup.exe as a plain - # release asset. Copied to a deterministic name so the checksum + upload steps below pick it up. - # Purely additive: the co-located viewer inside PerformanceMonitorDarling-*.zip is untouched. - - name: Create Darling Viewer Setup.exe (Velopack, unsigned) - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - dotnet tool install -g vpk --version 1.2.0 - New-Item -ItemType Directory -Force -Path releases/velopack-darlingviewer - vpk pack -u PerformanceMonitorDarlingViewer -v $version -p publish/DarlingViewer-velopack -e PerformanceMonitor.Darling.Viewer.exe -o releases/velopack-darlingviewer --channel darlingviewer - $setup = Get-ChildItem releases/velopack-darlingviewer/*Setup.exe | Select-Object -First 1 - Copy-Item $setup.FullName "releases/PerformanceMonitorDarlingViewer-$version-Setup.exe" - - - name: Generate checksums - shell: pwsh - run: | - $checksums = Get-ChildItem releases/*.zip, releases/*.exe | ForEach-Object { - $hash = (Get-FileHash $_.FullName -Algorithm SHA256).Hash.ToLower() - "$hash $($_.Name)" - } - $checksums | Out-File -FilePath releases/SHA256SUMS.txt -Encoding utf8 - Write-Host "Checksums:" - $checksums | ForEach-Object { Write-Host $_ } - - - name: Delete previous nightly release - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - run: gh release delete nightly --yes --cleanup-tag 2>$null; exit 0 - shell: pwsh - - - name: Create nightly release - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - shell: pwsh - run: | - $version = "${{ steps.version.outputs.VERSION }}" - $sha = git rev-parse --short HEAD - $body = @" - Automated nightly build from ``dev`` branch. - - **Version:** ``$version`` - **Commit:** ``$sha`` - **Built:** $(Get-Date -Format "yyyy-MM-dd HH:mm UTC") - - > These builds include the latest changes and may be unstable. - > For production use, download the [latest stable release](https://github.com/erikdarlingdata/PerformanceMonitor/releases/latest). - "@ - - gh release create nightly ` - --target dev ` - --title "Nightly Build ($version)" ` - --notes $body ` - --prerelease ` - releases/*.zip releases/*.exe releases/SHA256SUMS.txt - - # Darling.Tests has live-PostgreSQL tests gated on env vars that the per-PR/push build never - # sets, so regular CI runs only the ungated subset. This job stands up a throwaway PostgreSQL - # from the bundled pg-runtime and runs the full Darling suite against it nightly, so the - # Postgres + managed-bootstrap paths get real coverage. The one live-SQL-Server E2E stays - # skipped (no SQL Server on the runner) — expected. Gated like the build job so it only runs - # when dev actually changed (or on manual dispatch). - darling-pg: - name: Darling PostgreSQL tests - needs: check - # Same gating as the build job: manual dispatches always run, the scheduled - # re-dispatch (from_schedule=true) runs only when dev changed in the last 24h. - if: needs.check.outputs.has_changes == 'true' || inputs.from_schedule != true - runs-on: windows-latest - # Cold pg-runtime fetch + build + the full live-PG suite fits well inside an hour; a - # cluster that never comes up (pg_ctl -w) is the hang this bounds. - timeout-minutes: 60 - - steps: - # Scheduled runs always test dev (schedules execute from the default branch, so ref_name - # would be main). A manual dispatch tests the DISPATCHED ref — the only way to validate a - # branch's gated-pg test changes before merge; the artifact-publishing build job stays - # pinned to dev either way, so a branch dispatch can never ship branch binaries. - - uses: actions/checkout@v7 - with: - ref: ${{ github.event_name == 'workflow_dispatch' && github.ref_name || 'dev' }} - - - name: Setup .NET 10.0 - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - # Darling bundles a PostgreSQL 18 + TimescaleDB runtime (pg-runtime.zip). The fetch script - # pulls ~340MB of pinned EDB/TimescaleDB archives, so cache the assembled zip on the SAME - # key build.yml's release job uses (the fetch script's own content hash) — nightly and - # release share one cache entry, so a warm cache means no download here. - - name: Cache Darling pg-runtime.zip - id: cache-pg-runtime - uses: actions/cache@v6 - with: - path: Darling/artifacts/pg-runtime.zip - key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} - - - name: Build Darling pg-runtime.zip (cache miss only) - if: steps.cache-pg-runtime.outputs.cache-hit != 'true' - shell: pwsh - run: ./Darling/tools/fetch-pg-runtime.ps1 - - # One uniform path for hit and miss: we always have the zip (restored or freshly built), - # so always extract it — simpler than branching on the script's -KeepWork assembled tree. - # ExtractToDirectory mirrors the fetch script's own API and reads any zip it writes. - - name: Extract pg-runtime - shell: pwsh - run: | - $zip = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime.zip" - $dest = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime" - if (Test-Path $dest) { Remove-Item -Recurse -Force $dest } - Add-Type -AssemblyName System.IO.Compression.FileSystem - [System.IO.Compression.ZipFile]::ExtractToDirectory($zip, $dest) - if (-not (Test-Path "$dest\pgsql\bin\pg_ctl.exe")) { throw "pg-runtime missing pgsql\bin\pg_ctl.exe" } - - # The PREVIOUS-major runtime, for the upgraded-in-place fixture (#1706). Without it the gated - # store-upgrade E2E skips, and with it skips the ONLY behavioral coverage of the in-place - # 17-to-18 path: the runtime rescue, the TimescaleDB bridge, the data-directory swap, the - # revert, and the loopback override that stops pg_upgrade dialing ::1 against our IPv4-only - # listen_addresses. String assertions catch that constant being deleted; nothing but this - # catches the override ceasing to TAKE EFFECT, and the failure mode is a fleet-wide hang. - # Only the DOWNLOADS are cached (~340MB of pinned archives): the script re-verifies them by - # SHA256 and re-assembles in seconds, so a warm cache costs no network. - - name: Cache upgrade-fixture downloads (previous-major runtime) - uses: actions/cache@v6 - with: - path: Darling/artifacts/upgrade-fixture/work/downloads - key: upgrade-fixture-${{ runner.os }}-${{ hashFiles('Darling/tools/new-upgraded-store-fixture.ps1') }} - - # -SkipNew: the CURRENT runtime is already built/cached above as pg-runtime.zip, which is what - # DARLING_TEST_PGRUNTIME_NEWZIP points at. This step only needs the old side. - - name: Build previous-major runtime for the upgrade fixture - shell: pwsh - run: ./Darling/tools/new-upgraded-store-fixture.ps1 -SkipNew - - - name: Restore Darling.Tests - run: dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode - - - name: Build Darling.Tests - run: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore - - # Stand up a throwaway cluster from the bundled runtime. initdb TRUST auth is acceptable - # ONLY here: an ephemeral CI runner, loopback-only, throwaway data. The PRODUCT default is - # the opposite (scram-sha-256 + a generated credential; see DarlingManagedPostgres). The - # appended settings mirror DarlingManagedPostgres.BuildConfAppend (timescaledb preload — - # the Timescale-gated tests detect and use it — plus the port and loopback bind) and - # BuildWorkerSizingConfAppend (the two worker settings). Port 5541 is fixed and distinct; the - # managed-bootstrap E2E starts its OWN postgres on a random free port, so there is no collision. - # - # #1888: the worker settings are load-bearing, not garnish. PostgreSQL's default - # max_worker_processes = 8 cannot launch TimescaleDB's per-hypertable compression, retention - # and continuous-aggregate policy jobs, so without them this job tested a configuration no - # customer runs and made scheduler-racing failures luck-of-the-slot instead of reproducible. - # Values are the product's own derivation from the live hypertable count - # (TimescaleSupport.HypertableCount = the 50-collector catalog + collection_log = 51): - # timescaledb.max_background_workers = HypertableCount + 2 = 53 - # max_worker_processes = 3 + (HypertableCount + 2) + 8 = 64 - # Kept honest by CiClusterWorkerSizingTests (parses this file against the formula) and - # CiClusterWorkerSizingLiveTests (asserts the running cluster serves them). Must configure - # the cluster identically to build.yml's darling-pg job: that guard parses the appended - # settings out of BOTH files and requires the two sets to be equal, so a fix applied to one - # workflow and not the other — the ordinary way these hand-maintained copies drift — fails. - - name: Initialize and start throwaway PostgreSQL - shell: pwsh - run: | - $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" - $dataDir = "$env:RUNNER_TEMP\darling-pgdata" - $logFile = "$env:RUNNER_TEMP\darling-pg.log" - # The bootstrap superuser is named "darling", not "postgres": the V8 schema-split - # migration runs CREATE SCHEMA ... AUTHORIZATION darling (PgSchemaGenerator.OwnerRole), - # so a cluster without that role fails every fresh-store migration. Matching managed - # mode's shape (DarlingManagedPostgres also initdbs its owner as "darling"). - & "$bin\initdb.exe" -D $dataDir -U darling -A trust --encoding=UTF8 - if ($LASTEXITCODE -ne 0) { throw "initdb failed ($LASTEXITCODE)" } - Add-Content -Path "$dataDir\postgresql.conf" -Value "shared_preload_libraries = 'timescaledb'" - Add-Content -Path "$dataDir\postgresql.conf" -Value "port = 5541" - Add-Content -Path "$dataDir\postgresql.conf" -Value "listen_addresses = '127.0.0.1'" - Add-Content -Path "$dataDir\postgresql.conf" -Value "timescaledb.max_background_workers = 53" - Add-Content -Path "$dataDir\postgresql.conf" -Value "max_worker_processes = 64" - & "$bin\pg_ctl.exe" -D $dataDir -l $logFile -w start - if ($LASTEXITCODE -ne 0) { if (Test-Path $logFile) { Get-Content $logFile -Tail 50 }; throw "pg_ctl start failed ($LASTEXITCODE)" } - & "$bin\createdb.exe" -h 127.0.0.1 -p 5541 -U darling darling - if ($LASTEXITCODE -ne 0) { throw "createdb failed ($LASTEXITCODE)" } - - # DARLING_TEST_PG lights up the [Collection("live-postgres")] classes; DARLING_TEST_PGRUNTIME - # lights up the managed-bootstrap E2E. DARLING_TEST_SQL is intentionally unset — the one - # live-SQL-Server E2E stays skipped (no SQL Server on the runner). - - name: Run Darling PG tests - shell: pwsh - env: - DARLING_TEST_PG: "Host=127.0.0.1;Port=5541;Username=darling;Database=darling" - DARLING_TEST_PGRUNTIME: ${{ github.workspace }}\Darling\artifacts\pg-runtime - # #1706: the pair that lights up the upgraded-in-place store-upgrade E2E — a real - # previous-major store (hypertable, TOAST-sized plan XML, continuous aggregate, - # compressed chunk) upgraded through the production bootstrap and compared by ordered - # row checksum before and after. This is the [#1705] CI gap: every other job only ever - # sees a FRESH store, so nothing else can catch an upgrade path that breaks. - DARLING_TEST_PGRUNTIME_OLD: ${{ github.workspace }}\Darling\artifacts\upgrade-fixture\old\pg-runtime - DARLING_TEST_PGRUNTIME_NEWZIP: ${{ github.workspace }}\Darling\artifacts\pg-runtime.zip - run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build -- -trx TestResults/darling-nightly.trx - - - name: Stop PostgreSQL - if: always() - shell: pwsh - run: | - $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" - $dataDir = "$env:RUNNER_TEMP\darling-pgdata" - if (Test-Path "$bin\pg_ctl.exe") { & "$bin\pg_ctl.exe" -D $dataDir -m fast -w stop } - exit 0 - - - name: Upload PG log and test results on failure - if: failure() - uses: actions/upload-artifact@v6 - with: - name: darling-pg-failure - path: | - ${{ runner.temp }}/darling-pg.log - TestResults/ - if-no-files-found: ignore - - # ── Linux artifact + container image (#1804) ───────────────────────────────────────────────────── - # Runs AFTER the windows build job so the nightly release exists to upload into. Publishes the - # linux-x64 service tar.gz with its own checksum file (the windows job owns SHA256SUMS.txt; a - # cross-job rewrite of one file is a race), and pushes the service image to ghcr tagged :nightly. - # The bundled pg-runtime is deliberately absent from the linux artifact — the compose distribution - # pairs the service with the official timescale/timescaledb image, and managed mode stays Windows. - linux: - needs: build - runs-on: ubuntu-latest - timeout-minutes: 30 - - steps: - - uses: actions/checkout@v7 - with: - ref: dev - - - name: Setup .NET 10.0 - uses: actions/setup-dotnet@v6 - with: - global-json-file: global.json - cache: true - cache-dependency-path: '**/packages.lock.json' - - - name: Set nightly version - id: version - shell: bash - run: | - set -euo pipefail - base=$(grep -oPm1 '(?<=)[^<]+' Lite/PerformanceMonitorLite.csproj) - date=$(date +%Y%m%d) - echo "VERSION=${base}-nightly.${date}" >> "$GITHUB_OUTPUT" - echo "Nightly version: ${base}-nightly.${date}" - - - name: Publish service (linux-x64) - run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -r linux-x64 --self-contained false -o publish/DarlingService-linux - - - name: Package linux artifact + checksum - shell: bash - run: | - set -euo pipefail - version="${{ steps.version.outputs.VERSION }}" - mkdir -p releases - tar -C publish/DarlingService-linux -czf "releases/PerformanceMonitorDarling-linux-x64-${version}.tar.gz" . - (cd releases && sha256sum "PerformanceMonitorDarling-linux-x64-${version}.tar.gz" > SHA256SUMS-linux.txt && cat SHA256SUMS-linux.txt) - - - name: Upload linux artifact to the nightly release - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - run: gh release upload nightly releases/PerformanceMonitorDarling-linux-x64-*.tar.gz releases/SHA256SUMS-linux.txt --clobber - - - name: Build and push container image (ghcr, :nightly) - env: - GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} - shell: bash - run: | - set -euo pipefail - version="${{ steps.version.outputs.VERSION }}" - image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" - echo "$GH_TOKEN" | docker login ghcr.io -u "${{ github.actor }}" --password-stdin - docker build -f Darling/Dockerfile -t "${image}:nightly" -t "${image}:${version}" . - docker push "${image}:nightly" - docker push "${image}:${version}" - - # Keyless Sigstore signing: Fulcio issues a short-lived certificate against this job's OIDC - # identity and the signature lands in ghcr next to the image, logged in Rekor. No keys exist - # anywhere to manage or leak — the signature attests "built by this repository's workflow", - # which is the claim a container consumer actually wants verified. (SignPath's cosign support - # is edition-gated; this path has no subscription dependency.) Verify: - # cosign verify ghcr.io/erikdarlingdata/performancemonitor-darling:nightly \ - # --certificate-identity-regexp 'github.com/erikdarlingdata/PerformanceMonitor' \ - # --certificate-oidc-issuer https://token.actions.githubusercontent.com - - name: Install cosign - uses: sigstore/cosign-installer@v3 - - - name: Sign container image (keyless, Sigstore) - shell: bash - run: | - set -euo pipefail - image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" - digest="$(docker inspect --format='{{index .RepoDigests 0}}' "${image}:nightly")" - cosign sign --yes "${digest}" - - # GitHub-native provenance for the tarball (SLSA): verified with - # gh attestation verify PerformanceMonitorDarling-linux-x64-.tar.gz -R erikdarlingdata/PerformanceMonitor - - name: Attest the linux tarball (GitHub provenance) - uses: actions/attest-build-provenance@v4 - with: - subject-path: releases/PerformanceMonitorDarling-linux-x64-*.tar.gz +name: Nightly Build + +on: + schedule: + # 6:00 AM UTC (1:00 AM EST / 2:00 AM EDT) + - cron: '0 6 * * *' + workflow_dispatch: # manual trigger — and the vehicle the scheduled re-dispatch below rides + inputs: + from_schedule: + description: 'Set true by the scheduled re-dispatch so the 24h new-commit check applies. Leave false for manual runs, which always build.' + type: boolean + required: false + default: false + +permissions: + contents: write + # #1804: the linux job pushes the nightly container image to ghcr. + packages: write + # Sigstore keyless signing + GitHub provenance for the linux artifacts: id-token lets the job + # obtain its OIDC identity (Fulcio issues the short-lived signing cert against it), attestations + # lets attest-build-provenance store the tarball's provenance. Neither grants anything else. + id-token: write + attestations: write + +jobs: + # Scheduled workflows always execute the DEFAULT branch's copy of this file, while nightly + # artifacts deliberately build from dev's tree. That skew is how the 2026-07-26 nightly + # failed (run 30194606068): main's stale copy still read Dashboard/Dashboard.csproj, a path + # #1612 moved to deprecated/ on dev, so 'Set nightly version' died on a file missing from + # the tree it had just checked out — and the same trap bit before (#1550/#1551). The cure + # is structural, not another sync: on schedule this workflow does NOTHING but re-dispatch + # itself onto the dev REF, because a workflow_dispatch run executes the dispatched ref's + # copy of this file — dev's, current by definition. Once main carries this shape, its copy + # has exactly one job that must keep working, and that job references no tree paths at + # all; every future change to the real nightly logic lands on dev and takes effect the + # night it merges, no promotion to main needed. GITHUB_TOKEN can create workflow_dispatch + # runs (the Actions recursion guard exempts workflow_dispatch and repository_dispatch), + # and the dispatched run cannot loop back here because it arrives as workflow_dispatch, + # not schedule. Until main is synced once, the scheduled run still executes main's OLD + # copy and keeps failing nightly — the one-time sync is in the PR that introduced this. + redispatch: + if: github.event_name == 'schedule' + runs-on: ubuntu-latest + timeout-minutes: 5 + permissions: + actions: write + steps: + - name: Re-dispatch this workflow onto the dev ref + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: gh workflow run nightly.yml --repo ${{ github.repository }} --ref dev -f from_schedule=true + + # Everything below runs only in a workflow_dispatch run — the scheduled re-dispatch or a + # manual one — which executes from the dispatched ref (dev for the scheduled path). + # + # Each checkout resolves that ref EXPLICITLY rather than pinning `ref: dev`. Three of these + # were pinned, so a manual dispatch onto any other branch published dev's artifacts while + # reporting the dispatched ref as the run's head SHA — the run, the API and the release page + # all named a branch that was not what got built. That is worse than simply ignoring the + # input, because the mismatch is invisible: it surfaced as a release-candidate build stamping + # the PREVIOUS version, which is only noticeable if you happen to read the artifact name. + check: + if: github.event_name == 'workflow_dispatch' + runs-on: ubuntu-latest + timeout-minutes: 10 + outputs: + has_changes: ${{ steps.check.outputs.has_changes }} + steps: + - uses: actions/checkout@v7 + with: + ref: ${{ github.event_name == 'workflow_dispatch' && github.ref_name || 'dev' }} + fetch-depth: 0 + + - name: Check for new commits in last 24 hours + id: check + run: | + RECENT=$(git log --since="24 hours ago" --oneline | head -1) + if [ -n "$RECENT" ]; then + echo "has_changes=true" >> $GITHUB_OUTPUT + echo "New commits found — building nightly" + else + echo "has_changes=false" >> $GITHUB_OUTPUT + echo "No new commits — skipping nightly build" + fi + + build: + needs: check + # Manual dispatches always build (from_schedule defaults false); the scheduled + # re-dispatch sets from_schedule=true and builds only when dev changed in the last + # 24h — the same policy the schedule applied when it ran these jobs directly. + if: needs.check.outputs.has_changes == 'true' || inputs.from_schedule != true + runs-on: windows-latest + # Full pipeline (restore, tests, four publishes, cold pg-runtime fetch, vpk pack, + # release upload) is well under an hour; 90 minutes means hung-not-slow. Nightly ships + # unsigned, so unlike build.yml's release path there is no manual signing gate to wait on. + timeout-minutes: 90 + + steps: + - uses: actions/checkout@v7 + with: + ref: ${{ github.event_name == 'workflow_dispatch' && github.ref_name || 'dev' }} + + - name: Setup .NET 10.0 + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + - name: Set nightly version + id: version + shell: pwsh + run: | + $base = ([xml](Get-Content Lite/PerformanceMonitorLite.csproj)).Project.PropertyGroup.Version | Where-Object { $_ } + $date = Get-Date -Format "yyyyMMdd" + $nightly = "$base-nightly.$date" + echo "VERSION=$nightly" >> $env:GITHUB_OUTPUT + echo "Nightly version: $nightly" + + - name: Restore dependencies + run: | + dotnet restore Lite/PerformanceMonitorLite.csproj --locked-mode + dotnet restore Lite.Tests/Lite.Tests.csproj --locked-mode + dotnet restore Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj --locked-mode + + - name: Run tests + run: dotnet run --project Lite.Tests/Lite.Tests.csproj -c Release + + # #2501: -r win-x64 --self-contained, and it makes the portable ZIP SMALLER, not bigger. + # The RID-agnostic framework-dependent publish this replaced copied every platform its + # packages ship - 537 MB of runtimes\ on a 565 MB tree, of which only the 52 MB win-x64 + # folder can ever load on Windows (DuckDB.NET.Bindings.Full is most of it, SkiaSharp and + # SqlClient behind it). Pinning the RID drops ~485 MB of unloadable native payload, which + # is far more than the bundled .NET/WPF/ASP.NET runtime adds back. Measured on the same + # commit and SDK: 565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB. + # + # The size was the secondary reason. The primary one is #2489: a framework-dependent ZIP + # unzipped onto a stock Windows Server dies on the .NET host's own "You must install .NET + # to run this application" before a line of our code runs, so Lite cannot report it, and + # #2499 could only ship a READ-ME-FIRST.txt next to the exe. Self-contained there is + # nothing to install and nothing to fail. That matters most for the nightly ZIP, which is + # the UAT download and is not offered as a Setup.exe at all. + # + # Lite/PerformanceMonitorLite.csproj declares RuntimeIdentifiers so the committed lock + # file covers this RID; see the comment there before removing either half. + - name: Publish Lite + run: dotnet publish Lite/PerformanceMonitorLite.csproj -c Release -r win-x64 --self-contained -o publish/Lite + + - name: Publish Darling Service + run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o publish/DarlingService + + - name: Publish Darling Viewer + run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -o publish/DarlingViewer + + # Self-contained viewer publish that feeds the remote-seat Velopack Setup.exe (#1555), the same + # publish shape build.yml uses for the Dashboard/Lite Velopack packs. This is IN ADDITION to the + # framework-dependent "Publish Darling Viewer" above, which still feeds the co-located viewer\ + # folder inside PerformanceMonitorDarling-*.zip — that zip is unchanged. + - name: Publish Darling Viewer (self-contained for Velopack) + run: dotnet publish Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -r win-x64 --self-contained -o publish/DarlingViewer-velopack + + # Same cache key as build.yml's release path: the fetch script's content hash (the SHA256 + # pins live inside it). The ~340MB EDB/TimescaleDB fetch runs at most once per pin-set per + # branch; nightly and release runs share the assembled zip whenever the cache is visible. + - name: Cache Darling pg-runtime.zip + id: cache-pg-runtime + uses: actions/cache@v6 + with: + path: Darling/artifacts/pg-runtime.zip + key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} + + - name: Build Darling pg-runtime.zip + if: steps.cache-pg-runtime.outputs.cache-hit != 'true' + shell: pwsh + run: ./Darling/tools/fetch-pg-runtime.ps1 + + - name: Package artifacts + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + New-Item -ItemType Directory -Force -Path releases + + Compress-Archive -Path 'publish/Lite/*' -DestinationPath "releases/PerformanceMonitorLite-$version.zip" -Force + + + # Same layout as the release zip (build.yml "Package Darling (signed)"): service at the + # archive root with darling.sample.json alongside, viewer under viewer\, pg-runtime.zip + # beside the service exe where DarlingManagedPostgres extracts it on first run. Nightly + # zips are unsigned across the board, so this stages from publish/ instead of signed/. + $darlingDir = 'publish/Darling' + New-Item -ItemType Directory -Force -Path "$darlingDir/viewer" | Out-Null + Copy-Item 'publish/DarlingService/*' $darlingDir -Recurse + Copy-Item 'publish/DarlingViewer/*' "$darlingDir/viewer" -Recurse + Copy-Item 'Darling/artifacts/pg-runtime.zip' $darlingDir + + Compress-Archive -Path 'publish/Darling/*' -DestinationPath "releases/PerformanceMonitorDarling-$version.zip" -Force + + # Darling viewer remote-seat installer (#1555). Nightly ships it UNSIGNED like every other nightly + # artifact. Mirrors build.yml's release vpk pack (same pack id / exe / channel) but packs from the + # raw self-contained publish (no SignPath), and deliberately does NOT touch the Velopack update + # feed: the nightly GitHub release is deleted + recreated each night, so there is no persistent + # delta chain to `vpk download`/`vpk upload` from — we ship a standalone full Setup.exe as a plain + # release asset. Copied to a deterministic name so the checksum + upload steps below pick it up. + # Purely additive: the co-located viewer inside PerformanceMonitorDarling-*.zip is untouched. + - name: Create Darling Viewer Setup.exe (Velopack, unsigned) + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + dotnet tool install -g vpk --version 1.2.0 + New-Item -ItemType Directory -Force -Path releases/velopack-darlingviewer + vpk pack -u PerformanceMonitorDarlingViewer -v $version -p publish/DarlingViewer-velopack -e PerformanceMonitor.Darling.Viewer.exe -o releases/velopack-darlingviewer --channel darlingviewer + $setup = Get-ChildItem releases/velopack-darlingviewer/*Setup.exe | Select-Object -First 1 + Copy-Item $setup.FullName "releases/PerformanceMonitorDarlingViewer-$version-Setup.exe" + + - name: Generate checksums + shell: pwsh + run: | + $checksums = Get-ChildItem releases/*.zip, releases/*.exe | ForEach-Object { + $hash = (Get-FileHash $_.FullName -Algorithm SHA256).Hash.ToLower() + "$hash $($_.Name)" + } + # LF and no BOM, written explicitly (#2383). Out-File on Windows emits CRLF whatever + # the encoding, and `shasum -c` on macOS/Linux then takes the trailing \r as part of + # the FILENAME - every entry reports "No such file or directory / FAILED open or + # read", which is precisely what a tampered download looks like. The hashes were + # always right; the file just could not be used by the tool it exists for, on the two + # platforms where that tool is the default. A BOM breaks the first line the same way. + # + # This is the SECOND copy of that step. #2384 fixed the release workflow and left this + # one, so every nightly since has published a checksum file nobody outside Windows can + # verify - and the nightly is the build most likely to be checked by hand. + [IO.File]::WriteAllText( + "$PWD/releases/SHA256SUMS.txt", + (($checksums -join "`n") + "`n"), + [Text.UTF8Encoding]::new($false)) + Write-Host "Checksums:" + $checksums | ForEach-Object { Write-Host $_ } + + - name: Delete previous nightly release + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: gh release delete nightly --yes --cleanup-tag 2>$null; exit 0 + shell: pwsh + + - name: Create nightly release + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + shell: pwsh + run: | + $version = "${{ steps.version.outputs.VERSION }}" + $sha = git rev-parse --short HEAD + $body = @" + Automated nightly build from ``dev`` branch. + + **Version:** ``$version`` + **Commit:** ``$sha`` + **Built:** $(Get-Date -Format "yyyy-MM-dd HH:mm UTC") + + > These builds include the latest changes and may be unstable. + > For production use, download the [latest stable release](https://github.com/erikdarlingdata/PerformanceMonitor/releases/latest). + "@ + + gh release create nightly ` + --target dev ` + --title "Nightly Build ($version)" ` + --notes $body ` + --prerelease ` + releases/*.zip releases/*.exe releases/SHA256SUMS.txt + + # Darling.Tests has live-PostgreSQL tests gated on env vars that the per-PR/push build never + # sets, so regular CI runs only the ungated subset. This job stands up a throwaway PostgreSQL + # from the bundled pg-runtime and runs the full Darling suite against it nightly, so the + # Postgres + managed-bootstrap paths get real coverage. The one live-SQL-Server E2E stays + # skipped (no SQL Server on the runner) — expected. Gated like the build job so it only runs + # when dev actually changed (or on manual dispatch). + darling-pg: + name: Darling PostgreSQL tests + needs: check + # Same gating as the build job: manual dispatches always run, the scheduled + # re-dispatch (from_schedule=true) runs only when dev changed in the last 24h. + if: needs.check.outputs.has_changes == 'true' || inputs.from_schedule != true + runs-on: windows-latest + # Cold pg-runtime fetch + build + the full live-PG suite fits well inside an hour; a + # cluster that never comes up (pg_ctl -w) is the hang this bounds. + timeout-minutes: 60 + + steps: + # Scheduled runs always test dev (schedules execute from the default branch, so ref_name + # would be main). A manual dispatch tests the DISPATCHED ref — the only way to validate a + # branch's gated-pg test changes before merge; the artifact-publishing build job stays + # pinned to dev either way, so a branch dispatch can never ship branch binaries. + - uses: actions/checkout@v7 + with: + ref: ${{ github.event_name == 'workflow_dispatch' && github.ref_name || 'dev' }} + + - name: Setup .NET 10.0 + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + # Darling bundles a PostgreSQL 18 + TimescaleDB runtime (pg-runtime.zip). The fetch script + # pulls ~340MB of pinned EDB/TimescaleDB archives, so cache the assembled zip on the SAME + # key build.yml's release job uses (the fetch script's own content hash) — nightly and + # release share one cache entry, so a warm cache means no download here. + - name: Cache Darling pg-runtime.zip + id: cache-pg-runtime + uses: actions/cache@v6 + with: + path: Darling/artifacts/pg-runtime.zip + key: pg-runtime-${{ runner.os }}-${{ hashFiles('Darling/tools/fetch-pg-runtime.ps1') }} + + - name: Build Darling pg-runtime.zip (cache miss only) + if: steps.cache-pg-runtime.outputs.cache-hit != 'true' + shell: pwsh + run: ./Darling/tools/fetch-pg-runtime.ps1 + + # One uniform path for hit and miss: we always have the zip (restored or freshly built), + # so always extract it — simpler than branching on the script's -KeepWork assembled tree. + # ExtractToDirectory mirrors the fetch script's own API and reads any zip it writes. + - name: Extract pg-runtime + shell: pwsh + run: | + $zip = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime.zip" + $dest = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime" + if (Test-Path $dest) { Remove-Item -Recurse -Force $dest } + Add-Type -AssemblyName System.IO.Compression.FileSystem + [System.IO.Compression.ZipFile]::ExtractToDirectory($zip, $dest) + if (-not (Test-Path "$dest\pgsql\bin\pg_ctl.exe")) { throw "pg-runtime missing pgsql\bin\pg_ctl.exe" } + + # The PREVIOUS-major runtime, for the upgraded-in-place fixture (#1706). Without it the gated + # store-upgrade E2E skips, and with it skips the ONLY behavioral coverage of the in-place + # 17-to-18 path: the runtime rescue, the TimescaleDB bridge, the data-directory swap, the + # revert, and the loopback override that stops pg_upgrade dialing ::1 against our IPv4-only + # listen_addresses. String assertions catch that constant being deleted; nothing but this + # catches the override ceasing to TAKE EFFECT, and the failure mode is a fleet-wide hang. + # Only the DOWNLOADS are cached (~340MB of pinned archives): the script re-verifies them by + # SHA256 and re-assembles in seconds, so a warm cache costs no network. + - name: Cache upgrade-fixture downloads (previous-major runtime) + uses: actions/cache@v6 + with: + path: Darling/artifacts/upgrade-fixture/work/downloads + key: upgrade-fixture-${{ runner.os }}-${{ hashFiles('Darling/tools/new-upgraded-store-fixture.ps1') }} + + # -SkipNew: the CURRENT runtime is already built/cached above as pg-runtime.zip, which is what + # DARLING_TEST_PGRUNTIME_NEWZIP points at. This step only needs the old side. + - name: Build previous-major runtime for the upgrade fixture + shell: pwsh + run: ./Darling/tools/new-upgraded-store-fixture.ps1 -SkipNew + + - name: Restore Darling.Tests + run: dotnet restore Darling/Darling.Tests/Darling.Tests.csproj --locked-mode + + - name: Build Darling.Tests + run: dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-restore + + # Stand up a throwaway cluster from the bundled runtime. initdb TRUST auth is acceptable + # ONLY here: an ephemeral CI runner, loopback-only, throwaway data. The PRODUCT default is + # the opposite (scram-sha-256 + a generated credential; see DarlingManagedPostgres). The + # appended settings mirror DarlingManagedPostgres.BuildConfAppend (timescaledb preload — + # the Timescale-gated tests detect and use it — plus the port and loopback bind) and + # BuildWorkerSizingConfAppend (the two worker settings). Port 5541 is fixed and distinct; the + # managed-bootstrap E2E starts its OWN postgres on a random free port, so there is no collision. + # + # #1888: the worker settings are load-bearing, not garnish. PostgreSQL's default + # max_worker_processes = 8 cannot launch TimescaleDB's per-hypertable compression, retention + # and continuous-aggregate policy jobs, so without them this job tested a configuration no + # customer runs and made scheduler-racing failures luck-of-the-slot instead of reproducible. + # Values are the product's own derivation from the live hypertable count + # (TimescaleSupport.HypertableCount = the 62-collector catalog + collection_log = 63): + # timescaledb.max_background_workers = HypertableCount + 2 = 71 + # max_worker_processes = 3 + (HypertableCount + 2) + 8 = 82 + # Kept honest by CiClusterWorkerSizingTests (parses this file against the formula) and + # CiClusterWorkerSizingLiveTests (asserts the running cluster serves them). Must configure + # the cluster identically to build.yml's darling-pg job: that guard parses the appended + # settings out of BOTH files and requires the two sets to be equal, so a fix applied to one + # workflow and not the other — the ordinary way these hand-maintained copies drift — fails. + - name: Initialize and start throwaway PostgreSQL + shell: pwsh + run: | + $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" + $dataDir = "$env:RUNNER_TEMP\darling-pgdata" + $logFile = "$env:RUNNER_TEMP\darling-pg.log" + # The bootstrap superuser is named "darling", not "postgres": the V8 schema-split + # migration runs CREATE SCHEMA ... AUTHORIZATION darling (PgSchemaGenerator.OwnerRole), + # so a cluster without that role fails every fresh-store migration. Matching managed + # mode's shape (DarlingManagedPostgres also initdbs its owner as "darling"). + & "$bin\initdb.exe" -D $dataDir -U darling -A trust --encoding=UTF8 + if ($LASTEXITCODE -ne 0) { throw "initdb failed ($LASTEXITCODE)" } + Add-Content -Path "$dataDir\postgresql.conf" -Value "shared_preload_libraries = 'timescaledb'" + Add-Content -Path "$dataDir\postgresql.conf" -Value "port = 5541" + Add-Content -Path "$dataDir\postgresql.conf" -Value "listen_addresses = '127.0.0.1'" + Add-Content -Path "$dataDir\postgresql.conf" -Value "timescaledb.max_background_workers = 71" + Add-Content -Path "$dataDir\postgresql.conf" -Value "max_worker_processes = 82" + & "$bin\pg_ctl.exe" -D $dataDir -l $logFile -w start + if ($LASTEXITCODE -ne 0) { if (Test-Path $logFile) { Get-Content $logFile -Tail 50 }; throw "pg_ctl start failed ($LASTEXITCODE)" } + & "$bin\createdb.exe" -h 127.0.0.1 -p 5541 -U darling darling + if ($LASTEXITCODE -ne 0) { throw "createdb failed ($LASTEXITCODE)" } + + # DARLING_TEST_PG lights up the [Collection("live-postgres")] classes; DARLING_TEST_PGRUNTIME + # lights up the managed-bootstrap E2E. DARLING_TEST_SQL is intentionally unset — the one + # live-SQL-Server E2E stays skipped (no SQL Server on the runner). + - name: Run Darling PG tests + shell: pwsh + env: + DARLING_TEST_PG: "Host=127.0.0.1;Port=5541;Username=darling;Database=darling" + DARLING_TEST_PGRUNTIME: ${{ github.workspace }}\Darling\artifacts\pg-runtime + # #1706: the pair that lights up the upgraded-in-place store-upgrade E2E — a real + # previous-major store (hypertable, TOAST-sized plan XML, continuous aggregate, + # compressed chunk) upgraded through the production bootstrap and compared by ordered + # row checksum before and after. This is the [#1705] CI gap: every other job only ever + # sees a FRESH store, so nothing else can catch an upgrade path that breaks. + DARLING_TEST_PGRUNTIME_OLD: ${{ github.workspace }}\Darling\artifacts\upgrade-fixture\old\pg-runtime + DARLING_TEST_PGRUNTIME_NEWZIP: ${{ github.workspace }}\Darling\artifacts\pg-runtime.zip + run: dotnet run --project Darling/Darling.Tests/Darling.Tests.csproj -c Release --no-build -- -trx TestResults/darling-nightly.trx + + - name: Stop PostgreSQL + if: always() + shell: pwsh + run: | + $bin = "$env:GITHUB_WORKSPACE\Darling\artifacts\pg-runtime\pgsql\bin" + $dataDir = "$env:RUNNER_TEMP\darling-pgdata" + if (Test-Path "$bin\pg_ctl.exe") { & "$bin\pg_ctl.exe" -D $dataDir -m fast -w stop } + exit 0 + + - name: Upload PG log and test results on failure + if: failure() + uses: actions/upload-artifact@v6 + with: + name: darling-pg-failure + path: | + ${{ runner.temp }}/darling-pg.log + TestResults/ + if-no-files-found: ignore + + # ── Linux artifact + container image (#1804) ───────────────────────────────────────────────────── + # Runs AFTER the windows build job so the nightly release exists to upload into. Publishes the + # linux-x64 service tar.gz with its own checksum file (the windows job owns SHA256SUMS.txt; a + # cross-job rewrite of one file is a race), and pushes the service image to ghcr tagged :nightly. + # The bundled pg-runtime is deliberately absent from the linux artifact — the compose distribution + # pairs the service with the official timescale/timescaledb image, and managed mode stays Windows. + linux: + needs: build + runs-on: ubuntu-latest + timeout-minutes: 30 + + steps: + - uses: actions/checkout@v7 + with: + ref: ${{ github.event_name == 'workflow_dispatch' && github.ref_name || 'dev' }} + + - name: Setup .NET 10.0 + uses: actions/setup-dotnet@v6 + with: + global-json-file: global.json + cache: true + cache-dependency-path: '**/packages.lock.json' + + - name: Set nightly version + id: version + shell: bash + run: | + set -euo pipefail + base=$(grep -oPm1 '(?<=)[^<]+' Lite/PerformanceMonitorLite.csproj) + date=$(date +%Y%m%d) + echo "VERSION=${base}-nightly.${date}" >> "$GITHUB_OUTPUT" + echo "Nightly version: ${base}-nightly.${date}" + + - name: Publish service (linux-x64) + run: dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -r linux-x64 --self-contained false -o publish/DarlingService-linux + + - name: Package linux artifact + checksum + shell: bash + run: | + set -euo pipefail + version="${{ steps.version.outputs.VERSION }}" + mkdir -p releases + tar -C publish/DarlingService-linux -czf "releases/PerformanceMonitorDarling-linux-x64-${version}.tar.gz" . + (cd releases && sha256sum "PerformanceMonitorDarling-linux-x64-${version}.tar.gz" > SHA256SUMS-linux.txt && cat SHA256SUMS-linux.txt) + + - name: Upload linux artifact to the nightly release + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + run: gh release upload nightly releases/PerformanceMonitorDarling-linux-x64-*.tar.gz releases/SHA256SUMS-linux.txt --clobber + + - name: Build and push container image (ghcr, :nightly) + env: + GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} + shell: bash + run: | + set -euo pipefail + version="${{ steps.version.outputs.VERSION }}" + image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" + echo "$GH_TOKEN" | docker login ghcr.io -u "${{ github.actor }}" --password-stdin + docker build -f Darling/Dockerfile -t "${image}:nightly" -t "${image}:${version}" . + docker push "${image}:nightly" + docker push "${image}:${version}" + + # Keyless Sigstore signing: Fulcio issues a short-lived certificate against this job's OIDC + # identity and the signature lands in ghcr next to the image, logged in Rekor. No keys exist + # anywhere to manage or leak — the signature attests "built by this repository's workflow", + # which is the claim a container consumer actually wants verified. (SignPath's cosign support + # is edition-gated; this path has no subscription dependency.) Verify: + # cosign verify ghcr.io/erikdarlingdata/performancemonitor-darling:nightly \ + # --certificate-identity-regexp 'github.com/erikdarlingdata/PerformanceMonitor' \ + # --certificate-oidc-issuer https://token.actions.githubusercontent.com + - name: Install cosign + uses: sigstore/cosign-installer@v3 + + - name: Sign container image (keyless, Sigstore) + shell: bash + run: | + set -euo pipefail + image="ghcr.io/${{ github.repository_owner }}/performancemonitor-darling" + digest="$(docker inspect --format='{{index .RepoDigests 0}}' "${image}:nightly")" + cosign sign --yes "${digest}" + + # GitHub-native provenance for the tarball (SLSA): verified with + # gh attestation verify PerformanceMonitorDarling-linux-x64-.tar.gz -R erikdarlingdata/PerformanceMonitor + - name: Attest the linux tarball (GitHub provenance) + uses: actions/attest-build-provenance@v4 + with: + subject-path: releases/PerformanceMonitorDarling-linux-x64-*.tar.gz diff --git a/CHANGELOG.md b/CHANGELOG.md index d20e253d6..10cdf7eb5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,125 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [Unreleased] + +### Added +- **Store exposure admits both roles at once** ([#2665]) - `postgres.network.role` took one role, so a team could have an `admin` Viewer or read-only `viewer` ones on the LAN but not both, even though the store has always created and separately granted both roles and given each its own credential. What was single-valued was the one generated `hostssl` line. The field now accepts both - `"admin,viewer"`, separated by a comma, a `+`, or whitespace - and the reconciler writes one rule per role inside its own marked block, which is where it matters: the documented workaround was a second `hostssl` line added by hand outside the markers, and because the reconciler preserves everything outside them verbatim, that line silently kept the old CIDR when `allowFrom` was later tightened. Somebody narrowing access would reasonably have believed they had. Inside the block both rules narrow together. The guardrails hold as written: every generated line still names exactly one role and one CIDR, never `all` and never the superuser; the order is settled by the parser rather than by how the field was typed, so re-spelling `admin,viewer` as `viewer,admin` does not rewrite `pg_hba.conf` and reload the server; an unrecognised element rejects the whole value rather than opening the store with the half that parsed; and absent or blank still means `viewer` alone, so no existing configuration gains admin-capable access from the field learning to take a list. `--configure-network` offers both at the prompt (default still `viewer`) and `--print-viewer-connection` prints one paste-ready string per admitted role, each labelled with the seat it authenticates as. +- **PostgreSQL I/O and database trends** ([#2663]) - the two deepest series in the store had no shape over time: `pg_io_stats` and `pg_database_stats` had over a thousand snapshots each on a two-day-old rig and nothing differenced them. `get_pg_io_trend` follows one (backend_type, context) pair - a PAIR, because a hit ratio summed across contexts is meaningless when `bulkread` is a sequential scan deliberately bypassing the buffer pool - and reports read, write and extend rates per second, the interval's own hit ratio, and per-operation latency where the server measures it. `get_pg_database_trend` follows one database's temp-file spills, cache hit ratio, deadlocks and rollback share. That last one is where the hit ratio becomes usable at all: `pg_stat_database`'s counters are cumulative since the last reset, so the ratio computed from them raw is a lifetime average that barely moves, and a database that fell off a cliff an hour ago still reports 99% because of the weeks behind it. On the rig a database averaging **98.78% across the window had an interval at 78.49%**, and only the differenced series can say so. Both read the target's own `track_io_timing` before reporting a latency: it is **off by default** and was off on both rig targets, so `read_time_ms / reads` would have printed 0.000 ms - a statement about the configuration that, read as a latency, is the most reassuring wrong number on the page. Byte volumes say whether they are MEASURED (PostgreSQL 18's `read_bytes`) or ESTIMATED as operations x block size (pre-18), because the two are different quantities and 18 moves several blocks per operation. Both subjects are optional and each read says which it chose - and buffer HITS qualify a pair for that choice, so a fully cached server, which is the healthy state, still gets a chart instead of "nothing to follow". +- **The first PostgreSQL trend reads** ([#2663]) - the service shipped fourteen trend reads and not one worked on a PostgreSQL target, so "is this getting worse" - the question a monitoring tool exists to answer - had no PostgreSQL answer while every `get_pg_*` read described a single window. The data was already there: over a thousand stored snapshots of I/O and database stats on a two-day-old rig, and nothing differenced them. `get_pg_wait_trend` follows one wait event over time, per second; `get_pg_query_duration_trend` follows one statement's cost per execution, which is the regression read - a query that doubled halfway through the window still ranks where its average puts it in a single-window grid. Both take their subject as an optional parameter and choose a sensible one when it is omitted, and say which they chose. The wait trend's automatic choice deliberately skips the CPU class: `pg_wait_sampling`'s `Running` means the backend was NOT waiting and dominates any healthy server's profile, so defaulting to it would answer the opposite of the question. +- **PostgreSQL deadlocks, not just the count** ([#2661]) - we collected `pg_stat_database.deadlocks`, a number that goes up, and nothing else. PostgreSQL writes a complete report to its server log: the wait graph, every participant, and each one's full statement text. `get_pg_deadlocks` lists them with the victim, the lock modes and the resources; `get_pg_deadlock_detail` returns the graph as the server wrote it. In one respect this beats the SQL Server graph, which names the victim's statement and often leaves the other side as a handle. It needs **nothing configured on the target** - unlike plan capture there is no setting that suppresses a deadlock report, so it works on a managed fleet today; the only precondition is being able to read the log, which plan capture already established. The collector re-reads an overlapping tail on purpose so a report cut in half at one edge is whole in the next, and every row carries a hash of its graph so the same report is reported once, with a sighting count, rather than once per cycle. +- **PostgreSQL configuration, and what changed in it** ([#2658]) - `pg_settings` was never collected, so neither "what is `work_mem` set to on this server" nor "what changed last Tuesday" had an answer, and the second is the kind that cannot be recovered later at any price: a configuration history nobody recorded is not sitting on the server waiting to be read. `get_pg_server_config` reports the settings somebody actually chose, non-default first, with where each value came from and whether changing it needs a restart or a reload. `get_pg_server_config_changes` reports value changes between snapshots, old beside new. Both name `pending_restart` loudly - the state where `postgresql.conf` has been edited and reloaded but the running server is still on the old value, so the file and the server disagree with no symptom until a restart months later changes behaviour during someone else's incident. +- **Test an index from the predicate grid** ([#2612]) - right-click a row in PostgreSQL predicate statistics and ask whether the planner would actually use an index on that column. The command shipped with no caller: the only way to reach it was hand-writing a row into the command queue, which is how it was tested and is not a feature. It hangs off that grid and nowhere else, which is the shape it was scoped to - on demand only, never scheduled, driven from a row somebody is already looking at. The confirmation says what it costs the server before it runs (nothing executed, no index built, session reset), and a predicate whose estimate error is already large is flagged BEFORE the round trip, because an index does not fix a plan built on a wrong row count. +- **Azure SQL DB now reports every database's size, not just the connected one** ([#2643], raised from the field) - `sys.database_files` is database-scoped, so a Viewer pointed at `master` showed `master`'s two files and nothing else, which is correct and reads exactly like a broken collector. `sys.resource_stats` is a master-only view carrying `storage_in_megabytes` per database, so from a `master` connection the siblings now appear too - as one row each, labelled `(whole database)` with a NULL `file_id`, because that view has no per-file breakdown and a fabricated file name would make the grid look complete and be wrong. The sibling read runs through `sp_executesql`: the view does not exist in a user database and name resolution happens at parse time, so a guarded UNION still fails with 208 everywhere else. Verified against a live Azure SQL Database from both a `master` and a user-database connection. +- **Mark rows in grids** ([#2645], requested from the field) - right-click any FinOps grid and mark the selected rows **Done**, **To Do** or **Do Not Do**, so you can work through a result set and remember which rows you have dealt with and which you have decided against. Asked for on Index Analysis, where you decide index by index. Marks are held against the row objects, so they last exactly as long as the result set does: on a run-on-demand grid until you run it again, on a live grid until the next refresh. They are painted from `LoadingRow`, so no row model gained a property and no grid gained a column. +- **Test a PostgreSQL index without building it** ([#2612]) - the new `test_hypothetical_index` command plans one stored statement twice, with and without a candidate index the planner can see and nothing builds, and answers whether the planner would switch and by how much estimated cost. It closes the loop on `pg_predicate_stats`, which names columns that are filtered on heavily with nothing supporting them: without this every recommendation is an inference, and with it there is an answer. **On demand only, never scheduled** - no collector, no cadence, no sweep. Nothing is executed (`EXPLAIN` without `ANALYZE`), nothing is written, the candidate is visible only inside one session, and the session is reset before the call returns. A "no" is a real result and says so: on the verification rig a candidate on an already-served predicate left the plan and its cost unchanged, which is the answer that saves somebody from building an index. +- **The last eight PostgreSQL collectors are readable outside the Windows Viewer** ([#2629]) - `get_pg_predicate_stats`, `get_pg_index_bloat`, `get_pg_column_stats`, `get_pg_buffer_usage`, `get_pg_extensions`, `get_pg_lock_stats`, `get_pg_write_stats` and `get_pg_replication_stats`. Every PostgreSQL collector except `pg_plan_capture_readiness` now has an MCP tool and a web panel, and that one is the case the ratchet was built to tolerate - a single row of configuration state is a panel, not a question anyone asks an agent. The ratchet drops from nine to one. +- **Sampled waits and per-query OS CPU are readable from the MCP and the web dashboard** ([#2629]) - `get_pg_wait_sampling` and `get_pg_kernel_stats`. Both collectors were storing data reachable only from the Windows WPF Viewer: collected on every cycle on a Linux host and invisible from that host, and invisible to every agent. It mattered most for `pg_wait_sampling`, which #2625 had just started POINTING AT - a stock-PostgreSQL operator was told, correctly, that it answers what Aurora's wait instrumentation cannot, and sent to a panel the MCP could not read. Seven PostgreSQL collectors remain Viewer-only and a ratchet now stops that number rising. +- **Self-hosted PostgreSQL now answers "which queries cost the most"** ([#2625]) - `pg_statement_stats` read `aurora_stat_statements()` and gated itself Aurora-only, so on a non-Aurora target the single question a database monitor exists to answer had NO answer at all: the read returned "does not run on that engine, and never will". It now reads the vanilla `pg_stat_statements` view on any other PostgreSQL, chosen at query-build time - one collector, one table, identical ordinals - and reports the six Aurora-only columns as NULL rather than zero, because zero would be a claim about storage hardware the server does not have. Verified against a self-hosted PostgreSQL 17: 77 statement rows, and the read tool returns them with `storage_blks_read` and `max_exec_peakmem_bytes` honestly null. `toplevel` is guarded for PostgreSQL 13 and older, where nested tracking did not exist and every row IS top level. A target without the extension keeps the existing `42P01` -> ObjectMissing answer, which names the fix. +- **Permanent-gap messages now name the collector that DOES answer the question** ([#2625]) - "and never will" is right about the source and wrong to leave an operator with when a sibling covers it. On stock PostgreSQL the `pg_wait_stats` gap now adds that `pg_wait_sampling` answers the same question from a different source. Deliberately sparse: an entry is a promise the named collector answers substantially the same question, and a pointer at something merely adjacent is worse than none. +- **PostgreSQL plan capture reaches Aurora and RDS** ([#2538]) - `auto_explain` writes plans to the server log and nowhere else, and managed PostgreSQL has no filesystem to read: `pg_read_server_files` is not grantable and `pg_read_file` is denied. Those targets now get their plans from `DownloadDBLogFilePortion` instead. **One table, two transports**, and the route is chosen at DISPATCH rather than by gating the collector - gating would make the capability model report plan capture as a permanent gap on Aurora, which is false, and four capability tests said so. Both routes write through the SAME collector definition, so the column order, COPY command and standard prefix have one owner and cannot drift. **Three decisions worth knowing**: a cluster endpoint is resolved to its WRITER, because `DownloadDBLogFilePortion` takes an instance identifier and passing a cluster name through yields `DBInstanceNotFound` against a healthy cluster; the newest log is chosen by `LastWritten` rather than by NAME, because the filename embeds a date and sorting text puts `2026-08-09` after `2026-08-25`, which would pin the collector to a stale log forever without ever erroring; and reader and custom endpoints are REFUSED with a reason, because they round-robin across replicas so plans pulled through one would be attributed to whichever replica answered. **No credentials are stored** - the SDK's default chain finds the instance profile the service already runs under, so nothing can leak from config. The resume marker is in memory on purpose: a restart re-reads a bounded tail, which is harmless because plan rows dedup on `(queryid, plan_hash)`, and persisting it would add a schema rung that could disagree with reality after a log rotation. Redaction is not re-derived anywhere - it lives once in `PgPlanLogParser`, shared by both routes. +- **`get_pg_plans`: the plan itself, not a pointer to one** ([#2567]) - an agent consuming the MCP has no viewer to follow a reference into, so a read that answers with an id has not answered. Returns captured PostgreSQL plans grouped by shape and ranked by total time, with the plan emitted as **navigable JSON** rather than an opaque string. **`queryid` is a STRING on the wire** (#2548): it is a signed int8 spread over the whole 64-bit range, so most ids exceed 2^53 and any consumer decoding JSON numbers as IEEE-754 doubles rounds one - and since queryid is an equality join key, a rounded value does not approximate the answer, it matches nothing. The tool also **accepts** it as a string for the same reason. **The empty answers are the work**: 'no plan' has three unrelated causes with three unrelated remedies, and this names which one it is rather than guessing in prose - capture not configured (naming the specific unsatisfied facet from `pg_plan_capture_readiness`, with its remedy), capture configured but the statement never crossed `auto_explain.log_min_duration` (the healthy answer), or the plan aged out under `plan_content_retention_days`. **`captures` is named for what it counts** - the collector reads an overlapping tail of the server log so one execution can be seen twice; `get_pg_top_queries.calls` remains the authority on how often a statement ran. Nothing here re-derives redaction: the plan is stripped at collection, so there is no un-redacted copy in the store for a read to leak. +- **PostgreSQL execution plans, captured and redacted** ([#2566], part of [#2538]) - schema **v99** adds `collect.pg_plan_capture`, reading `auto_explain`'s JSON output out of the server log. **Its own table, NOT `query_plan_dim`**: that one holds SQL Server plan XML and `plan_xml_compression = none` is a documented contract with direct-SQL consumers who read the column AS XML - dropping JSON rows into it would break their dashboards with no schema change to notice, no error, and no version to key off. **No query text and no literals reach the store, and that is the load-bearing part.** `auto_explain` emits `Query Text` verbatim and `auto_explain.log_parameter_max_length = 0` does NOT suppress it (that setting covers bind parameters only - measured on #2565); literals also sit inside the plan tree in `Filter` and its relatives. So the text is dropped outright - statement identity is `query_id`, whose normalised text already lives in `pg_statement_stats` - and every remaining string is redacted before storage. **Redaction is deliberately asymmetric**: quoted literals go from every string, but bare numbers only from condition fields, because a blanket numeric strip would rewrite a relation genuinely named `transactionitems1` into `transactionitems?` - destroying identity to hide a value that was never there. Verified against 25 real captured plans: **zero literals survived**, and the digit-bearing name survived intact. `plan_hash` is taken of the REDACTED plan so the same shape recurs to the same hash whatever values it ran with. The log read is **bounded to a 4 MB tail** because #2565 measured 772 MB of log in twenty seconds at capture-everything, and truncated blocks at the window edge are skipped rather than stored half-parsed. **Self-hosted in practice**: reading the log needs `pg_read_server_files` AND an explicit `GRANT EXECUTE ON FUNCTION pg_read_file` - measured, the role alone does NOT carry it, because the function's ACL is `postgres=X/postgres` - and on Aurora or RDS there is no filesystem at all, so the panel stays empty by design and says so. +- **Which PostgreSQL columns are filtered on, and where the planner is wrong about them** ([#2603]) - `pg_index_usage_stats` records indexes that were USED; schema **v98** adds `collect.pg_predicate_stats` from `pg_qualstats`, which records predicates that were EVALUATED **including the ones with no index behind them**. That is the index-candidate signal, and it is invisible to every other collector because nothing records a scan that had no index to record. **The counts are a SAMPLE and the rate is stored on every row**: `pg_qualstats.sample_rate` defaults to `1/max_connections` - 0.01 on the verification rig, where `pg_qualstats()` returned ZERO rows on a server that had just run the queries it was meant to record. A small count means the sampler fired rarely, not that the predicate is rare, and the counts are deliberately NOT scaled up by the rate because three observations at 1% is not three hundred. **`worst_estimate_error_ratio` is a different problem from selectivity** - measured at **57.9x** on one predicate beside 1.04x on another: the first means the plan was built on a wrong row count, which an index will not fix. **Runs per database, and that is load-bearing**: the function returns cluster-wide rows keyed by `dbid`, but `lrelid`/`lattnum` are OIDs meaningful only inside their own database while `pg_class` and `pg_attribute` are per-database catalogs - unscoped, the join either silently DROPS other databases' rows or resolves them against whatever local object shares the OID and reports a confident wrong column name. Measured: two cross-database rows present, neither colliding locally that day, so the failure would have been silent loss and different OID luck would have produced the wrong name. The extension is itself per-database, so databases without it degrade to a named non-fatal skip. +- **OS CPU and disk per query on PostgreSQL** ([#2603]) - `pg_stat_statements` reports elapsed time, which cannot separate a query that was WAITING from one burning processor. Schema **v97** adds `collect.pg_kernel_stats` from `pg_stat_kcache`: the kernel's own user CPU, system CPU and bytes that reached the disk, keyed by the same `queryid` so the two halves compose. **`exec_read_bytes = 0` does NOT mean the query read nothing** - the counters come from `getrusage` and measure I/O that reached the DEVICE, so a read served by the OS page cache is genuinely zero, which was the case on every row of the verification rig while writes were not. Zero reads with high CPU is a cached workload working properly; it is not comparable to a logical-read figure, and the columns are named `_bytes` so nothing in a SQL Server codebase reads them as one. **Only top-level statements are collected**: with `pg_stat_statements.track = 'all'` a nested statement appears both on its own row and inside its caller's, and summing them double-counts every function body on the server - the rig ran `track = 'top'` where that filter is inert, which is exactly why it is pinned rather than remembered. Times convert seconds to milliseconds once, in the collector, because a column named `_ms` holding seconds is wrong by a thousand and looks entirely plausible. Ranked by **CPU**, not bytes. Cumulative like the wait profile, but this table can do better than inferring a reset from a counter moving backwards: `stats_since` is collected, so the read detects a reset from the stamp changing AND from the backwards check. **Carries `database_name`, and that is not a contradiction of v95** - the function exposes `dbid` and `pg_database` is a SHARED catalog, so the name resolves correctly from any connection; `pg_wait_sampling` omits it because its catalog genuinely cannot support the claim. +- **PostgreSQL wait analysis stops being Aurora-only** ([#2603]) - `PgWaitStatsCollector` reads `aurora_stat_system_waits()`, so every self-hosted, on-prem and plain-RDS target had **no wait data at all**, which is backwards: self-hosted is where the extension story is richest. Schema **v96** adds `collect.pg_wait_sampling` from the `pg_wait_sampling` module, attributing each wait to the `queryid` that waited so it joins `pg_statement_stats` - the "which query waited on what" question this product answers on SQL Server. **Every design choice came from running it.** `event_type = 'Activity'` is excluded at the source: measured on an idle PostgreSQL 17 the entire top of the raw profile was background processes waiting for work (`AutovacuumMain`, `LogicalLauncherMain`, `WalWriterMain`, `CheckpointerMain`), accumulating samples forever precisely BECAUSE nothing was happening, while attributed waits returned nothing at all - rank that and every healthy server reports autovacuum's idle loop as its top wait. The filter is `IS DISTINCT FROM`, never `<>`, because a NULL event means the backend was **not waiting** and `NULL <> 'Activity'` is NULL rather than true, which would silently discard every on-CPU sample; those are kept and labelled **`CPU`/`Running`** instead of stored as a blank row that sorts near the top of a count-ranked grid (23 samples across 11 backends arrived exactly that way). **`sample_count` is a tally, not a duration** - `profile_period_ms` travels beside it because a count is uninterpretable alone, and the samples-to-milliseconds multiplication happens once in the read where the header says *Est. Wait*. The profile is **cumulative**, so a counter going BACKWARDS is a reset (`pg_wait_sampling_reset_profile`, or a restart nobody chose) rather than a negative wait: the read uses the newest value whole in that case and flags `counter_reset`, where `GREATEST(new - old, 0)` would have quietly reported zero waits across a restart. **No `database_name`, deliberately** - the profile is cluster-wide and version 1.1 exposes no database column, so claiming that attribution would be the same scope error v95 removed from three other tables. Absent module degrades to a named non-fatal skip, and the Extensions panel says whether installing it would help. +- **Three PostgreSQL collectors could not say which database their rows described** ([#2599]) - found by two of our own collectors contradicting each other on a live Aurora target. `pg_table_bloat_stats`, which runs per database, reported a `pg_extension` row for `pgstattuple`; `pg_extension_availability`, which ran once per target, reported the same extension as merely `available`. Both reads were correct and they were looking at different databases, and neither row said which. **`pg_extension` is per-database while `pg_available_extensions` is cluster-wide**, so that collector was mixing two scopes in one `state` column - its own doc comment already spelled out why that is wrong, which is the part worth keeping: the documentation was right and the code did the thing it warned against. Schema **v95** adds `database_name` to `collect.pg_column_stats`, `collect.pg_index_bloat` and `collect.pg_extension_availability`, and `pg_extension_availability` now runs per database. **The reads were the user-visible half**: all three used `DISTINCT ON` over a key that did not include the database, so on any cluster carrying the same schema in two databases - the ordinary multi-tenant shape - the rows collapsed to one and the newest `collection_time` silently decided which database the grid was describing. The invariant is now pinned rather than remembered: a test walks the catalog and fails the build for any collector that returns true from `RunsPerDatabase` without declaring a `database_name` column. It found a third offender the investigation had missed - `pg_index_bloat`, added the same day. +- **PostgreSQL index bloat, MEASURED rather than estimated** ([#2561]) - and the estimator this issue proposed is the one thing that could not work. #2561 suggested porting the ioguix btree estimator, which derives an expected page count from `pg_stats` column widths. Measured: **`pg_stats` returns ZERO rows to a `pg_monitor`-only login**, because the view filters on `has_column_privilege` - the same trap that produced #2542's 88.59% reported against a true 0.50%. Meanwhile `pgstatindex` **does** run for `pg_monitor`, because pgstattuple grants EXECUTE to `pg_stat_scan_tables` which `pg_monitor` includes. So under exactly the permission set this product runs as, the exact function works and the estimator is blind. Schema **v94** adds `collect.pg_index_bloat`, collected daily per-database on primaries, matching `pg_index_usage_stats` so a read can join the two halves of one question - is this index earning its keep, and is it wasting space doing it. **`avg_leaf_density` is stored RAW and never converted to a bloat percentage**: measured across seven freshly built indexes it sits between 89.98 and 91.48, and between 87.07 and 90.81 after `REINDEX`, so `100 - density` invents about ten points of bloat on a perfect index and there is no constant to subtract because the healthy value varies per index. The read derives a reclaimable-bytes estimate against a **90% floor** rather than a full page, and presents it beside the raw figure rather than replacing it. **Ranked by reclaimable BYTES, never by density** - a tiny index at 20% density looks alarming and is worth kilobytes next to a large one at 70%, verified against seeded data where a 10 GB index at 45% showed 5.37 GB reclaimable while a 64 KB index at 20% showed 50 KB. **Only b-tree indexes are candidates, and the filter is fenced with `OFFSET 0`**: `pgstatindex` RAISES on anything else - verified on GIN, BRIN and hash, all `relation "x" is not a btree index` - so one GIN index would otherwise fail the whole collection every cycle; the planner did filter first without the fence in testing, but correctness should not depend on plan shape when the failure is total. **Nothing is silently skipped**: the function reads every page, so indexes above a 20 GB ceiling are recorded with NULL measurements and a stated reason, and the read sorts them FIRST - unknown is not zero, and they are the likeliest big win. The function is qualified `public.pgstatindex`, not `pg_catalog.` - it is an extension function and lives where pgstattuple was created, and `pg_catalog.pgstatindex(oid)` does not resolve at all (verified). +- **What is resident in PostgreSQL shared buffers, with the join every published example gets wrong** ([#2544]) - the last named slice. A hit ratio says how often the pool worked; this says what is IN it, which is what answers "does the working set fit" and "which relation is holding the memory something else needs". Schema **v93** adds `collect.pg_buffer_usage`, collected **hourly**. **Three correctness traps, all measured**: (1) `relfilenode` is NOT `oid` - the join every example writes, `pg_class.oid = pg_buffercache.relfilenode`, silently loses any table that has ever been rewritten, and after a single `VACUUM FULL` the naive join reported **0** buffers for a table holding **6,667**; `pg_relation_filenode(oid)` is the correct key. (2) Mapped catalogs carry `relfilenode = 0` in `pg_class` - measured on `pg_class` and `pg_proc` - so joining the raw column drops those too, which the same function fixes. (3) The pool is CLUSTER-wide while `pg_class` is per-database, the fourth catalog in this effort whose scope does not match its name, and worse here than elsewhere: a filenode from another database can collide with a local OID and resolve to a confidently WRONG name. So the relation name is resolved only when the buffer belongs to the connected database, and foreign buffers are KEPT with their database named and a NULL relation - dropping them would understate how full the pool is, which is the one number this exists to report. Pool totals ride on every row, taken from a window over the SAME scan rather than a second query or `pg_buffercache_summary()`, because sampling a moving pool twice puts the disagreement straight into the percentage a reader computes. **Hourly for two reasons that agree**: the full view is a scan of every buffer, measured at 6.1 ms for a 512 MB pool and scaling linearly to roughly 780 ms at 64 GB of `shared_buffers`; and a server WITHOUT the `pg_buffercache` extension records an `ObjectMissing` outcome each cycle, which at a minute grain would be thousands of rows a day of noise. That miss is now actionable rather than dead: `pg_extension_availability` (#2545) reports whether the extension is `available` and one `CREATE EXTENSION` away, so the collector that cannot run and the panel that says how to make it run are the same story. +- **PostgreSQL replication connections, with the measure that actually catches a stalled standby** ([#2544]) - the replication slice, and distinct from `pg_replication_slots` in a way worth stating: a SLOT is a promise to retain WAL and exists whether or not anybody is attached (an abandoned one is the classic way to fill a disk with nothing connected to it), while this is the live CONNECTION. A server can have a slot with no standby, a standby with no slot, or both. Schema **v92** adds `collect.pg_replication_stats`, sampled per minute on every PostgreSQL target including standbys - a cascading replica's downstream is as worth watching as a primary's, and recovery state changes on failover without a dispatch gate noticing. **Both the byte distance and the time lag are stored, because the time lag understates a replay stall badly.** Measured twice against a real standby holding `pg_wal_replay_pause()`: 26 MB behind reporting 209 ms, and 33.7 MB behind reporting 2.8 seconds. The lag columns do move - they time the round trip of the most recently replayed record - but they scale with nothing a reader can act on, so the byte distance is the proportionate measure and the one to alert on. **Four distances rather than one, because they localise the fault**: in that same stalled run `sent`, `write` and `flush` were all ZERO and only `replay` was behind, so the WAL had been shipped, written and fsynced and the problem was purely apply - while `state` still read `streaming`. No single column would have shown that. Every distance is measured from `pg_current_wal_lsn()` rather than from `sent_lsn`, so a sender that has itself fallen behind is visible instead of being used as the baseline. The read returns the latest sample AND the window's WORST, because a replica that drifts hundreds of megabytes behind each afternoon and recovers by evening reads as perfectly healthy in every single sample - and is the one most likely to be useless when somebody needs to fail over to it. Standbys are identified by `(application_name, client_addr)` rather than name alone, since an unconfigured replica reports the default `walreceiver` and two of them would otherwise average into one row; the join uses `IS NOT DISTINCT FROM` because `client_addr` is NULL for a standby on a Unix socket and an equality join would silently drop it. A standby appearing in fewer samples than were taken has been DISCONNECTING, which every other column reports as healthy, so the sample count is surfaced too. Zero rows on a replica is CORRECT rather than a fault - `pg_stat_replication` is the primary-side view - and the panel says which of those it is looking at. +- **PostgreSQL column statistics, with the customer data deliberately left behind** ([#2543]) - a plan says what the planner DID; these say why it thought that was reasonable, which for the commonest class of PostgreSQL problem (a bad row estimate producing a nested loop over millions of rows) is the difference between the symptom and the cause. Schema **v91** adds `collect.pg_column_stats`, collected daily per-database and retained a year. **The central decision is what is NOT collected.** `most_common_vals` and `histogram_bounds` hold raw column values - measured on a realistic table they came back as customer names, an identifier fragment and a list of live email addresses - so collecting them would copy customer data into the monitoring store under OUR retention, the same exposure as the `auto_explain` literals on #2538. Neither is collected, and a test asserts it against the SELECT list rather than the whole query so the comment explaining the exclusion cannot satisfy it. **Every finding the issue wanted survives without them**: `most_common_freqs[1]` is a frequency with no value attached, and skew is the entire parameter-sensitivity signal - measured, a status column whose top value covers 59.85% of the table is parameter-sensitive, and knowing WHICH value adds nothing to that conclusion. `most_common_vals` is touched only through `cardinality()`, which reads the array length and never its contents, so "how many values dominate" is answerable without copying any of them. `histogram_bounds` was also the largest column by bytes (3,251 against 228 and 140), so dropping both is cheaper as well as safer. Two traps handled: `n_distinct` is stored as a floating type and **never an integer count**, because negatives are a RATIO of row count and `-1` means distinct-is-about-every-row - an integer column would let a read print "-1 distinct values" for a unique key; and a **1 MB size floor** bounds what is the widest fan-out of any collector here (columns x tables x databases), since statistics on a tiny table cannot produce a misestimate anyone notices. **Zero rows has two causes and the panel says so**: `pg_stats` filters on `has_column_privilege`, so a monitoring login without SELECT sees nothing for a table that exists - measured, a `pg_monitor`-only role gets ZERO rows where a superuser gets all of them, and `pg_statistic` underneath is permission-denied outright, which is why Datadog ships a SECURITY DEFINER helper. Row-level security empties the view the same way. This slice deliberately installs no helper object; it collects what the granted role can see and reports the shortfall rather than presenting missing statistics as clean ones. +- **PostgreSQL lock state by mode, type and relation** ([#2544]) - the locks slice, and it does NOT duplicate the blocking read even though it looks like it should. `PgBlockingCollector` reads `pg_blocking_pids()` and stores blocked/blocker PAIRS; it never touches `pg_locks`. So the product could say *pid 4821 is blocked by pid 3390* and could not say **what lock, in what mode, on which relation** - which is the half that decides what to do. Schema **v90** adds `collect.pg_lock_stats`, sampled per minute on every PostgreSQL target including standbys, where a lock conflicting with WAL replay stalls recovery with no primary-side equivalent. **The mode IS the remedy**, measured on a real pile-up: one `AccessExclusiveLock` granted on a relation with two `AccessShareLock` requests queued behind it is a DDL blocking every reader and the fix is to kill it or wait; the identical pair shape with `RowExclusiveLock` is ordinary write concurrency needing no action at all. It also sees waiters the pairs view structurally cannot - a lock waiting on a PREPARED TRANSACTION has no live backend for `pg_blocking_pids()` to name, so it is invisible there while being exactly the thing nobody can find. Aggregated by (database, locktype, mode, granted, relation) rather than one row per lock, so the row count is bounded by CONTENTION rather than by concurrency. **A third catalog whose scope does not match its name**: `pg_locks` is CLUSTER-WIDE while `pg_class` is PER-DATABASE, so a lock held in another database resolves to an OID with no name - measured, not assumed. Both columns are stored and the row is kept, because dropping it would hide real contention and naming it from the connected database catalog would name the wrong table; `relation_oid` is `bigint` and not `integer`, since OIDs are unsigned 32-bit and one past 2^31 lands negative in a signed int. Wait time comes from `pg_stat_activity.state_change` rather than `query_start` - a backend waiting on a lock has been in its current STATE since it began waiting, whereas `query_start` also covers the work it did beforehand and would overstate every wait - and is NULL on granted rows rather than 0, because 0 reads as "granted instantly", which is a measurement where NULL is the absence of one. The read carries the CAPTURE COUNT on every row: these are samples, so three ungranted rows means something different in 60 captures than in 4, and queue size is `max(backend_count)` and never a SUM - a queue of three persisting across forty captures is three backends, not a hundred and twenty. +- **PostgreSQL extension availability: the third capability axis, and the only actionable one** ([#2545]) - engine KIND (#2536) and engine EDITION (#2511) both answer "this collector cannot run here, permanently". Extensions answer something better: "this read needs pg_stat_kcache, which this server offers and has not installed" is a **setup step**, not a wall. Schema **v89** adds `collect.pg_extension_availability`, collected daily on every PostgreSQL target and retained a year, because "when did this extension appear" is asked months later - usually right after a plan changed shape and nobody can say why. **Four states rather than a boolean**, since a boolean collapses the actionable one into the hopeless one: `installed`, `outdated` (a newer `default_version` exists - its own state because a stale extension can be missing columns a collector reads, which surfaces as a confusing 42703 rather than as a capability gap), `available` (one `CREATE EXTENSION` away), and `absent` (the server does not offer it, so an OS-level install or simply not on the menu). All four were exercised against a live PostgreSQL 17, including forcing `outdated` by installing pg_stat_statements at 1.10 against a default of 1.11 and watching `ALTER EXTENSION ... UPDATE` return it to `installed`. **The scope trap this is built around**: `pg_available_extensions` is CLUSTER-WIDE while `pg_extension` is PER-DATABASE, which their names do not say - measured on one cluster reporting pgstattuple installed in one database and not in another while the available list said yes in both. So `installed_version` is a claim about the connected database only, and the column header, the panel note and the migration comment all say so; reporting "not installed" from the maintenance database about an extension living in the application database is the obvious way to get this wrong. The row set is a **derived** half (a full outer join of both catalogs, so nothing is missed for not being on a list) plus a small **enumerated** roster that exists only because absence is not a row in any catalog and cannot be derived - a server's other extensions are still recorded, they simply carry no advice from us. **auto_explain and pg_wait_sampling are deliberately excluded and it is pinned**: both are preload-only modules that never appear in `pg_available_extensions` on ANY server including ones running them, so listing them would manufacture a permanent false `absent` - exactly the defect shipped in #2564 and fixed in #2584. They are detected instead by `shared_preload_libraries` plus their own GUCs, which `pg_plan_capture_readiness` already does. Unlike the column-statistics axis (#2543, where `pg_stats` returns **zero** rows to a `pg_monitor` role and needs a SECURITY DEFINER helper), both catalogs here read fine for `pg_monitor` - measured at 45 and 2 rows - so this ships without touching the monitored database. +- **PostgreSQL write-side collection: checkpoints, background writer and WAL** ([#2544]) - the first slice of the second-tier gaps, and the one whose difficulty is entirely the VERSION SURFACE. Schema **v88** adds `collect.pg_write_stats`, collected per minute on any PostgreSQL 14+ target including standbys. One collector for three views because they are one story: a requested checkpoint climbing against timed ones means `max_wal_size` is too small, which shows as WAL volume in the same breath and as buffers the background writer could not clean - split across three collectors the numerator and denominator of every useful ratio would sit in different tables with different collection times. All three sources are cluster-wide singletons, so a snapshot is genuinely ONE row. **The version problem is worse than the 16-to-17 view split the issue described, and it was measured across four majors rather than read from release notes**: 17 took **seven of `pg_stat_bgwriter`'s eleven columns** - five renamed into the new `pg_stat_checkpointer`, and `buffers_backend`/`buffers_backend_fsync` to `pg_stat_io`, which is a different view with a different shape and **no successor here at all**, so a rename table alone would silently drop the two columns that say whether backends are writing their own buffers. Then **18 REMOVED four columns from `pg_stat_wal`** (`wal_write`, `wal_sync`, `wal_write_time`, `wal_sync_time`) while adding `num_done` and `slru_written` to the checkpointer - so a collector written and tested against 17 raises **42703 undefined column** on an 18 target and fails every cycle. The stored shape is therefore the **UNION and never varies**: a column a major does not supply is NULL, never zero and never absent, so one store can hold a 16 and an 18 target and mean the same thing in both rows. Every gate is a `>=` floor rather than an equality, pinned by asserting a major that does not exist yet behaves like the newest known one. Post-17 names are canonical (`num_timed`, not `checkpoints_timed`). `wal_bytes` is `numeric(38,0)` and not `bigint`, because upstream types it `numeric` precisely so cumulative WAL volume may exceed 2^63. **All three `stats_reset` stamps are stored**, since `pg_stat_reset_shared` takes a target and the families reset independently - and the read differences each family only while its OWN stamp held still, because a difference taken across a reset reports an enormous positive number that looks like a catastrophe and means nothing. Verified by running the shipped SQL against live PostgreSQL 16, 17 and 18: identical 26-column shape on all three, with the NULLs landing exactly where the major lacks the column. The read was verified against seeded data too, including a mid-window WAL reset where the WAL family blanks while the checkpointer and background-writer families still difference correctly, and a single-sample window that returns nothing rather than a confident zero. +- **A captured PostgreSQL plan is an orphan without `%Q`, so plan-capture readiness now checks for it** ([#2538]) - measured on PostgreSQL 17 while working out which capture mechanism is viable, and it is the kind of thing that makes a feature look like it is working while producing nothing usable. `auto_explain` with `log_format=json` emits **no query identifier at all**, even with `compute_query_id=on` - the JSON body carries `Query Text` and `Plan` and nothing that identifies the statement. The id appears in exactly one place, the **log line prefix**, and with `%Q` present it equals `pg_stat_statements.queryid` exactly (verified both sides: `-4828029293864693941`). So without `%Q` every captured plan is an orphan that cannot be joined to the statement it came from, and nothing about the configuration looks wrong. The new `plan_attribution` facet reports it, and needs **no migration** - the row-per-facet shape means a new precondition is a query change, not a schema change. Matched with `strpos` rather than `LIKE`, because `LIKE '%%Q%'` parses as wildcard-wildcard-Q-wildcard and reports success on any value containing the letter Q; and case-sensitively, because `%q` and `%Q` are **different** prefix escapes - `%q` truncates the prefix in non-session processes and is common in real configurations. Both wrong answers were reproduced against a live server before the right one was: a prefix of `%m [%p] %q%u@%d QUEUE ` correctly reports unsatisfied. +- **The web session cookie can now carry WHO is holding it, and the signature covers it** ([#2550]) - the prerequisite for per-user sign-in, landed on its own because it is the part that is provable without an identity provider. The session cookie minted after the token exchange was `{expiryUnix}.{base64url(HMAC)}` - an expiry and nothing else - so even with OIDC standing in front of the dashboard, the subject established at sign-in was discarded at the redirect and every request after it arrived anonymous again. That is why `updated_by` on a web-authored custom view is the constant `web`: not an oversight in the write path, but the plain fact that nothing in the request carried a name. The cookie now has a second shape, `{expiryUnix}.{base64url(subject)}.{base64url(HMAC)}`, and **the HMAC covers the whole prefix before the final dot** rather than the expiry alone. That detail is the entire security content of the change: signing only the expiry and parking the subject beside it would leave identity unauthenticated, so any holder of a valid cookie could rewrite the subject and be served as anybody - a strictly worse position than the shared token this exists to improve on, because it would LOOK like identity while providing none. Pinned both directions: a rewritten subject fails, and a signature lifted from another principal's genuine unexpired cookie fails on this one. The two shapes cannot be confused for each other even though they share a separator, because the signature is always the last segment and the signed payload always everything before it - re-presenting a 3-segment cookie as a 2-segment one would need the subject to be a valid HMAC over the expiry, and the reverse needs the key. A smuggled fourth segment is refused outright rather than parsed, so two distinct cookies can never resolve to one subject. The shared-token seat reports **null, not empty string** - "the shared token did this" and "somebody signed in whose name we could not read" are different facts and a provenance stamp has to keep them apart - and an over-long subject is **refused rather than truncated**, because two subjects sharing a prefix collapsing into one seat would attribute one person's writes to another, which is worse than a failed sign-in. The subjectless form is byte-for-byte what was minted before, so cookies already in browsers keep validating instead of signing everyone out on upgrade. No producer sets a subject yet; that arrives with the OIDC wiring in [#2550]. +- **PostgreSQL targets now say whether plan capture is even possible** ([#2564]) - the shippable first slice of #2538, and the failure it fixes is SILENCE. A PostgreSQL target has no execution plans today, and somebody evaluating this against DBM on Aurora cannot tell whether the product does not support plans, their server is not configured for them, or something is broken - three different actions, and we said nothing. Schema **v87** adds `collect.pg_plan_capture_readiness`, collected hourly on every PostgreSQL target including standbys (a replica can carry a different parameter group from its writer, and gating to writers would hide exactly that). It does not capture a plan; it records whether capture is POSSIBLE, as **one row per facet** rather than a wide row - because each facet has a different remedy and collapsing them produces the single "plans unavailable" that tells nobody what to do. `library_loaded` is a custom CLUSTER parameter group plus a writer reboot on Aurora/RDS, not a `SET`, which is the misconception worth heading off. `capture_threshold` is the trap the issue was filed about: `auto_explain.log_min_duration = -1` means loaded and capturing **nothing**, indistinguishable from not-loaded from the outside but a completely different fix - so `-1`, an absent setting, and a real threshold are three distinct states with three distinct remedies. `extension_available` separates "could this server ever" from "is it doing it", which is the difference between a configuration task and a platform limitation. Two things came out of running the query against a live PostgreSQL 17 rather than reasoning about it: the remedy text had to become CONDITIONAL, because one fixed sentence per facet meant a store with the threshold at 0 reported "log_min_duration = -1 means capturing NOTHING" beside `is_satisfied = true` and contradicted itself; and `log_min_duration = 250` comes back as **`250ms`**, the GUC carrying its unit, which is why `observed` is stored as the server's own text and never cast - the `track_activity_query_size` defect this codebase has already paid for once. It deliberately does NOT claim plans are being captured: the library being loaded, capture being configured, and the plans being readable by us are three separate facts, and the last is #2566/#2567's. +- **HTTPS for the web dashboard, so the access token stops crossing the LAN in the clear** ([#2562]) - `grep -c UseHttps` across `Darling/` returned **0**: the dashboard was plain HTTP in every mode, which meant the whole token-to-cookie exchange - constant-time compare, HMAC-signed HttpOnly `SameSite=Strict` cookie, CIDR allowlist, Host-header anti-rebind, all of it carefully built - rode an unencrypted transport the moment the listener left loopback. An on-path attacker on the segment could lift the `?token=` or the cookie and get the entire collected store plus the Custom Views write path. A new **opt-in** `web.network.tls` block takes either a PKCS#12 bundle (`pfxPath`, with its password in `encryptedPfxPassword` (DPAPI), in `pfxPassword` as a `file:`/`env:` reference, or absent for an unprotected bundle) or a PEM pair (`certPath` + `keyPath`, which is what the compose distribution mounts) - **one form or the other, never both**, which is refused rather than resolved by precedence because silently serving one certificate while the operator watches the other expire looks exactly like working right up until it does not. TLS applies to the **network listener only**; the loopback listeners stay plain HTTP on purpose, since the certificate names the LAN address rather than `localhost` and loopback traffic never reaches the segment being protected. That also settles redirect-versus-refuse by mechanism: one port cannot speak both schemes, so a plain-HTTP client fails the handshake, and adding a second HTTP port to redirect from would reopen the cleartext surface the feature exists to close. The session cookie's `Secure` flag became **per-request** (`context.Request.IsHttps`) for the same reason - one host now serves both schemes, so a hardcoded `true` would mint a cookie the loopback browser refuses to send back and loop its login forever, while a hardcoded `false` would let a cookie minted over TLS ride an `http://` downgrade. Loading **fails closed** exactly as an undecryptable token does (Critical, then loopback-only) for a certificate that is missing, unreadable, ambiguous, not yet valid, or **expired** - never a silent downgrade to serving the LAN over HTTP - and because an expired certificate takes the dashboard down, the service warns for the **30 days** before that happens and logs subject, thumbprint and expiry at every start. Both loaders carry the **intermediate chain**, which is not incidental: `CreateFromPemFile` materializes only the FIRST certificate in a PEM and `LoadPkcs12FromFile` only one certificate of a bundle, so a leaf-plus-intermediate file - exactly what the config documents and what an internal CA hands you - loaded as a bare leaf and Kestrel served an incomplete chain, failing the handshake on any client that had not independently cached the intermediate. Measured with `openssl s_client -showcerts` before and after: a PEM holding two certificates put **one** on the wire, and now puts two. A self-issued root in the bundle is deliberately dropped rather than sent - a client that does not trust it is not persuaded by receiving it - and the PKCS#12 leaf is found by which certificate holds the private key, because bundle ordering is not a contract. Two things worth knowing that only turned up by running the real thing: the private key needs `MachineKeySet` on Windows (an ephemeral key loads fine and then fails every handshake, because Windows cannot use one for TLS server authentication) while macOS refuses `EphemeralKeySet` outright, and a certificate built straight from PEM carries an ephemeral key that has to be round-tripped through an in-memory PKCS#12 before Kestrel can serve it. The certificate must also carry an **iPAddress SAN** for the listen IP, because the anti-DNS-rebind Host allowlist accepts only `localhost`, a loopback literal, or that exact IP - so the DNS-name certificate an internal CA issues by default can never match, and the service now says so at startup instead of leaving it to a browser warning. `--configure-network`'s exposure summary reports the TLS state alongside listen/allowFrom. MCP is deliberately unchanged: its no-TLS rationale is that a self-signed certificate breaks real MCP clients, which is an argument about MCP clients rather than about the wire, and does not transfer to a surface whose only client is a browser. Prerequisite for [#2550] (OIDC), which most identity providers will not do over a non-HTTPS `redirect_uri`. +- **The sessions behind a pinned xmin horizon, and the ones that only look like it** ([#2540]) - `pg_xmin_horizon` already named the CLASS of thing holding the horizon back, and when that class is `session` it emitted exactly one row: a pid and a formatted detail string. Enough to know a session is the cause, not enough to do anything about it - so the product could show vacuum falling behind, show the bloat that produced, show where it ends, and give no way to see what started it. Schema **v86** adds `collect.pg_session_states`, sampled every minute on any PostgreSQL target including standbys, and `get_pg_session_states` reads it rolled up per BACKEND rather than per sample - a session parked idle in transaction for two hours is in a hundred and twenty captures, and a raw read would report one problem a hundred and twenty times while pushing everything else under the row limit. The grouping key is the collector's synthetic `(backend_start, pid)` identity, the same construction `pg_blocking_edges` uses, so a pid reused by a second backend inside the window is two findings and not one merged average - and so the same backend can be followed across both collectors. +- **`idle in transaction` is not automatically starving your vacuum, and this is the finding the read exists to make** ([#2540]) - the obvious version of this feature reports a state string and a duration, and the obvious version is wrong about half the sessions it flags. Four idle-in-transaction shapes were created on a live PostgreSQL 16.15 instance and **two of them pin nothing at all**: a READ COMMITTED transaction that has only read reports `backend_xmin` NULL because it released its snapshot when the statement ended, and one whose `UPDATE` matched **zero rows** reports both columns NULL because no transaction id was ever assigned. Only the transaction that actually wrote (holding `backend_xid`) and the REPEATABLE READ one (holding `backend_xmin`) pin the horizon, and neither column alone sees both cases. So `horizon_age` is stamped at collection from the GREATER of the two ages, and is **`-1` rather than `0`** when the session holds neither - 0 would read as "holding the newest possible xid", which is the opposite finding. The tool gates its causal claim on that column and not on the clock: a session that pinned nothing is told so in those words and explicitly refused the vacuum argument, because terminating it reclaims not one dead row. It is not waved away either - at five minutes it gets a different argument on its own terms, since a transaction idle that long is holding a connection and its locks whatever it is doing to the horizon. +- **A held horizon is a share of samples, not a flag** ([#2540]) - every write transaction is momentarily the oldest holder on the instance, because that is what a transaction IS, so a boolean "was the holder" would fire on ordinary traffic and be ignored within a day. `horizon_holder_samples` is a count against `sample_count`, and the read only calls a session a sustained holder past half the samples that saw it - reporting both raw numbers beside the verdict so a caller can disagree with the threshold rather than having to trust it. One sighting in a hundred and ninety-eight in a hundred are opposite findings and the stored data can tell them apart. The oldest holder is also force-included in every capture regardless of the duration floor, and is chosen over the FULL activity set rather than the reportable subset: a young transaction can be the oldest holder, and crowning a filtered survivor would name the wrong session precisely when the real one fell below the floor. +- **No statement text is stored, and that is a decision rather than an omission** ([#2540]) - `pg_stat_activity.query` is the statement as submitted, literal parameter values inline; a probe session on the live rig came back carrying its literal argument verbatim. `pg_blocking_edges` stores that text because blocking is an exceptional condition where the rows are rare and the text is the finding. This table fills on a DURATION FLOOR that a perfectly ordinary application can cross, so the same column would mean routinely accumulating user data in the store to answer a question that does not need it. What it stores instead gives up nothing: **`query_id`**, PostgreSQL's own normalised statement fingerprint, which joins to `pg_statement_stats` whose text is already normalised to `$1` placeholders - full actionability, zero literals - plus **`command_tag`**, the leading SQL keyword matched against a closed whitelist. A whitelist and not a substring, because plenty of ORMs prepend a comment block, so the first token of a real statement can be a comment opener followed by anything the application chose to put in it and a leading-N-characters rule would carry a `WHERE` clause's literals straight out. +- **The permission that gutted this feature silently, caught before it shipped** ([#2540]) - the same check [#2542] made for `pg_stats`, and the answer is worse here because there is no error. Measured against a least-privileged role on a live target: without `pg_monitor` PostgreSQL does **not** refuse the read - for every backend the login does not own it returns the row with all but SIX columns NULL. Measured column by column on PostgreSQL 16.15 rather than enumerated from memory: what survives is `pid`, `application_name`, `datname`, `usename`, `backend_xid` and `backend_xmin`, while `state`, `state_change`, `xact_start`, `query_start`, `backend_start`, `wait_event_type`, `wait_event`, `backend_type`, `client_addr`, `leader_pid` and `query_id` all come back NULL and the query text is replaced by an insufficient-privilege literal - **so the two that survive are exactly `backend_xmin` and `backend_xid`**. So the horizon still reads as pinned and every column that could explain it is gone. On one capture over the same nine backends the privileged role saw four sessions idle in transaction and the unprivileged one saw **zero**, because `is_idle_in_transaction` is derived from a state that is NULL. `state_is_redacted` is therefore stamped per row from the privilege literal and **not** from `state IS NULL` - background workers legitimately report a NULL state under full privilege (`checkpointer`, `walwriter` and the autovacuum launcher all do, measured), so NULL alone cannot tell a permission problem from a background process. Redacted rows outrank every severity with a band of their own rather than being painted healthy or critical, the same treatment `estimate_unavailable` gets on the bloat surface. Unlike that one, plain `pg_monitor` is enough here - no `pg_read_all_data` needed. +- **The denominator, and a defect only a live run could find** ([#2540]) - `pg_stat_activity` is a point-in-time view and PostgreSQL records nothing about session state unless something asks, so a transaction that opened and closed between two samples is genuinely invisible - which the read says out loud rather than letting "no long transactions" be read as "none happened". Capture counts come from `collection_log` and not from the data, because this is an exception surface like `pg_blocking_edges` where zero rows is the HEALTHY answer and probing the table would report a well-monitored server as uncollected ([#2508]). Every capture also carries its own session counts and its pre-limit reportable count, so two idle-in-transaction sessions out of six connections is distinguishable from two out of four thousand, and a capture truncated by the row cap is self-evident. The defect: the first draft of the collector query contained a literal SQL comment opener **inside** a comment. PostgreSQL block comments **nest**, so it opened a nested comment that never closed and the entire collection failed to parse - every cycle, on every target. No text assertion could see it and nothing about reading the C# reveals it; running the shipped string against a live instance is what caught it, and a nesting-depth test now pins it. Version floors were measured the same way: `pg_attribute` read for the view on a live PostgreSQL **13.23** instance shows 21 columns including `leader_pid`, `backend_type`, `backend_xid` and `backend_xmin`, and **no `query_id`** - so that one column is substituted with a typed NULL below 14 rather than gating the collector off, and the PG13 form was executed against that instance. +- **PostgreSQL index usage, with the half that decides whether an index can actually go** ([#2541]) - nothing collected `pg_stat_user_indexes`, so SQL Server users had `get_index_usage` and PostgreSQL users had nothing. An unused index costs MORE here than on SQL Server: it costs the same write throughput and space **plus** it participates in every `VACUUM`, and index cleanup is a large share of vacuum work - so a dead index on a hot table slows the maintenance that [#2530]'s vacuum surface exists to watch. Schema **v84** adds `collect.pg_index_usage_stats`, collected daily per database on writers only, and `get_pg_index_usage` reads it. The scan counts are the easy half. The half that makes it shippable is that **"unused" is not "droppable"**, and the tool refuses to say otherwise: every row carries whether the index backs a primary key, a unique or exclusion constraint, or the table's replica identity, whether it is partial, an expression index, or INVALID, and its full `pg_get_indexdef` text - because the single most common zero-scan index on any schema is a unique index enforcing a constraint, and enforcing uniqueness is not a scan of the kind `idx_scan` counts. Advice derived from the counter alone tells somebody to drop their primary key. It also refuses to call anything unused on a window it has fewer than two samples of, since PostgreSQL records no index creation time anywhere and how long we have been watching is the only evidence an index is old enough to judge. +- **The scan count only a monitoring store can produce** ([#2541]) - `pg_stat_user_indexes.idx_scan` is cumulative since the statistics were last reset, which is the only figure a live query can give: an index with nine million lifetime scans and none in the last ninety days reads as heavily used. The read reports **both** - the lifetime counter, so it cannot appear to disagree with psql, and `scans_in_window` differenced across the stored samples, which is the number that decides anything. That difference is clamped at zero and the clamp is not decoration: measured against a live store, a fixture whose counter runs 9000 -> 50 -> 100 across a statistics reset returns **-8900** unclamped, and a negative scan count sorts straight to the top of a list ranked least-used-first. The clamp alone would then hide the reset, so `stats_were_reset_in_window` travels beside it, taken from the server's own `stats_reset` timestamp with `IS DISTINCT FROM` rather than `<>` - `stats_reset` is NULL until a database's first-ever reset, so `<>` evaluates to NULL there and misses precisely the reset the counter arithmetic cannot see either. Retention is **90 days** because the retention window IS the evidence: an index can only be called unused for as long as we have been watching it, and 30 days cannot clear a monthly report. It still cannot clear an annual one, which the read says rather than glossing. +- **PostgreSQL table bloat: the damage, next to the cause chain we already collected** ([#2542]) - `pg_autovacuum_stats`, `pg_wraparound_stats` and `pg_xmin_horizon` between them collected the entire CAUSE chain for bloat and nothing that measured what it cost, so a user could watch vacuum fall behind and had no way to see the damage - which is the question they ask immediately afterwards. Schema **v85** adds `collect.pg_table_bloat_stats`, collected hourly per database on writers only, and `get_pg_table_bloat` reads it. **Hourly matches `pg_autovacuum_stats` deliberately rather than by copying**: this measures the damage whose cause that one measures, and "vacuum fell behind at 14:00 and bloat grew" is only a sentence the data can support if both are sampled on the same grain. +- **The bloat measurement decision, settled by measurement rather than by reasoning** ([#2542]) - three options, and the cheap one wins on a margin that was measured, not assumed. `pgstattuple` is exact and reads the **whole relation**: 11.8 ms warm on a 54 MB fully-cached table, i.e. roughly 4.6 GB/s in the best case, so a 1 TB table costs minutes per table per cycle even entirely in memory - and it is an extension, absent unless somebody installed it. `pgstattuple_approx` measured 0.58 ms on the same table, but that is its BEST case and not a representative one: the pages it must still read are exactly the recently-written ones, so on the churning table you actually care about it degrades toward the full scan, and it shares the availability problem without solving the cost one. The **statistics-based estimate** ships: it reads `pg_class` and `pg_stats` only, never touches the relation, and so costs by table COUNT rather than by table SIZE - **44 ms for a 2,001-table / 854 MB database against 860 ms for `pgstattuple` over the same tables**, and that 19x is a floor on the ratio because those tables were tiny and fully cached. Erik's inclination in the issue was the estimate on cadence with `pgstattuple` available on demand for one table; the on-demand half is **not implementable as a tool**, because the MCP surface is strictly read-only over collected rows and has no ad-hoc path back to a monitored server. So each row instead records whether `pgstattuple` is installed on that database and ships the exact command that would settle it, which is the reachable form of the same idea. +- **The bloat estimate is suppressed, not captioned, when its inputs cannot be trusted** ([#2542]) - a bloat number that is quietly wrong gets used to justify a `VACUUM FULL` on a production table, so how wrong it can be was measured rather than hedged. With current statistics the estimate was within about **2 percentage points** of `pgstattuple` on every table tested (churn-bloated 75.14 vs 74.82, delete-bloated 60.73 vs 59.52, clean 0.58 vs 0.50). With **stale** statistics it was wrong by **81 percentage points**: two byte-identical 8,998-page tables with a true bloat of 10.93% estimated 92.64% and 11.01%, differing only in whether `ANALYZE` had run after a widening UPDATE, with nothing in the arithmetic to show it. The stale table had modified 200% of its rows since its last analyze and the fresh one 0%, so `mods_since_analyze` is carried as the trust signal - and it is a sound proxy rather than a coincidence, because width statistics can only go stale THROUGH modifications. Above a fifth of the table modified since the last analyze, the number is withheld and the row says why; the threshold is set low deliberately, because the failure is asymmetric - a needlessly cautious estimate costs a second look, an over-trusted one costs a table rewrite. Every field carrying the figure is named `_estimate` in the store as well as in the read, so the qualifier cannot be lost between the two. +- **The permissions trap that would have shipped a tool telling everyone to VACUUM FULL everything** ([#2542]) - `pg_stats` is filtered by `has_column_privilege(..., 'select')`, and **`pg_monitor` - the one grant the PostgreSQL runbook says Darling needs - does not confer SELECT on user tables**. Measured against a `pg_monitor`-only role on a live PostgreSQL 16 target: **zero** `pg_stats` rows visible, and the estimator did not fail. It returned a confident **88.59%** for a table whose true bloat is 0.50%, **95.03%** for one that is really 74.82%, and **22.57%** for one that is really 0.46% - every row an argument for rewriting a production table. `estimate_unavailable` is therefore load-bearing rather than advisory: it is TRUE in exactly that state and the read suppresses the number outright, because a footnote under a large red percentage is not a safeguard. Adding `pg_read_all_data` (PostgreSQL 14+; the role does not exist on 13, where explicit `GRANT SELECT` is the only route - both verified) restored byte-identical agreement with the superuser's numbers, and the runbook now carries that grant with the reasoning. What survives the suppression is the **dead-tuple fraction**, computed from the server's own counters with no width model and no SELECT needed - narrower than bloat, since it counts dead tuples and not the free space a completed vacuum left behind, but measured rather than inferred, so a permissions gap degrades the answer instead of removing it. +- **A Storage tab on both front ends, and neither read is MCP-only** ([#2541], [#2542]) - the two new reads land together on a new **Storage** tab in the web dashboard and in the WPF viewer, with the same id and the same grouping, because both answer one question: where the space went and whether it is earning its keep. Deliberately NOT on the Vacuum tab despite bloat being what vacuum lag costs - that tab is the cause chain read in causal order, and dropping the damage into the middle of it breaks the sequence that makes those three panels one story; the bloat panel points back at it instead. Both ratchets that would have caught an MCP-only ship were green before the panels existed and are green after: the web pin derives its "must have a home" set from the read dispatch, the viewer pin derives its own from `CollectorCatalog`. The viewer grid renders a suppressed estimate as a **dash** rather than a number, re-deriving the suppression rule rather than trusting a flag, so the two surfaces cannot disagree about whether a figure is publishable. +- **Version floors for both collectors, measured against six live majors rather than read from documentation** ([#2541], [#2542]) - every column was checked by listing the actual catalogs on PostgreSQL **13, 14, 15, 16, 17 and 18**, and both shipped queries were then executed on all six. Exactly one floor exists: **`last_idx_scan` is PostgreSQL 16+**, proven by running the gated-on form against 13/14/15 and getting `ERROR: column i.last_idx_scan does not exist` - which fails the whole collection every cycle, not just that column. It is substituted with a typed NULL rather than omitted, so the row shape does not change across a mixed-version fleet, and NULL is the honest value because a PostgreSQL 15 server genuinely does not know when an index was last scanned. Everything else - `pg_stat_user_indexes`, `pg_statio_user_indexes`, `pg_index`'s droppability columns, `pg_constraint.conindid`, `pg_stats`, `pg_class.reltuples`/`relpages` - was confirmed present on all six, so that is stated as checked rather than assumed, and the bloat query carries **no version gate at all** because it produced identical output on every major. The MAXALIGN detection the estimate depends on names `aarch64` and `arm64` explicitly (Graviton RDS and Aurora are aarch64) and stores the value it actually used, so a platform where the detection is wrong shows up in the data instead of skewing every row silently. +- **Both new collectors are fan-outs, and the cost of that was budgeted rather than discovered** ([#2541], [#2542]) - `pg_stat_user_indexes` and `pg_stats` are both scoped to the connected database and PostgreSQL has no cross-database read, so these are the second and third per-database fan-outs after `pg_autovacuum_stats` - and on PostgreSQL a fan-out means one **connection** per database per cycle. Index usage runs **daily**, inherited from `index_object_stats` (the SQL Server collector answering the same question) rather than from the PostgreSQL sibling: "has anything scanned this index" is a structural question, and an hourly sample would record the same catalog facts 24 times a day at 24x the connections. Both carry a size floor as the cost control - 64 KB for an index, 1 MB for a table, below which there is nothing to model, since the whole argument for dropping an unused index is the write, space and vacuum cost it carries and all three scale with its size. Measured on a 2,001-table / 6,005-index database, the floors removed 2,000 tables and 2,001 indexes. An INVALID index is kept regardless of size, and that escape hatch was verified live: a 16 KB invalid index came through a 64 KB floor. +- **A fourth kind of nothing: `precondition`, for a setup step somebody can actually change** ([#2546]) - the miss vocabulary had three words and none of them fitted a Query Store that is switched off, a capture session that is not running, a PostgreSQL extension that was never created, or a login refused msdb. `empty` claims we looked and there was nothing; `unavailable` sends the reader to collection health, where they find a collector that is running and doing its best; `not_collected` is the PERMANENT answer and tells them to stop looking. The store already held every one of these facts - the runners classify a denied grant, a missing source object and a disabled feature out of the general ERROR bucket and write the remedy into `collection_log.error_message`, and the hourly `query_store_health` collector records `actual_state` per database - and **no read reported any of it**. Two Query Store reads were literally guessing in prose ("Query Store may not be enabled on target databases"), which is equally true of a server where it is on and the window simply reached past retention. Now the read states it, names the databases, and carries the remedy for the state they are actually in - OFF, READ_ONLY and ERROR each get a different one, because telling an operator to "turn it on" about a database whose Query Store is already on and broken is an ALTER that silently no-ops - or quotes what the server itself said, including the `CREATE EXTENSION` a 42P01 already stores. Crucially it is evaluated at READ time, not at gate time: an `AppliesTo` gate is decided once when the connection is made and would go on reporting a precondition after somebody satisfied it, which is the worst possible direction to be wrong in. Wired on five reads per SKU (`get_deadlocks`, `get_blocked_process_xml`, `get_running_jobs`, `get_long_query_completions`, `get_query_store_top`) plus Darling's `get_pg_top_queries`, always AFTER the engine-capability answer so a permanent gap still wins - an order a `??` chain cannot enforce on its own, so a source scan pins it across both trees. +- **The WPF viewer renders PostgreSQL tabs at a PostgreSQL target, and stops rendering nineteen empty SQL Server ones** ([#2530]) - `git grep IsPostgres` across `PerformanceMonitor.Darling.Viewer` returned **nothing**, so an Aurora target opened on the SQL Server correlated-lane Overview and got the whole nineteen-tab strip - tempdb, Trace Flags, Query Store, Plan Viewer, Always On - nearly all of it permanently empty and several tabs meaningless for the engine. It now gets **six**: Overview, Activity, Vacuum, Waits, I/O and Replication, the same six the web dashboard chose in [#2547], with the same ids and the same grouping so the two front ends do not teach one engine two shapes. All **nine** PostgreSQL collectors land on exactly one tab, and that placement is DERIVED from `CollectorCatalog` in both directions - a tenth collector turns the pin red naming itself, which is the check that did not exist while eight of them shipped MCP-only for three releases. The `ViewerCollectorCoverageTests` allow-list, which carried all nine as tracked debt, is now **empty**. +- **The two Aurora-only PostgreSQL panels are SHOWN on stock PostgreSQL, not hidden** ([#2530]) - `pg_wait_stats` and `pg_statement_stats` both gate on `IsAurora`, so on a stock target they can never have content. Each panel prints `CollectorEngineCapability.NotCollectedMessage` instead - the same sentence the MCP surface and the web dashboard print, naming the server, the engine, the collector and the exact `aurora_stat_*` surface, ending "and never will". The defect being closed is UNEXPLAINED emptiness, not emptiness; hiding them would also make the tab strip a different shape on two PostgreSQL servers in one fleet. The Overview tab is built from the CATALOG rather than from `collection_log` for the same reason: a gated-off collector writes no log row at all, so a log-driven grid would have dropped precisely the row an operator most needs explained. +- **One copy of the PostgreSQL read SQL, shared by the MCP surface and the desktop viewer** ([#2530]) - the nine `DarlingPg*Reader` classes moved from the service's MCP folder into `PerformanceMonitor.Darling.Storage`, which both front ends reference. Copying them into the viewer would have meant a second copy of, among others, a 200-line recursive blocking walk whose revisit guard, root attribution and truncation flag were each a separate review finding - and the copy that drifts is never the one being read. No query text changed; every existing reader test still pins the same constants. +- **The viewer's server rows carry the engine discriminator, from BOTH server reads** ([#2530]) - `DarlingServer` gained `EngineKind`, `EngineEdition`, `IsPostgres` and `EngineDescription`, and both `ServersSql` and `ManagedServersSql` select `engine_kind` and `sql_engine_edition`. Both, because the sidebar uses `ManagedServersSql` on any seeded store - i.e. every real deployment - so a column added to only the other one would have left every PostgreSQL target on the SQL Server tab set while a unit test passed. Only a POSITIVE claim switches: a null or unrecognised token keeps the SQL Server tabs and shows NO engine badge, because the tabs such a server gets are a default rather than a finding. +- **PostgreSQL temp-file spills, cache hit ratio, deadlocks and the rollback ratio - one collector, one read** ([#2539]) - nothing touched `pg_stat_database`, so four separate questions had no answer on a PostgreSQL target. The one that mattered: **temp files are work `work_mem` could not hold**, which is the most common reason a PostgreSQL query is slow for a cause its plan shape does not show - and a spilling query looks identical to an expensive one in the statement statistics. `pg_statement_stats` does carry per-query `temp_blks_written`, but it reads `aurora_stat_statements()` and is **Aurora only**, so on stock PostgreSQL there was no temp-file evidence anywhere; on Aurora it is a top-N over statements `pg_stat_statements` still retains, and this is the ground truth those are a share of. Schema **v83** adds `collect.pg_database_stats`, the new `pg_database_stats` collector reads the view per minute, and `get_pg_database_stats` differences it across a window. Every major and every target shape: the newest column selected arrived in PostgreSQL 9.2 (verified live on 13, 16 and 17), it is core rather than Aurora, and a standby is included deliberately because a sort on a read replica spills exactly the way it does on a writer. +- **A PostgreSQL statistics reset is reported AS a reset** ([#2539]) - `pg_stat_database`'s counters are cumulative since the last `pg_stat_reset()`, so a reset silently zeroes them, and the two obvious ways to difference them are both wrong: last-minus-first goes **negative**, max-minus-min **spikes** to the pre-reset lifetime total. Both look like measurements. The read clamps per interval so neither can happen, and then says so: `stats_reset_count` from the server's own per-database `stats_reset` moving, `counter_rewind_count` from a counter falling below its predecessor, and a per-database note that the window totals are LOWER BOUNDS. Both signals are carried because neither sees every reset - a reset followed by enough activity to climb back past the old value inside one collection interval leaves every difference positive, and only the timestamp knows. The sharpest case is a database's FIRST-EVER reset, where `stats_reset` moves NULL -> timestamp: the first sample of a series is excluded by its ROW_NUMBER rather than by its `LAG` being NULL, because `LAG(stats_reset)` is NULL both when there is no previous row AND when the previous row had never been reset - the ordinary state - so a guard on it would miss the one reset nothing else can see. Rows are stored **per database** for the same reason: `stats_reset` is per database, so an aggregate would have to pick one timestamp for rows that legitimately disagree, and a single database's reset would corrupt the cluster's delta with nothing left in the data to say so. +- **The web dashboard renders PostgreSQL tabs at a PostgreSQL target, and stops rendering twelve empty SQL Server ones** ([#2530]) - the per-server page's `SERVER_TABS` was a FLAT list applied to every server, so an Aurora instance got Wait Stats, CPU, Memory, Blocking, File I/O, Queries, Configuration, Config Changes, tempdb and the rest - nearly all permanently empty, several meaningless for the engine - while the eight `get_pg_*` reads that DO have data were reachable only through MCP. The registry is now two registries, and `serverTabsFor(card)` chooses from the fleet card's server-derived `is_postgres`. A PostgreSQL target gets **six** tabs: **Overview** (freeze headroom, xmin horizon, autovacuum backlog and replication-slot vitals over findings and collection health), **Activity** (`get_pg_blocking` with its capture denominator, top query shapes, and [#2539]'s per-database temp-file spills directly beneath them, because a spill is the explanation a statement's own stats cannot give), **Vacuum**, **Waits**, **I/O** and **Replication**. Six against twelve is the design and not a shortfall: parity was explicitly not the constraint, because bloat, wraparound, the xmin horizon and autovacuum have no SQL Server analogue and are what actually pages a PostgreSQL DBA. **Vacuum is one tab on purpose** - an old xmin horizon starves vacuum, starved vacuum falls behind on freezing, freezing falling behind ends in wraparound; read separately each looks survivable and together they are one escalating story, so the three reads sit in that causal order under a note that says why. **Waits is SHOWN on stock PostgreSQL** even though `aurora_stat_system_waits()` means it can never fill there: [#2532] taught that read to answer `not_collected` naming the server, the engine, the collector and the exact Aurora surface and to say the gap is permanent, so the panel explains itself - and hiding it would have made the tab set change shape between two PostgreSQL servers in one fleet and made the one Aurora-only capability invisible from a stock instance. A **NULL** `engine_kind` keeps the SQL Server set, unchanged: only a positive PostgreSQL claim switches, because absence is not evidence for either engine. The bar and the grid wait for the ONE `/api/fleet` read the page already made rather than flashing the wrong twelve tabs and correcting them. The card also gained `engine_description` (the engine's name in words, from `MonitoredEngineKind`), so the header can badge the engine without any surface owning a second copy of the vocabulary - and `vizStat` gained the `emptyText` guard `vizLine` already had, because several PostgreSQL reads answer their HEALTHY case with a data body carrying prose and none of the summary keys, which used to render as a row of em-dashes. +- **The store records what ENGINE a monitored target is** ([#2530]) - `collect.servers` recorded `sql_engine_edition` and `sql_major_version` and nothing that said a target was PostgreSQL. `SERVERPROPERTY` does not exist on PostgreSQL, so a PostgreSQL target landed at engine edition **0** - byte-identical to a SQL Server that has never completed a connect. Two different facts, one representation, and no reader could tell them apart. Schema **v82** adds a nullable `collect.servers.engine_kind` carrying `sqlserver`, `postgres` or `aurora-postgres`; a token rather than an `is_postgres` boolean because Aurora versus stock PostgreSQL is already a fact the collectors gate on (the `aurora_stat_*` surface has no equivalent in any core PostgreSQL version) and a boolean cannot grow a third value. Nullable with no default and no backfill: a row written before the rung genuinely does not know, and defaulting to `sqlserver` would have every PostgreSQL target assert it was SQL Server until its next connect. The registry upsert stamps it on every connect, so it populates itself within one collection cycle. The `/api/fleet` card and `get_fleet_overview` now carry `engine_kind` plus the derived `is_postgres` / `is_aurora`, which `docs/uat-onboarding.md` §3.4 named as the blocker for any PostgreSQL panel. **No UI branches on it yet** - a PostgreSQL target still renders SQL Server tabs on both SKUs; that is the next wave of [#2530]. +- **Ten viewer surfaces the browser and MCP could not reach now have read endpoints** ([#2484]) - `/api/read/*` is machine-derived from the MCP tool catalog, so a read with no tool was absent from BOTH surfaces: the WPF viewer on a Windows desktop was the only way to that data, and an agent asked about it got nothing. Every one of them was a missing ENDPOINT rather than missing collection - the store already held the rows. Now served on both SKUs: `get_collection_log` (the raw per-run log under `get_collection_health`'s seven-day rollup, which is what you reach for when the rollup reads HEALTHY and collection still looks wrong), `get_current_waits_trend`, `get_blocking_stats` (severity, where the existing trends only counted incidents - ten one-second blocks and one ten-minute block are the same count and a different problem), `get_health_parser_significant_waits` (the ninth member of a family whose other eight already had one), the three Performance Trends siblings, `get_query_store_regressions`, `get_query_heatmap`, `get_lock_wait_trend` and `get_daily_summary_range` (without which the web calendar could only ever show today). One correction came out of it: the viewer's execution-count trend was a DUPLICATE of a number `get_query_duration_trend` already returned rather than a missing sibling, so it did not become an eleventh read - what it got instead was `executions_per_second`, because that count shipped truncated to a `long` and a server running 0.4 executions a second reported **zero** and read as idle. +- **An in-place upgrade can now remove the files the new build stopped shipping, and always names them ([#2529])** - an upgrade is an OVERLAY: `Expand-Archive -Force` writes what the new build ships and deletes nothing else, so a file the old version had and the new one dropped stayed in the install tree forever, and [#2185]'s layout report could not see it either - that report walks top-level DIRECTORIES, so a stale DLL in the install root, or inside `viewer\` or `runtimes\`, is structurally invisible to it. **The measurement came before the code**, because the whole question was whether this has ever actually happened: diffing the file lists of consecutive release zips says the Lite package dropped **44 shipped files across twelve consecutive releases** - 43 of them in ONE step, where a target-framework move stranded every `runtimes\*\lib\net8.0\` assembly including two copies of `Microsoft.Data.SqlClient.dll`, and one lone assembly in a later step - while the Darling package has dropped none in the four steps it has existed for. So it is rare, bursty, tied to a packaging or framework change rather than to ordinary development, and when it happens it arrives forty files at a time **in a .NET probing directory**, where a stale assembly is a candidate for loading rather than inert clutter. **The authority is a manifest derived from the payload, never a list anybody maintains**: after each successful copy `upgrade-darling.ps1` records the files that copy laid down, and the next upgrade diffs its own payload against it - so every path it can name provably came out of one of our own zips. That disposes of the entire false-positive class the obvious approach has. It cannot warn about `darling.json`, the DPAPI blobs, the `.bak-*` config copies, the `_rollback_manual_*` backups or `pg-runtime`, because none of those was ever in a payload and so none of them is ever in the manifest - a property of where the list comes from, not a list of exceptions somebody keeps up to date, which is [#2525] again with a new subject. **It fails safe and says so**: no manifest yet, one that will not parse, one that disagrees with its own `file-count`, or a source it cannot read all mean remove NOTHING, on its own line, with the reason. `-RemoveStaleFiles` is the switch that turns naming into deleting and it is **off for the first release** - the check reports on every upgrade either way, and a delete that runs inside a monitoring host's install directory should spend a few deploys showing operators its answer first. Directories emptied by their own removal are pruned with NON-RECURSIVE deletes, deepest first, which is not tidiness: the layout report identifies a satellite-resource directory structurally, an EMPTY one fails that test, and leaving the shell of `de\` behind would turn one stale file into a permanent startup warning. Guards at three depths, each proved RED by reverting only itself: the `pg-runtime` **prefix** rule (a prefix, not the two names we ship, because an enumeration one entry out of date here costs the store), an empty payload selecting nothing rather than everything, an empty shipped-set refusing the delete outright, case-INSENSITIVE matching so a respelled `App.js`/`app.js` is never deleted after the copy just wrote it, and containment resolved against the filesystem rather than the string. Pinned by handing the shipped delete a deliberately POISONED list holding the store, the config, its backups, a credential blob, a rollback backup and a relative escape, then asserting against the **disk** rather than the function's own account of what it did. `upgrade-darling.ps1` also now has a parse-and-AST pin in CI, because `$x | ForEach-Object { } -join ', '` binds `-join` as a PARAMETER, parses clean, and throws at runtime - which in this script means after the service has been stopped. **That pin went red on its first CI round, against lines nobody had touched, and it was right**: `upgrade-darling.ps1` did not parse under Windows PowerShell 5.1 at all - the default `powershell.exe` on Windows Server and the one SSM runs. The file is BOM-less UTF-8, 5.1 decodes such a file with the machine's ANSI code page, and the third byte of a UTF-8 em dash becomes `”` in CP1252 - which PowerShell honours as a **closing double quote**. So every em dash [#2525] had placed inside a double-quoted message terminated its string early, and the script failed to parse before doing anything. Reproduced byte for byte, down to the same "The '<' operator is reserved for future use" CI reported. The five em dashes already in `install-darling.ps1` and `fetch-pg-runtime.ps1` were harmless only because they happened to sit in COMMENTS, where a mis-decoded character is never parsed - not a property anyone can maintain by eye. Fixed by making all three shipped scripts **ASCII-only** rather than by adding a byte-order mark: a BOM works and then decays, one careless save from being gone and invisible when it is, while "no byte above 127" holds under every code page and is a rule anybody can check. Pinned both ways, each proved red by planting a single em dash back into one double-quoted string. Never shipped in a release - [#2525] is in this same unreleased section. +- **A supported in-place upgrade script, because the step that had no script was the one accumulating 5.48 GB** ([#2525]) - `install-darling.ps1` registers a service and `uninstall-darling.ps1` removes one; laying a new build over an old one was a procedure that lived in people's heads and in ad-hoc SSM scripts, and the part of it nobody remembered by hand was deleting the backups from the last twenty deploys. A dogfood box was found carrying **46 `_rollback_manual_*` directories, 5.48 GB, the oldest three weeks old** - and the [#2185] install-location report, doing exactly its job, named every one of them on every start. **46 warnings for something our own procedure created is a guard that has stopped guarding by being too loud**: a real layout problem - a stray DLL, a half-extracted upgrade - arrives as warning 47 in a list of 46 identical ones, which is the same failure as a pin that never bites wearing the opposite clothes. So both halves changed. `upgrade-darling.ps1` ships in the zip and owns retention **at deploy time**, keeping the newest `-KeepRollbacks` (3 - enough for "the release, the one before, and the one before that"; the fourth can only roll back to a version nobody wants) and pruning the rest AFTER taking the new backup, so the tree never passes through a moment with fewer rollback points than retention promises. It verifies the source zip's SHA256, refuses to copy while anything is still running out of the install tree - **naming what, never killing it**, because the bundled PostgreSQL lives under `pg-runtime` and a blanket kill takes the store down - refuses to run from the install directory itself (the copy would overwrite the script PowerShell is reading), confirms `darling.json` is byte-identical afterwards, and is safe to re-run at every step: a backup taken in the last `-BackupWindowMinutes` is REUSED rather than replaced, so a re-run after a failed extract cannot overwrite the good pre-upgrade copy with a copy of a half-upgraded tree. And the service now RECOGNISES the convention rather than deleting anything - it never deletes what it did not create - reporting the whole set on **one** line with a count, a total, the oldest, and `upgrade-darling.ps1 -PruneOnly`, informational within retention and a warning past it. That half matters even where retention is running, because the boxes carrying the backlog got it before any script pruned anything and no upgrade removes it. The naming convention is a single shared constant with the script's own predicate executed against the service's in CI, because two spellings of one convention is how this happened in the first place - and the drift would be silent, since each of those 46 warnings was true. Two defects were found by running the script's functions against planted trees rather than reading them: a backup dated in the FUTURE (clock skew) read as "recent" and would have suppressed the deploy's backup entirely, and a failure while composing a LOG LINE was caught by the delete's handler and reported as a failed delete - two removed and two failures, for the same two directories. Nothing here could be tested on real Windows from the machine that wrote it; the pins run in CI. +- **The web dashboard gets a per-query drill-down, which is what makes `get_query_trend` reachable there at all** ([#2520]) - the read answers "is *this* query getting worse", the question you ask second and the one that decides whether you act, and it keys on a **required** `query_hash` plus a **required** `database_name`. Every panel on the per-server page fetched with nothing but a server and a window, so there was no `query_hash` anywhere on the surface to send and `/api/read/get_query_trend` could not be called from a browser - the ONLY read in the catalog whose absence was missing UI rather than a stated product boundary (the other nineteen each have one: a dedicated endpoint already serves the data, a write path the web deliberately lacks, a payload shape the table renderer cannot draw, or a desktop capability a web imitation would be worse than). The Queries tab's Top Queries table now carries a picker, in the shape the Wait Stats tab established: **the options are the rows of the table directly above it**, and an option's VALUE is that row's index into the array the table rendered - so the query trended is the same array element the reader is looking at, not a name that matched. That is deliberately not a "list every query" read: `get_wait_types` is unread on the Wait Stats tab for exactly this reason, because a full distinct set would offer queries absent from the table and make the two disagree. The chart carries avg CPU and avg elapsed only - both are milliseconds, so one y-domain holds them honestly, and both survive the hourly rollup [#2353] falls back to past the raw tier's four days; executions, DOP and the plan hash go in a grid of the per-collection snapshots below it, where the read's NULL reads as the blank it is instead of flattening a chart to a zero nobody measured. When the rollup answers, its `aggregate_note` and the truncated-window disclosure are RENDERED rather than dropped - the page's range reaches 30 days, so without them those columns simply go blank and a blank column reads as "nothing to see" rather than "not measured at this resolution". Rows carrying no `query_hash` (history predating the column) are shown in the table and not offered in the picker, because a request the read answers with a 400 is worse than an absence. A new pin asserts the category rather than the instance: every REQUIRED parameter of every read the page fetches is one the page actually sends - the existing pin only asked whether the keys sent were bound, which a read fetched with NO keys passes vacuously. +- **The analysis family can be anchored at a past window too, and `analyze_server` refuses to persist when it is** ([#2506]) - [#2495] anchored 57 Darling reads and 50 Lite reads and deliberately left four out, and they are the four an incident investigation reaches for first: `analyze_server`, `get_analysis_facts`, `compare_analysis` and `get_analysis_findings`. Their window is not built in the tool; it is built inside the shared analysis engine, which the tools hand a bare `hours_back`. **`as_of` now reaches the engine** - same parameter name, semantics and description text as the other 107, from the same shared constant - so `analyze_server` and `get_analysis_facts` collect and score facts over the anchored window, `compare_analysis` hangs **both** of its windows off the anchor (`baseline_hours_back` has always been measured from the comparison window's end), and the anomaly detector's hour-of-day x day-of-week baseline moves with the window - which matters more than the window itself, because a run pinned to now compares Tuesday 03:00 against whatever baseline bucket today happens to be in. `get_analysis_findings` is the one whose window is on **analysis time**, so anchoring it asks what a scheduled pass was *saying* then rather than re-analyzing that window now; its read also gained an **upper** bound, without which an anchor could only ever move the window's start earlier and every anchored read would still have returned everything up to now. +- **`analyze_server` accepts the anchor and does not persist an anchored run** ([#2506]) - it is the one tool in the family that WRITES, and an anchored run persisted normally would be worse than no anchor at all. A finding row's identity, for every consumer there is, is its `analysis_time` - the moment the pass ran, not the window it looked at: the viewers' Recommendations tab reads `MAX(analysis_time)` and calls the result the server's current state, and `get_analysis_findings` filters on `analysis_time` and then collapses on `(story_path_hash, incident_id)` to produce occurrences / first_seen / last_seen / peak_severity. So a backdated pass stamped now would become "what is wrong with this server" for every human looking at the viewer and would inflate the very occurrence stats an operator uses to judge whether a live incident is getting worse - caused, invisibly, by somebody else's exploratory read. Recording the window on the row does not fix either: `time_range_start` / `time_range_end` are **already** persisted and already returned, and no consumer filters on them. So an anchored pass is exploratory by definition - its findings come back in full, and neither the row nor the completion event that feeds notification is emitted. The rule is **derived** in the engine (`AnalysisContext.PersistFindings` is false exactly when the window was anchored) rather than a flag a caller sets, because there is no legitimate caller for "anchored AND persist" and so there should be no way to express it. The result says which it was, in `persisted` / `persistence_note`. +- **Every windowed read can be anchored at a past incident, not just at now** ([#2495]) - every MCP read that takes `hours_back` measured it backwards from *now*, so no caller could ask what a server was doing during last Tuesday's incident. Widening `hours_back` until Tuesday falls inside is not the same question: for an aggregate read a wider window is a **different answer**, not the same answer with more rows - it changes what a top-N returns, what an average is taken over, and how much a capped read truncates. **`as_of`** is now an optional ISO-8601 UTC anchor on **57 Darling reads and 50 Lite reads** (identical name, semantics and description on both SKUs, from one shared constant), moving the window's END off now while `hours_back` stays its LENGTH - so "the four hours around 03:00 last Tuesday" is one call instead of a 170-hour window filtered by hand. It reaches the web surface too: `/api/read/{name}` is machine-derived from the MCP catalog, so each read's descriptor carries `as_of` and a test pins the two surfaces together in both directions. **Backward compatible by construction** - the anchor defaults to now, so a caller sending only `hours_back` gets byte-for-byte the window it always got, and `ValidateWindow` returns exactly what `ValidateHoursBack` returned when no anchor is sent. Following the `ValidateTop` convention, an anchor we cannot use is **refused** rather than quietly replaced: an unparseable value, or one in the future (past a five-minute client-clock allowance), comes back as a message, because a read that silently reverts to now is indistinguishable from a correct one. An anchor older than anything the store still holds is **not** refused - retention is per-deployment, per-server and per-collector, so a hardcoded floor would be a guess; it returns the read's own `empty` / `unavailable` status, which is unambiguous because the caller chose the window. Three reads that bounded only the LOWER edge (`get_cpu_utilization`, `get_tempdb_trend`, `get_alert_history`) now bound both, which for alert history is a correctness fix as well: its cap is applied by the database, so trimming after the read would have spent the whole `LIMIT` on rows newer than the anchor. Left unanchored on purpose, and stated in the instructions: the analysis family (`analyze_server`, `get_analysis_facts`, `compare_analysis`, `get_analysis_findings`) windows inside the analysis engine rather than at the read, and `get_pvs_stats` / `get_fleet_overview` mix a latest-snapshot measurement with a windowed one, where anchoring only the windowed half would return a result whose two halves describe different instants. +- **Five ready-made dashboards, so a new user's Custom Views page is not an empty one** ([#2480]) - no custom view ships seeded, so the first thing a new user meets at `#/views` is a blank page and a blank canvas over an 82-read catalog. The first-run hero already knew what would help and rendered it as three **inert** chips - "Top waits by server", "CPU trend over time", "Slowest procedures by database", no handler, no href - which is the [#2437] defect shape (a promise rendered as a caption) on the one page nobody arrives at with context. Those chips are now the real templates: **Server health at a glance**, **CPU investigation**, **Blocking and deadlocks**, **Memory pressure** and **Configuration review**, each created in one click against a server picked beside them, and they sit in the existing "New from template" menu alongside the notebook seeds ([#1563] D7) rather than in a second affordance. **Templates, not seeded rows**: seeding would need a migration rung, a `StorageVersion` bump, four pinned test files and a Viewer probe sentinel for content that is not schema - and a seeded row RESURRECTS itself on the next upgrade after the user deletes it, with no reset path if an editing seat breaks one. A template is created only when asked for, and what lands is an ordinary view the user owns, edits and can delete for good. The two halves of the menu behave differently on purpose: a notebook template links to the composer pre-filled because its composed panels re-scope live from the notebook's own controls, while a dashboard template is v1 READ panels whose `server` param is **static** (`renderView` threads variables and range into composed panels only), so it is created against the server chosen in the menu - and the server name goes into the view NAME, so two servers' copies of one template do not collide on the unique-name constraint. **Honest on day one**, which for a starter dashboard is the whole game - it is the first screen a UAT tester opens, on the store with the least data it will ever have. Every table and chart panel carries its own empty-state sentence (a count pin, not a spot-check), and the two reads that CANNOT be honest on a fresh install are deliberately absent rather than merely unused: analysis findings (the pass writes nothing for 24 hours) and Query Store (a target with it off has nothing, ever). The chart case is the one that bites without looking like it - `get_blocking_trend` and `get_deadlock_trend` answer an IDLE server with `trend: []` and no `{status,message}` envelope, so those panels say an empty trend means none happened rather than inheriting a sentence about collection. Every empty and failure state is stated rather than blank: a fleet with no servers says why the dashboards are unavailable instead of rendering an empty picker, and a refused create surfaces the backend's own message verbatim (with a 409 getting its own sentence, since a second click of the same template is the likely failure) because flattening it to "could not create" would hide the one line that says which panel is wrong. Pinned by the same invariant the built-in pages carry - every read exists in the shipped dispatch, every parameter key is one its read binds, every viz is in the shipped vocabulary - and each of the five definitions was fed to the service's own `ValidateDefinition`, the authority that would refuse the POST. +- **The web dashboard's server page gets the viewer's tabs: twelve sections, 61 of the 82 served reads, and a time range** ([#2475]) - the service dispatched **82** reads at `GET /api/read/{name}` and the built-in pages reached **23** of them; `#/server/{name}` was one scroll of seven panels against the desktop viewer's **65** `TabItem`s. The gap was never backend work - `panels.js` has been a generic renderer over those reads since #1562 - so this is descriptors over the unchanged `renderPanel` seam: twelve sub-tabs (Overview, Wait Stats, CPU, Memory, Blocking, File I/O, Queries, Configuration, Config Changes, Activity, System Events, Collection Health) carrying ~70 panels, reaching **61** reads. The header carries the WHY beneath the band badge - `Warning` has three unrelated causes (a real metric breach, a server awaiting its first collection, a collector error), so a badge reading "Warning" with no way to ask why is [#2422] rebuilt on a new surface; the fleet's own reason string and the fleet card's severity chips are rendered there, the chips through `fleet.js`'s own `metricBands` so there is one implementation rather than two, and neither is re-derived in the browser (R1). Every web grid that renders query text now puts it immediately right of its time/identity anchor ([#1949]) and the pin that enforced that on one array enforces it on all seven. The tab id rides in the hash (`#/server/{name}/{tab}`) so a section is deep-linkable and survives the 60s refresh, and an unknown or absent id resolves to Overview, which is what keeps every existing `#/server/{name}` link working. A page-level range picker (1h/4h/12h/24h/7d/30d) is the twin of the viewer's toolbar presets - **not** persisted, because a page that reopens on a 30-day window is slow for a reason the reader cannot see; panels whose read takes no window at all say "latest snapshot" rather than inheriting a label that would misdescribe them. Two reads were previously **unreachable from a browser** because they require a parameter no UI collected - `get_wait_trend` needs a `wait_type` and `get_perfmon_trend` a `counter_name` - and both now have a picker, the wait one seeded from the rows of the table directly above it so the picker and the table cannot disagree. **No fifth viz kind was added**: the property-grid shape that tempted one is served by `stat` (the reads returning a flat object) and `table` (the ones already returning rows), and a fifth kind that lived only in `panels.js` would be a page-only special case, while doing it properly means `KnownVizList`, `derive.js` and an editor config arm - composer surface this change does not need. What the browser genuinely cannot do is **stated in the tab where a reader goes looking for it** rather than left as a page that quietly lacks a feature: plan analysis, the query heatmap, cached-plan retrieval and actual-plan re-execution need a plan renderer and a command back to the monitored server, and the block-chain view and interactive deadlock graph need a graph viewer - the Blocking tab hands over the captured blocked-process-report and deadlock-graph XML verbatim instead of pretending. Every data panel supplies its own empty-state sentence and both helpers THROW without one. vizTable's generic "No rows in this window" reads as a fault on a collector that is off, opt-in, or daily; and the chart case was worse - `get_blocking_trend` and `get_deadlock_trend` answer an IDLE server with `trend: []` and no `{status,message}` envelope at all, so a perfectly healthy server was told its blocking chart did not have "enough data points to chart yet". `vizLine` now renders a descriptor's `emptyText` at exactly ZERO rows and still falls through at one (where the chart's own sentence is the true one) and when no `emptyText` was authored, so every stored view predating this is unchanged. A read feeding several panels on one tab is fetched ONCE (`fanout`) rather than per descriptor - `readTool` has no cache, so `get_collection_health`, which rolls up seven days of collector logs and computes sweep pressure, was running three times to open its own tab; six such duplicates existed across five tabs and a pin now refuses a seventh. Guarded by an invariant rather than by spot-checks: every read name the module mentions must exist in the shipped dispatch, every parameter key must be one its read actually binds (an unknown query key is silently ignored, so `limit` sent to a read binding `top` quietly returns the default), every viz must be in the shipped vocabulary, and no `get_pg_*` read may appear until the fleet payload can tell a PostgreSQL target from a SQL Server one - it carries `engine_edition`, not a `CollectorTargetEngine`, so a PostgreSQL panel today would render on all 42 SQL Servers, permanently empty. + +### Changed +- **The SQL-Agent collectors no longer gate on msdb access, so running the GRANT we advise actually does something** ([#2559]) - `HasMsdbAccess` is a **grant**, not an engine capability, but it was probed once at connect and cached for the connection's life. Three collectors gated dispatch on it, so the sequence a user actually follows was broken end to end: read our advice, run the `GRANT`, and nothing happens until the service restarts, with nothing to indicate why. `running_jobs`, `job_history` and `agent_status` now attempt regardless and fail into `PERMISSIONS`, which is a first-class outcome here rather than a defect - error **916 is already in `SqlServerPermissionErrors`**, so the run is classified as a permission denial and never as an ERROR, and `CollectorHealthClassifier` bands a collector that has only ever been denied as `NO_PERMISSIONS`, a check that runs **before** FAILING and STALE. So a server that will never have the grant does not read as broken, raises no alert, and the grant now takes effect on the next cycle instead of the next reconnect. The cost this trades for is three fast-failing statements per cycle on such a server - a compile-time permission check with no execution behind it. The fleet number that made the call: across both stores over three hours, **all 84 servers already dispatch all three collectors with zero non-SUCCESS runs**, so `HasMsdbAccess` never fires here and removing it changes nothing on this fleet while fixing the journey for anyone it does fire on. `HasMsdbAccess` is kept and still probed - it is honest information for a connection surface - it just no longer decides dispatch. One pin was **deleted rather than left passing**: the PostgreSQL-leak guard asserted no SQL Server gate reads `HasMsdbAccess`, which is now true for a reason that has nothing to do with the filter it was written to test, so it could no longer go red and was reading as coverage it did not provide. +- **`get_pg_top_queries` and `get_pg_blocking` return their int8 IDENTITIES as strings, because a JSON number was rounding them** ([#2548]) - PostgreSQL's `queryid` is a signed `int8` derived from a hash, so its values are spread over the whole 64-bit range and most of them sit past 2^53. Serialized as a JSON **number**, every parser that decodes numbers as IEEE-754 doubles - `JSON.parse`, `json.loads`, most agent tooling - silently rounded one: `-4185925123159566327` came back as `-4185925123159566300`, verified against real `JSON.parse`. The value was never wrong on the wire; it was unrecoverable after parsing, which for an identity is the same thing. **This is a breaking change to the response shape and it is the right trade**: `queryid` is the ONLY identity a PostgreSQL statement has, and every use of it - `SELECT ... FROM pg_stat_statements WHERE queryid = ...`, matching a row on our screen against one on the instance, quoting it in a ticket - is an equality join. A rounded metric is still approximately true; a rounded key matches nothing. **`get_pg_blocking`'s `root_backend_id` had the same defect in a worse form**, and it is fixed in the same change rather than noted for later. The collector builds that id by CONCATENATING the backend's start epoch with its zero-padded pid, so every value is a 17-digit integer around 1.79e16 - roughly 2x past 2^53, where adjacent doubles are 2 apart. Half of all backend ids are therefore odd and have no representation at all, and an unrepresentable one does not round to nothing: it rounds onto its even NEIGHBOUR, which is a different backend. Measured on four adjacent pids, two of the four collided; over 200, 100 lost precision and 99 landed on an id belonging to another backend. That is a worse failure than `queryid`'s - joining to the wrong row rather than to none - in the field the tool's own description tells a reader to prefer for comparing a root blocker across captures. The fix is scoped to those two: `database_id` and `user_id` are PostgreSQL `oid`s (unsigned 32-bit, structurally unable to reach the range), `root_pid` is an `int`, SQL Server's `query_hash` / `plan_hash` already reach this surface as `0x...` text, and Query Store's `query_id` / `plan_id` are sequential `bigint` IDENTITY columns whose real values are nowhere near 2^53 - so none of them has a defect to fix and none was touched. Web rendering needed no change: the Query ID column already rendered the raw value, so it simply started showing the true digits. Guarded by `PgInt64IdentityWireShapeTests`, which pins both fields as strings, pins `database_id` and `root_pid` as numbers so "stringify everything" cannot pass, and asserts its own fixture is genuinely out of double range - the assertion that stops the guard being vacuous. +- **Fifteen reads stopped answering "nothing happened" and "nothing was collected" with the same sentence** ([#2485]) - a read that returns no rows was saying so identically whether the window was genuinely quiet or the collector had never run, and those want opposite responses: widen the window, versus go find out why collection is not running. They now use the miss vocabulary honestly - `empty` for a true negative, `unavailable` for data that was never collected - with the second saying outright that it is NOT an all-clear and naming what to check. The design rule that came out of it is worth stating: **for an EDGE table the denominator is whether we LOOKED, not whether we ever FOUND anything**. Blocking and deadlocks only write a row when something went wrong, so probing the data reports a server that has been collected perfectly for months and simply never blocked as uncollected - a false alarm sending someone to fix collection that works. Those reads count successful collector runs instead. A parity pin now asserts each shared sentence appears in both source trees, because every one of these messages exists twice and nothing previously stopped one copy drifting. +- **`get_collection_health` serves what a HEAVY run costs, not only what runs cost on average** ([#2460]) - `query_store` on one dogfood server reported `avg_duration_ms: 13,834` over 1,155 runs. 958 of those runs carried the `enumeration yielded 0 items` note, and an empty enumeration costs **36 ms** - measured on a control server that yields nothing on all 1,551 of its runs and pays 36 ms for every one. Back that out and the remaining 197 PRODUCTIVE runs cost **~80,900 ms EACH**: more than the entire 60,000 ms sweep budget, on their own, once every few cycles. 13,834 ms describes neither population - it is an 83/17 blend that happens to land in a range reading like a plausible single number, and everything downstream inherited it, [#2459]'s brand-new `peak_cycle_ms` included, which understated that server's worst body by ~67,000 ms. Every collector row now carries `p95_duration_ms` and `max_duration_ms` beside `avg_duration_ms`, and `peak_cycle_ms` is built from the p95 rather than the mean. **The store already held this** - `duration_ms` has been written per run since the schema's second rung, and nothing had ever read it as anything but a mean - so this is two aggregates over a table already being grouped, no new collection and no new column. p95 rather than max for the number a decision is made from, because a max is one run and a single pathological cycle would make a collector read as permanently terrible for a week; the max is served BESIDE it as a fact, and comparing the two is what tells a routine tail from a one-off (avg ~ p95 ~ max is one population, avg << p95 is two, p95 << max is one bad run). p95 also scales itself to the sample: over 3,500 runs it discards the outlier, and over the six runs a daily collector gets in a week it lands on the max, which is right, because with six samples there is no outlier anyone can afford to throw away. Each collector is charged the p95 **floored at its mean**, so the aligned cycle can only ever rise - a p95 CAN sit below a mean (99 runs at 10 ms and one at 1,000,000 ms), and taking it unconditionally could have retracted a `BODY_OVERRUN` [#2446] correctly caught. **The verdict is untouched and still amortizes the mean**: sustained demand over a window IS the mean, and a rate built from a tail claims work the server never sustains. On the measured server the aligned body goes from 73,408 ms to 140,507 ms and `peak_collector` changes from `index_object_stats` (the larger MEAN, once a day) to `query_store` (the larger heavy RUN, every five minutes) - which is the collector that was actually overrunning those bodies. `peak_cycle_note` now states the mean beside the heavy run and the gap between them; `heaviest_collectors` rows carry both new statistics and take `pct_of_sweep_budget_per_run` from the heavy run, since a "per run" percentage computed from a mean that describes no run is the defect itself. Both SKUs, one query shape, and the same fixture pinned against live DuckDB and live Postgres so the two engines have to agree - including that both ignore a NULL `duration_ms` in an ordered-set aggregate, which neither query states. +- **`sweep_pressure` answers the single-sweep question as well as the sustained one** ([#2446]) - a server logging six "collection body has not completed after 60-69s of execution - skipping relaunch" warnings in three hours reported `busy_percent: 20.4`, `verdict: OK`, every collector HEALTHY - and both numbers were right about what they measured. [#2296]'s amortized model asks "does this server's total demand fit its cadence on average"; an operator reading a skipped relaunch is asking "did THIS sweep overrun". Those diverge exactly when one collector's single run approaches the budget while its amortized cost is negligible: `index_object_stats` took 37,207 ms of a 60,000 ms body and, at a 1440-minute cadence, contributed 26 ms/min to the verdict. The block now also carries `peak_cycle_ms` / `peak_cycle_percent` - what the body costs on the cycle where every scheduled cadence comes due together, the collectors' averages added WITHOUT being amortized - and `peak_cycle_risk` (FITS / BODY_OVERRUN). That cycle is not a hypothetical worst case: the shipped cadences are strictly nested (1 | 5 | 60 | 1440), so alignment is a guaranteed periodic event, and on that server it costs 73,408 ms against the 60,000 ms budget, which is what the pinned fixture reproduces. `peak_collector` names the collector that owns the most of one sweep and `peak_cycle_note` explains it, because `heaviest_collectors` ranks by amortized contribution and therefore ranks the offending collector out of sight by construction - that list now also carries `amortized_ms_per_minute` and `pct_of_sweep_budget_per_run` per row, so the two costs sit side by side. **The verdict is deliberately unchanged.** A once-daily 37-second collector is not saturation; calling it SATURATED would spend the word on a case whose lever is the schedule's shape rather than the capacity that verdict recommends, and an operator who learns to discount SATURATED loses the signal [#2296] built. Separate field and separate vocabulary, so neither can be read as the other and a fleet scan can filter on either. Measured across the dogfood fleet: the two servers that logged skipped relaunches read OK/BODY_OVERRUN (122% and 109% of budget) while a quiet one read OK/FITS at 19%. The decision stays in the shared SweepPressureClassifier (PerformanceMonitor.Common) with the same table pinned in both suites, and both SKUs' tools serve the identical shape. +- **Lite's portable ZIP is self-contained, which HALVED it** ([#2501]) - `Publish Lite` is now `-r win-x64 --self-contained` in both `build.yml` and `nightly.yml`, so neither Lite artifact has a .NET prerequisite any more and the failure [#2489] documented stops existing: a tester who unzips onto a stock Windows Server no longer meets the .NET host's bare `You must install .NET to run this application` before a line of our code runs. **The size went the opposite way from what bundling a runtime suggests.** The old publish was RID-agnostic, so it copied every platform its packages ship - **537 MB of `runtimes\` on a 565 MB tree** (osx 130, linux-x64 116, linux-arm64 70, win-arm64 56, then win-x86, musl, loongarch64 and riscv64), of which only the **52 MB `win-x64`** folder could ever load on Windows. `DuckDB.NET.Bindings.Full` is most of it, SkiaSharp and SqlClient behind it. Dropping ~485 MB of unloadable native payload beats the cost of bundling .NET, WPF and ASP.NET Core by roughly two to one: measured on one commit and one SDK, **565 MB tree / 212.7 MB zipped becomes 277 MB / 114.2 MB**. It matters most for the **nightly** ZIP, which is the UAT download and is not offered as a `Setup.exe` at all. **A RID-specific publish needed two more files than the flag.** `Lite/packages.lock.json` had only a `net10.0-windows7.0` target, and a RID restore adds `net10.0-windows7.0/win-x64` to it - after which the `dotnet restore --locked-mode` that BOTH workflows run before the publish fails `NU1004: the project's runtime identifiers have changed`, because locked mode compares the PROJECT's RID set (empty) against the lock file's (win-x64). Reproduced locally; that is a red CI run on every PR, not the future `--no-restore` trap it was filed as. The fix is `win-x64` in `PerformanceMonitorLite.csproj`, so the project itself asks for that graph and one committed lock file satisfies the RID-less locked-mode restore and the RID publish alike; `RuntimeIdentifiers` (plural) sets no RID on the build, so a plain `dotnet build` stays RID-agnostic and `Lite.Tests` is untouched. **SignPath needed nothing** - the `Lite` artifact-configuration slug already receives both shapes today, and the signed re-zip reads `signed/Lite/*`, inheriting whatever shape `publish/Lite` has. Auto-update is unaffected; the ZIP is not a Velopack channel. `LiteRuntimePrerequisiteDocsTests` went red on the flag alone (3 of its 7 facts) and was rewritten to state every claim BOTH ways round: [#2499]'s version asserted only that the docs DID name the runtimes, so two of its facts stayed green while the prose went stale. It now also derives the lock file's RID coverage from the `-r` flags in the workflows, and every new assertion was proven red with its fix reverted. + +### Fixed +- **Ten PostgreSQL MCP tools were never registered with the host, so no agent could call them** ([#2659]) - `get_pg_write_stats`, `get_pg_buffer_usage`, `get_pg_extensions`, `get_pg_lock_stats`, `get_pg_index_bloat`, `get_pg_column_stats`, `get_pg_kernel_stats`, `get_pg_predicate_stats`, `get_pg_replication_stats` and `get_pg_wait_sampling` were implemented, documented, dispatched by the web dashboard and counted in the instructions census, and `tools/list` answered 116 tools where the census claimed 126. Registration is per class and explicit, and nothing failed when a class was left out: the inventory pin checks tool NAMES, which exist either way, and the tab pin exists to stop a read shipping reachable only through MCP - this was the exact inverse. A reflection-derived pin now asserts every `[McpServerToolType]` class is registered, so it fails when someone adds a class rather than when an agent next reaches for the tool. +- **PostgreSQL 18 silently lost every I/O byte figure, and the estimate it replaced was off by an order of magnitude** ([#2655]) - 18 removed `op_bytes` from `pg_stat_io`, and both byte figures `get_pg_io_stats` serves were derived from it, so on 18 they came back null with no note, no status and nothing to distinguish a version change from a collector that had stopped. The columns 18 replaced it with are better than what was lost: `op_bytes` was the per-operation block size that the read multiplied by a count to ESTIMATE volume, while `read_bytes`/`write_bytes`/`extend_bytes` are measured totals. 18 also introduced vectored reads, so one entry in `reads` can cover several blocks and the old estimate undercounts - measured through the running service against a real 18.6 target, one combination reported 4,742 reads against 448,724,992 bytes where the estimate would have said 38,846,464, an 11.6x undercount, and three combinations ran 10x to 16x. They are now collected and served, with `bytes_source` on every row and on the envelope saying whether a figure was measured or estimated, because the two are not comparable and must never share a name silently. +- **A PostgreSQL column that the server's VERSION removed read as a missing measurement** ([#2653]) - PostgreSQL 17 removed `buffers_backend` and `buffers_backend_fsync` from `pg_stat_bgwriter` outright, and `get_pg_write_stats` returned both as bare nulls under a note that went on explaining `buffers_backend` as a live backpressure signal. The collector was right - it emits NULL for them deliberately from 17 on - but nothing anywhere recorded the target's PostgreSQL major, so no read could tell a column the version does not have from one nothing collected. Seven PostgreSQL collectors gate on that version and the read layer had no access to it at all. The registry now carries it, stamped on every connect like the engine kind next door, and this read spends it: on 17 and later it names the removal, says it is not a measurement gap, and points at `get_pg_io_stats`, where the fact actually lives now. +- **Self-hosted PostgreSQL had no query TEXT, so `test_hypothetical_index` could never work there** ([#2651]) - the statement-text store read `aurora_stat_statements()` with no vanilla path, so off Aurora `collect.pg_statement_text` was never populated. Two things failed silently: `get_pg_top_queries` returned `query_text: null` on every row forever, while that field's own documentation says null means "not captured YET" - true on Aurora, a lie here; and #2612's `test_hypothetical_index` resolves its statement from that table, so it always answered "no statement text is stored" and blamed a refresh cadence for a missing source. Fixing it exposed a second defect immediately: `pg_stat_statements` keys on `(queryid, userid, dbid, toplevel)`, so one queryid returns once per user and database, and the upsert - which keys on `(server_id, queryid)` - met those duplicates as `21000: ON CONFLICT DO UPDATE command cannot affect row a second time` and abandoned every batch. Both fixed and verified against a real self-hosted PostgreSQL: 47 statement texts stored where there were zero, and the hypothetical-index command then answered end to end for the first time - cost 1,059.34 to 438.71, a 58.6% reduction. +- **Azure SQL DB captured almost no deadlocks, and could not say so** ([#2641], reported from the field) - the only source was a DATABASE-scoped Extended Events session we create, read from a 4 MB ring buffer. A connection to `master` therefore captured only `master`'s deadlocks, and the reporter's fifty user databases were invisible; the buffer is also memory-resident, so an Azure failover empties it unannounced. Azure's own file-backed telemetry - `sys.fn_xe_telemetry_blob_target_read_file` - is now read alongside it, and it carries the USER database's name so one read from `master` covers the whole logical server. It is MASTER-SCOPED: called from a user database the identical statement returns zero rows with no error, which is why the arm is guarded on `DB_NAME()` rather than on the engine. Verified against a live Azure SQL Database by running the shipped query text: from `master` it returns the deadlock with the user database attributed, from that user database it returns the ring-buffer row unchanged. +- **Lite forgot the time range you picked** ([#2640], reported from the field) - the picker wrote nowhere, so "Last 7 days" came back as four hours after a restart. The settings key it should have written (`default_time_range_hours`) already existed and was already read at startup; only the Settings window ever wrote it. Picking a range now persists it, and opening a second server tab in the same session uses it too. Custom Range is deliberately not persisted - restoring a window that ended two days ago is worse than restoring nothing. +- **Lite's Database Sizes grid looked broken on Azure SQL DB** ([#2640]) - it is headed "All Servers" and on Azure shows one database's files, because the collector reads `sys.database_files` on the CONNECTED database by design (#1631 removed `master` from its path, since a database-level firewall rule makes `master` unopenable). Correct, and indistinguishable from a collector that only found `master` - which is what the reporter saw and reasonably filed. The grid now names the scope and says the sibling databases are unreachable rather than missing. +- **A missing-extension message said CREATE EXTENSION without naming which database** ([#2638]) - extensions are per-database, and on the fleet `pg_buffercache` was installed on the cluster in a DIFFERENT database from the one the collector connects to while the collector reported it missing. Both statements were true; an operator who checked the obvious database would have found it already there and concluded the collector was broken. The message now names the connected database, says outright that an extension installed elsewhere on the same cluster is invisible from there, and degrades to the old wording when the database is unknown rather than inventing a name. +- **`get_index_usage` hid whole databases behind a fixed 200-row cap and never said so** ([#2636], reported from the field) - rows sort unused-first across the WHOLE server and the cap was a hardcoded 200 with no way to scope or raise it, so on an instance with 200+ unused indexes in one legacy database that database consumed the entire answer and every Active index everywhere else was invisible. The reporter's database had healthy collection, full retention, zero errors and zero returned rows, which reads exactly like a broken collector. It now takes `database_name` and `limit`, and every response carries `matching_index_count`, `truncated` and a sentence saying which rows were dropped and why. A database filter that matches nothing while the server has rows elsewhere gets its own message naming the count, so a typo'd or excluded database is never reported as an uncollected one. The unused-first ordering STAYS - it is right for the question the tool exists for, and reversing it would only trade the defect for its mirror image. +- **RDS plan capture reported SUCCESS "no new plans" when the AWS call was DENIED** ([#2633]) - the ingestor caught every failure, warned, and returned zero rows, and the runner turned zero into a `collection_log` row asserting the log had been opened and held nothing. Measured on the PostgreSQL monitoring host: the row said SUCCESS while the app log said `rds:DescribeDBLogFiles` was denied by IAM - nothing had been read. A regression against the route it replaced, since the `pg_read_file` path answers the same situation with PERMISSIONS and names the grant. An authorization refusal now degrades to PERMISSIONS with a message saying the grant is on the MONITORING HOST's IAM role rather than the database login, and that nothing was read; every other failure stays loud, because a permanent-sounding status on a transient fault is how an outage gets read as a configuration choice. +- **`IsAwsRds` was never set on a PostgreSQL target, so plain RDS PostgreSQL could never capture plans** ([#2633]) - it is probed with a T-SQL detection query on the SQL Server path only, so `pg_plan_capture`'s `IsAurora || IsAwsRds` dispatch had an unreachable half and a managed non-Aurora instance fell to the `pg_read_file` route, where there is no filesystem to read and the failure names a database grant that would never have helped. Now derived from the endpoint. Aurora was unaffected because `IsAurora` carried the routing, which is exactly why the fleet could not show it. +- **A summary count taken over a capped result read as a fact about the server** ([#2629]) - caught in the new `get_pg_extensions` before it shipped: at a 50-row limit it reported `installed: 10` for a server with 13, because 50 rows was all it had looked at. State totals, the measured-index count and the ungranted-lock count are now WITHHELD when the result is truncated, with `truncated: true` and a sentence saying why, rather than renamed to something nobody wants to read. +- **`pg_wait_sampling` excluded only `Activity`, so `Client`/`ClientRead` was 100% of the profile** ([#2630]) - its Aurora sibling excludes `Activity`, `Client` and `Timeout`, with a rationale measured on production, and the sampler restated a shorter list. On the first target profiled with real client connections, `ClientRead` was 2,717,290 of 2,717,989 samples and every real event rounded to zero; the unfiltered profile held 17,864,575 samples of those three types against 2,150 of everything else. Both collectors now splice ONE definition of what does not count as a wait - they answer the same question from different sources, and #2625 tells operators to read one instead of the other. The predicate coalesces before it compares, so the CPU row - a backend that was NOT waiting, and this collector's distinctive signal - survives a filter that would otherwise silently discard it as NULL. +- **An unknown column `format:` rendered as raw text instead of failing** ([#2629]) - the web grid renderer falls through silently, so a column declaring a format that does not exist still renders, looks populated, and is simply wrong. Caught while writing one. Every format `server-tabs.js` declares is now pinned against the renderer's own vocabulary. +- **`--enable-web` on a non-Windows host reported the wrong reason and hid the path that works** ([#2626]) - the platform guard fired before the mode guard, so a macOS or Linux operator following the service's own advice was told `--enable-web requires Windows (DPAPI + firewall)`. True of the verb, misleading about the situation: both Windows dependencies are managed-store concerns, and a non-Windows deployment is necessarily bring-your-own, where the flags live in the operator's own PostgreSQL. The refusal now names that path and hands over the statement - `UPDATE config.config_service SET web_enabled = true;` - which the service applies within one sweep with no restart. A managed store, or a config that cannot be read, still gets the platform sentence, because then it is the honest one. +- **A per-database collector that failed in SOME databases reported SUCCESS with no note** ([#2623]) - both runners tolerate a per-database failure by design, and both escalate only when EVERY database failed. In between sat a hole with no evidence in it: the cycle recorded SUCCESS, whatever the surviving databases produced, and a note composed solely from enumeration probe failures - which a thrown exception is not. When the failing database is the only one holding data, that is SUCCESS with zero rows, which is exactly the shape of a target that genuinely has nothing to report. It is how the payload-arity bug above stayed alive: all three broken collectors run per database, all three failed in the one database that had data, and only the one that also fails in `postgres` tripped the all-failed escalation. A partial loss now composes a note naming the skipped databases, the count, and the first error, merged with any probe-failure note rather than replacing it. Skipping the database is still right; skipping it quietly is not. +- **Three PostgreSQL collectors wrote fewer payload values than they declared columns** ([#2599]) - `pg_extension_availability`, `pg_column_stats` and `pg_index_bloat` all gained a `database_name` COLUMN in v95 whose value was never written, so every row they produced was rejected at `EndPayload`. The runtime check existed and was not enough: it only fires when a collector RUNS, all three are daily, and on the Aurora fleet two of them return zero rows for unrelated reasons - so the store simply had nothing from them, which looks exactly like a quiet server. Found by pointing the service at a **self-hosted PostgreSQL with 40 tables and 160 indexes**, three schema versions after the bug shipped. Now pinned at build time across the whole catalog: a counting writer runs every collector's `WritePayload` and asserts the write count equals the declared column count, with a coverage assertion so collectors the probe cannot exercise can never quietly become most of the catalog. +- **`pg_index_bloat` had never returned a single row** ([#2617]) - found by dogfooding v99 on a live Aurora target, where the collector failed every cycle with `Exception while reading from stream` and `rows_ever = 0` for its entire life. **The wrong dimension was bounded.** `pgstatindex` reads every page of the index it is pointed at, and the 20 GB `MeasureCeilingBytes` bounds one INDEX while nothing bounded the STATEMENT: measured on that target, **1,517 indexes totalling 461 GB** in a single query on the default command timeout, which never finished and dropped the connection mid-read. A cycle work budget now measures the **200 largest** indexes per run - a count rather than a byte cap, because the cost is per page and index sizes are wildly uneven, so a byte cap would measure three indexes on one server and four hundred on another with no way for an operator to predict which. Everything past the budget is still **returned** carrying `skipped_reason`, never dropped: an index missing from the result reads as one that does not exist, while an index present with a stated reason cannot be mistaken for healthy. A 300-second `CommandTimeoutSecondsOverride` means a slow single index now yields a CLASSIFIED timeout instead of an unclassified stream failure. The local rig that verified #2561 had two indexes on one table, so the question of total work never arose there - which is the class of defect only a real instance shows. +- **The README documented an upgrade script that no released build contains** ([#2593]) - reported by somebody upgrading 3.3 to 3.5 who could not find the procedure written down anywhere, and they were right: `upgrade-darling.ps1` landed on 2026-08-22, three days AFTER 3.5.0 was tagged, so neither their install nor the 3.5.0 zip contains it - while the README on the default branch has been describing it as the supported path to anyone reading the repo. The upgrade section now says plainly that the script does not exist in 3.5.0 or earlier, and a new section gives the manual procedure it automates. That procedure also states the thing the reporter had to infer: **nothing needing preservation lives in the install directory** - the managed store, the DPAPI credential blobs and the logs are all under ProgramData, so only `darling.json` has to travel, and renaming the install folder is a BINARY rollback rather than a data one. Their own plan turned out to be better than ours on one axis and it is written down as such: extracting into a clean directory cannot accumulate files a newer build stopped shipping, which is exactly what `-RemoveStaleFiles` exists to clean off an overlay-upgraded tree (#2529). +- **Tag management had no discoverable entry point in the Darling Viewer, and none at all on a viewer with no servers** ([#2595]) - reported, and correct on both counts. The only door was the group-header context menu, which means right-clicking a row labelled **Untagged** - a place nobody looks for tag management. And that header only renders when at least one server is untagged (`FleetView` emits it under `untagged.Count > 0`), so on a freshly registered viewer with no servers there was no right-click target at all and tags were unreachable entirely. Lite has carried a visible **Manage Tags** button in the same sidebar footer all along, so this was a parity gap rather than a missing feature; the viewer now has the same button, next to Manage Servers. The reporter's second observation - that clicking "Assign Tags" does nothing - is WPF behaving normally: that item is a submenu parent, so it opens on hover rather than doing anything on click, and with no tags defined its submenu holds only "No tags yet" and a nested "Manage Tags..." that was the sole remaining door. `ManageTags_Click` already carried its own read-only backstop (#2008), so the new button needed no gating of its own. Pinned by a test that locates the handler INSIDE the sidebar footer and checks it is a `Button` - asserting mere presence would have passed against the broken build, because the handler was wired, just only inside a `ContextMenu`. The pin strips XML comments before scanning, because the new button carries a comment that names the handler and an unstripped scan finds the explanation instead of the wiring. +- **The plan-capture remedy recommended the one setting measured to be catastrophic** ([#2565]) - the `capture_threshold` facet told anyone sitting at `log_min_duration = -1` to "set it to 0 to capture every statement, or a millisecond threshold". Then #2565 measured what 0 costs: on PostgreSQL 17 under pgbench with 8 clients, capturing every statement cost **31 percent of throughput** and wrote **772 MB of server log in 20 seconds**, while the same instrumentation at a 10ms threshold cost nothing measurable and at 100ms measured slightly above baseline. On Aurora that log is bounded by `rds.log_retention_period` (3 days on this fleet) and read back through the RDS API, so at that rate plans age out before anything can collect them. The remedy now says **set a millisecond threshold, not 0**, and explains why in one sentence rather than leaving it as taste. A server already sitting at 0 gets its own arm rather than the generic everything-is-fine sentence - it is still `is_satisfied = true`, because capture genuinely is working, but the detail names the cost and says to move off it. Both arms now also note that the threshold is a **dynamic** parameter on Aurora/RDS and applies without a reboot, unlike the library itself - measured across all 51 cluster parameter groups, where all 14 `auto_explain.*` parameters are dynamic. One more thing came out of checking how to compare against 0 safely: `current_setting` renders this GUC in the largest unit that divides evenly, so -1 and 0 come back bare, 250 comes back `250ms`, **1000 comes back `1s`** and 60000 comes back `1min`. Anything downstream that read this as a number would break on a perfectly ordinary one-second threshold, which is why `observed` is stored as the server's own text and never cast. Verified against a live server in all four states. +- **`extension_available` said no on every PostgreSQL server in existence, including ones actively running auto_explain** ([#2564]) - shipped hours earlier in the same release and wrong from the first row. The facet asked `pg_available_extensions`, and **`auto_explain` is a preload-only module**: no `CREATE EXTENSION`, no control file on disk, so it never appears there. Caught by running the shipped query against a container with the module loaded and serving a 250ms threshold, where the collector reported `library_loaded = true` and `extension_available = false` two rows apart, and told the reader that plan capture was "a platform limitation rather than a configuration step". That is exactly the send-half-the-readers-to-the-wrong-fix failure the collector was written to prevent, aimed at everybody rather than at half. It now reports availability it can **prove** - catalogued, or already loaded, since a loaded library is proof of its own availability - and a negative says plainly that it is inconclusive, names the reason, and points at the cluster parameter group's allowed values for `shared_preload_libraries`, which is where the real answer lives on Aurora/RDS and which SQL cannot see. The catalog check is kept rather than dropped, because a managed provider shipping a control file would still be a true positive. +- **The Viewer's connection self-test probed a different connection string than startup opens** ([#2578]) - `MainWindow` builds its data service as `new ViewerDataService(settings.ConnectionString, appSettings.ConnectionTimeoutSeconds)`, and that constructor rewrites the string through `ApplyConnectionTimeout`. The headless self-test used the RAW darling.json string, so the two diverged whenever a Connection timeout preference existed - which is always, since the preference defaults to 5 and is clamped to 5..60. A diagnostic that validates a different string than the application opens cannot do the job it exists for: it can report `[PASS]` on config, tcp, tls, auth and schema while startup times out against the same store, which is exactly the field report that found it, and it makes that outcome look like a contradiction rather than the expected consequence it is. The probe now reads the preference the same way `MainWindow` does, applies the same transform, and prints a note naming the preference when it changed the string - so the report says which connection string was actually tested. Unreadable viewer settings do not stop the probe; it falls back to the raw string and says so, because a tested-something result beats a refused one as long as the reader knows which. +- **A gated-off collector logged a fake SUCCESS on SQL Server targets** ([#2579]) - the dispatch loop drops wrong-DIALECT collectors before they run, and does the same for right-dialect collectors whose `AppliesTo` gate excludes the target - but that second skip was scoped to PostgreSQL. On a SQL Server target a gated-off collector still dispatched, the runner's own check returned `CollectorRunResult(0,0,0)`, and that was recorded as SUCCESS: `duration_ms=0`, `sql_duration_ms=0`, `rows_collected=0`, no error, no note. The scoping was deliberate and documented, left as "its own decision" because it changes a shipping SKU's log semantics for a handful of Azure-gated collectors. On an AWS RDS fleet it is not a handful - 84 instances x `agent_status` and `running_jobs` x a 5-minute cadence is **~24,000 rows a day** claiming success for collectors deliberately not running. The reason this is worse than noise: such a row is byte-identical to a real one, so nothing downstream can tell them apart, which is exactly the shape the miss vocabulary exists to prevent everywhere else - and it read convincingly enough as working collection to produce a filed issue and an opened PR built on it before the 0ms durations gave it away. Both were withdrawn. The skip now applies on every engine at **all three** dispatch loops - the scheduled sweep, the on-load loop that runs on every connect and reconnect, and `snapshot_now` - since fixing only the sweep would have left the other two landing fake successes at exactly the moments an operator is watching. The wrong-dialect drop above each is untouched, and `--test-connection` still names which collectors do not apply to a target and why, so the silence is not silent. +- **Rollback backups kept an unhardened copy of `darling.json`** ([#2574]) - the upgrade script copies the install root's files aside as `_rollback_manual_` before laying down a new build, and `Copy-Item` into a new directory takes the DESTINATION's inherited DACL rather than the source file's. So the copy landed with whatever the install root grants, measured on a real box as `BUILTIN\Users : ReadAndExecute` - the inherited-from-`C:\` DACL #1647 called out. That matters because of what #1647 established: `darling.json` holds every monitored server's `encryptedPassword` plus the MCP and web dashboard tokens, all DPAPI **LocalMachine** scope with an entropy constant published in this open-source repo, so anything that can READ a copy can decrypt the lot. The live file was hardened; up to three retained copies of it beside were not, on the same box. The backup now restricts the `darling.json` and `*.dpapi` files it takes to SYSTEM and Administrators, inheritance disabled so the root's ACE cannot flow back in - and deliberately WITHOUT the Interactive grant the live file carries, because the Viewer and the CLI verbs read the live config while nothing reads a backup except a human recovering, who is an administrator by then (the same line `install-darling.ps1` already draws for `.bak-*`). Best-effort and non-fatal: a backup taken but not hardened beats an upgrade that refuses to proceed, and the operator is told which. **Retention was not the problem** - the script already keeps a bounded number (`-KeepRollbacks`, default 3) and prunes the rest. A fleet sweep found 34 backups across three boxes, 18 of them with a `darling.json` readable by `BUILTIN\Users`, but every one predates `upgrade-darling.ps1` itself, which first shipped two days earlier along with that pruning - historical residue from however those boxes were upgraded before the script existed, now removed. The script is the only documented upgrade path, so this fix covers every future one. +- **The upgrade-path gate was climbing from a store nobody has** ([#2572]) - `MigrationUpgradeLadderLiveTests` is the #2119 gate added after a rung replayed on a 3.3.0 store referenced `query_plan_gz` three rungs early and killed every 3.3.0-to-3.4.0 upgrade with 42703 at service start. It builds a store from the previous release's own frozen ladder and runs the current one over it. The generator's instructions say to regenerate that fixture **at each release cut** - and both the 3.4.0 and 3.5.0 cuts shipped without it, so with 3.6.0 about to ship seven new rungs the only tested upgrade population was still a **3.3.0** store (top rung V39) while every real user upgrades from **3.5.0** (V79). Not simply less coverage: the stale fixture replays MORE rungs, 47 against 7, which for the specific defect #2119 exists to catch is broader. What went untested is the POPULATION - a rung that assumes the shape a 3.5.0 store actually has can pass against the 3.3.0-derived fixture and fail in the field. The v3.5.0 fixture is regenerated (78 rungs, top V79), and the old one is **kept rather than retired**, against the tool's own instruction: retiring at each cut shrinks the replay surface every release until it covers nothing but the current cycle's own rungs, which is the opposite of what #2119 wanted. The climb is now a Theory over every fixture in the tree, so a release cut adds one without editing the test. And because this decayed silently for two releases - the contract lived in a csproj comment and a generated file header, both prose, neither executable - there is now a guard on the guard: a test derives the newest released version from CHANGELOG.md's own headings and fails until a fixture exists beside it, printing the exact command to generate one. Derived rather than pinned to a constant, because a constant would need remembering at precisely the moment the fixture needed remembering. It runs without a live PostgreSQL, unlike the climb it protects, since a gate is a poor place for the check that says the gated thing is pointed at the right store. +- **`get_running_jobs` still claimed a server had no jobs running when we were never allowed to look** ([#2559]) - #2546 built the `precondition` vocabulary precisely to stop that affirmative claim, and #2557 wired it into this read. It did not cover the case #2559 is about, because that case produces no evidence to read: `DarlingRuntimePrecondition.StatusAsync` reports what the collector's last run RECORDED, and a collector whose `AppliesTo` gate is off never runs - the runner returns before writing any `collection_log` row, deliberately, since a per-cycle fake row was thousands of rows a day of noise. So a login with no msdb access at all (`HAS_DBACCESS('msdb') = 0`) left nothing behind, the precondition reader found nothing, and the read fell through to *"No running SQL Agent jobs found"* - an assertion about the server's Agent, on a server nobody was permitted to query. What #2557 DID fix is the adjacent case: a login that can enter msdb but is refused SELECT on the job tables, where the collector dispatches, is denied, and records that denial. The two are indistinguishable to a user, which is why this read as fixed. The gap is now closed **without a schema change and without touching dispatch**: a collector with no `collection_log` row at ALL, on a server that is demonstrably collecting other things, is gated off by construction - there is no third way to get there, and retention cannot fake it, because a collector whose rows have aged out has not run inside the window either. The message names the candidate gates rather than guessing (neither `HAS_DBACCESS` nor the RDS flag is persisted on the registry) and, unlike the ordinary precondition, tells the operator a **reconnect** is required. **That also fixes a contradiction that has shipped since #2546**: every precondition appended a shared epilogue promising "nothing to restart on the monitoring side", while the `SESSION_MISSING` arm's own sentence said a dropped session "stays missing until the next connect" - both claims in the same operator-facing string. The epilogue is now made only by the arms that can honour it; the connect-scoped arms say what they actually require. The MCP instructions carried the same blanket promise to every connected agent and now say to read the message instead, because telling somebody to retry a connect-scoped precondition sends them round a loop that never terminates. The dispatch half of #2559 - whether to drop the gate and pay a doomed query every 5 minutes forever - is deliberately untouched and stays open. Landed on BOTH SKUs: the gate lives in the shared collector definition, so a Lite login with no msdb access reproduced the identical claim, and the parity pin now holds the corrected instruction text plus a per-SKU wiring assertion so one tree cannot grow the arm without the other. +- **A wildcard `listen` refused every LAN request with a 400 — which is what the compose distribution ships** ([#2569]) - the DNS-rebinding Host guard compares the request's `Host` header against the configured listen IP, and with `"listen": "0.0.0.0"` the only Host that can ever match is the literal string `0.0.0.0`, which no client sends. So a wildcard-bound web dashboard or MCP endpoint answered `localhost` and refused everything else - before any auth ran, since the guard is the first middleware in both bind modes. `Darling/compose/darling.sample.json` sets exactly that for **both** surfaces and the quickstart then tells you to browse to `http://:5153`, so a deployment following the documented steps got a dashboard that 400s. It has plausibly gone unnoticed because a specific listen IP - what a Windows install naturally uses - is the configuration that works. Measured through a real Kestrel pipeline rather than inferred: `Host: 192.168.1.205` -> **400**, `Host: localhost` -> 200; after the fix, 200 and 200. **A wildcard listen now accepts any IP literal**, because a wildcard names no single address to compare against. The obvious alternative - enumerate this machine's interfaces at bind time - is worse than it looks and would NOT have fixed the case that motivated this: inside a container the local interfaces are the container's (`172.18.0.3`), while the address the browser used is the host's published one, which the container cannot see; NAT, port forwarding and multi-homing break it the same way, and silently, as a 400. **The rebinding defence is unchanged**, which is what makes this a fix rather than a loosening: a rebind arrives with a HOSTNAME by construction - the victim's browser loaded `http://evil.com`, so it sends that Host and treats the reply as same-origin - and a hostname still never passes, nor does an IP-shaped prefix on an attacker domain (`192.168.1.205.evil.com`). An admitted IP literal still has to clear the surface's CIDR check and its token. A SPECIFIC listen IP is untouched and still rejects every other literal, pinned so the wildcard arm cannot leak into it. Fixed once in `HostHeaderGuard`, so Darling's web dashboard, Darling's MCP host and Lite's MCP host all get it. +- **Editing a REGISTERED server in `darling.json` was silently ignored, and the warning about the adjacent case made that worse** ([#2552]) - `WarnAboutFileOnlyServersAsync` compared server_ids and names and nothing else, then returned early once every file server was registered. So for a server the store already had, **every per-server setting in the file was dead text** - `host`, `database`, `auth`, `username`, `encryptMode`, `trustServerCertificate`, `excludedDatabases`, `port`, `engine`, the display name - read, parsed, validated, logged as "Loaded configuration ... 1 server(s)", and then not used. The field report is the loop that makes it expensive: a PostgreSQL target refused a self-signed certificate, the operator applied the documented fix (`"trustServerCertificate": true`), restarted, and got a **byte-identical error** - the store row still said `f`, with `created_at == modified_at`. A connection failure is exactly the class of problem an operator fixes by editing config and restarting, and this was the one class of edit that produced an unchanged error with no explanation. [#2252]/[#2254] already warned - well - about the ADJACENT case (a server ADDED to the file and never registered), which teaches "adding a server to the file does not register it" and invites precisely the wrong inference: that the file still drives the servers the store already knows about. The service now names, once per start, every registered server whose file entry disagrees with its registry row, every field it disagrees about, and both values. **The store still wins** - store-authoritative is the deliberate design and none of it changed; the defect was that the disagreement was invisible. **The comparison runs through the same folds the connect path and the collectors apply**, because a check that warns about differences that do not exist is the same silence with extra steps: `encryptMode` through the connection builder's own fail-closed fold (so `strict` vs `Strict` is not drift, and neither is a typo against `Mandatory`), `engine` through `TargetEngine` (so `aurora` and `postgres` are one engine), a blank `database` through the engine's implicit default (`master` / `postgres`), PostgreSQL port `0` as its `5432`, `excludedDatabases` as the `NOT IN` SET the collectors splice (order, repetition, case and whitespace change nothing), and `monthlyCostUsd` numerically so `1200` and `1200.00` are one figure. Fields that are inert on the target are not compared at all - `readOnlyIntent` / `multiSubnetFailover` are SqlClient concepts the Npgsql builder never sees, a PostgreSQL port is not a SQL Server concept, and a username is only in the connection string under SQL auth - and the gate reads the STORE's engine, because the store is what the service connects with. **No credential is compared or printed, by construction rather than by intention**: `encrypted_password` is not in the SELECT list, so there is no blob in memory for a later edit to leak - and it must not be compared anyway, since a file entry legitimately carries an `env:`/`file:` reference or a dev plaintext password while the store row carries a DPAPI blob, which is the supported shape the read-time backfill exists to serve, not a disagreement. The remedy sentence names the Viewer's **Manage Servers** window and explicitly rules OUT `add_servers`, which cannot do this - it skips an already-monitored server as a `duplicate` before it validates anything, so sending an operator there would send them somewhere that silently does nothing. The `--test-connection` caveat is appended only when a CONNECTION-relevant field drifted: that verb probes the FILE's settings, so it can report PASS for a connection the service will never make - but it does not exercise the display name, the exclusion list, the cost figure or the delivery override, and a warning must not claim more than it can support. `darling.sample.json`'s `servers` block now carries the seed-once paragraph the `web.network` / `mcp.network` blocks have always had, and the list of servers is far more likely to be edited than either. **A control-plane-DISABLED server gets its own line, at Information** (raised in review): it is still registered, so it still pairs and its drift is still real, but "the registry is what the service uses" is not a true sentence about a server nothing is connecting to - so that line drops the claim, the remedy and the `--test-connection` caveat, and says only what is true of a paused server. Filtering the store read to `is_enabled = TRUE` was the obvious fix and is the one option that must NOT be taken: that read also answers "is this file entry REGISTERED", which a pause does not change, so filtering there would report a paused server as never-monitored and advise re-adding it - the [#2158] defect. Dropping it silently was rejected more narrowly, since [#2552] is a defect about silence being expensive and the drift is exactly what an operator walks back into on re-enable. **The deliverable is not the field but the invariant**: the covered set is DERIVED from both ends - `MonitoredServer`'s `[JsonPropertyName]` keys on one side, and the shipped comparison's own output when driven with two entries differing in every field on the other - so a per-server setting added to darling.json later is compared here or turns a test red naming itself. It cannot go back to being silently dead text, which is the category this belongs to rather than the one field that was reported. **Four folds are engine-ASYMMETRIC**, all four found by review asking what the message CLAIMS rather than what the code does: `encryptMode` really is three behaviours on SQL Server (three `SqlConnectionEncryptOption` values) but the Npgsql builder branches on `OPTIONAL` alone, so `Strict` and `Mandatory` are ONE connection on a PostgreSQL target and reporting them would report a difference the connection cannot express - the comparison collapses them there while the message still prints what each side actually says, so it can never show a value neither holds; `trustServerCertificate` is inert on a PostgreSQL target whose stored mode is `Optional`, because that branch picks `SslMode.Prefer` without ever reading the flag; and `database` and `excludedDatabases` are compared case-SENSITIVELY on PostgreSQL and case-insensitively on SQL Server, because PostgreSQL matches `pg_database.datname` byte for byte - `ReportingDB` and `reportingdb` are two databases there, and the exclusion list is live on both engines (`PostgresTargetProvider.BuildDatabaseListPlan` splices the same `NOT IN` filter to choose the per-database fan-out). The first two would have been false POSITIVES; the last two false NEGATIVES, which is the silent direction this whole entry is about. The exclusion set's inner comparer is the one that nearly shipped un-guarded: the outer comparison catches a case difference either way, so what only the comparer decides is whether a store excluding **both** `Scratch` and `scratch` is RENDERED as excluding one - a test that passes with and without the fix is worse than none, so it was rebuilt until it bit. Darling only - Lite has no control-plane store to disagree with. +- **`get_pg_top_queries` had never returned a row on any engine: its SQL did not parse** ([#2554]) - [#2219] added `LEFT JOIN collect.pg_statement_text AS t` to make the read carry statement text. That put `t.queryid` in scope beside `differenced.queryid`, which made the unqualified `queryid` references ambiguous, and PostgreSQL answers that with **42702 at PARSE time** - before a single row is examined. So the read threw on every call from #2219 onward, on **Aurora** as much as anywhere, and a store with zero rows failed in exactly the same way a full one did. Found by driving all nine `get_pg_*` reads against a live PostgreSQL 16 target; eight answered, this one returned `42702: column reference "queryid" is ambiguous`. **The obvious one-word fix does not work**, which is worth recording: qualifying only the `GROUP BY` still fails, on the select-list `queryid`, which is the reference PostgreSQL names first - measured before fixing, not assumed. Both are now `differenced.queryid`. **Nothing caught it because nothing executed it.** Roughly a dozen tests assert things about this query's TEXT and every one of them passed throughout; a substring assertion cannot resolve a name, only a server can. So the deliverable is not the qualifier but `DarlingPgReadSqlParsesLiveTests`, which `PREPARE`s **every** shipped PostgreSQL read against a real PostgreSQL. `PREPARE` runs the whole front end - name resolution, ambiguity detection, type inference - and stops before execution, so it needs no fixture and no seeded rows, which is precisely why it catches a defect that zero rows was never a defence against. The reads are discovered by reflection rather than listed, so a new one is covered the day it lands, and the discovered COUNT is asserted too - a filter that quietly stopped matching would otherwise turn the guard into a test that passes by finding no work to do. All twelve shipped read constants across the nine readers now parse. +- **A read that throws now gets the same honest capability answer as one that comes back empty** ([#2554]) - `DarlingEngineCapability`'s gate sat inside `if (rows.Count == 0)`, so it could only speak when the query SUCCEEDED and returned nothing. On a stock-PostgreSQL target - where `pg_statement_stats` can never run at all, its `AppliesTo` being Aurora-only - the parse error above surfaced as a raw SQL error where the sibling `get_pg_wait_stats` gives "does not run on that engine ... and never will". The gate is now consulted on the **throw** path as well. Deliberately NOT moved ahead of the read, which is what it looks like it should be: that helper's contract is explicit that every call site asks AFTER its read came back empty, so that a server whose registry row says one engine while its collected rows say another - a re-registration, a restored database - still gets its DATA rather than a confident explanation of why it cannot have any. Asking first would trade this defect for that one across every read; on the throw path there is no data to prefer, so the same reasoning points the other way. It also stays narrow: `NotCollectedStatusAsync` returns null unless the collector provably cannot run on this server's engine, so a genuine fault against an **Aurora** target still reads as an error rather than being dressed up as a capability gap - which is the assertion that keeps the fix from being an exception swallower, and is pinned as such. The shape is repo-wide (all nine PostgreSQL tools have exactly one `catch` and only this one consults the gate there), but with the parse defect fixed no other read can currently throw, so the other eight are recorded rather than changed. +- **On a PostgreSQL target, 55 more reads stopped telling you to go and check a collector that will never run** ([#2532]) - [#2511] wired the engine-capability helper into the reads whose collectors are gated off on Azure SQL Database, which was the right scope for the edition axis: a collector that runs on every SQL Server had nothing to say. The engine-KIND axis [#2530] added changes that arithmetic completely - on a PostgreSQL target **every** SQL Server read is a permanent gap - so `get_database_config`, `get_wait_stats`, `get_top_queries_by_cpu` and the rest still answered "unavailable - the config collector may not have run yet" about a collector that does not run on that engine at all. True, and pointing at the wrong cause. A human will get PostgreSQL tabs when the [#2530] UI waves land; an **agent** over MCP has no tabs and asks by name, which is half the product's surface. **72 Darling reads and 71 Lite reads** now ask on the miss path (up from 17 on each), covering every server-scoped read served by an identifiable collector. Deliberately NOT wired, with reasons: the rollups that read a dozen collectors and can name none of them (`analyze_server`, `get_analysis_facts`, `compare_analysis`, `get_daily_summary`, `get_daily_summary_range`, `get_server_summary`), the reads served by `collection_log`, which a PostgreSQL target genuinely populates (`get_collection_health`, `get_collection_log`), and everything that resolves no server at all (`list_servers`, `get_fleet_overview`, `get_store_metrics`, the alert and custom-view tools). `CapturePathByCollector` gains 28 noun phrases so the message names a surface - "the sys.dm_os_wait_stats cumulative totals" - rather than "the data this read is served from". Every call sits AFTER the read came back empty, so a server whose registry row disagrees with its collected rows still gets its data. +- **Eight PostgreSQL reads told a SQL Server target its vacuum was healthy, and stock PostgreSQL was told to go and check a collector that will never run** ([#2532]) - [#2530] taught the miss vocabulary the engine-KIND axis but wired only the *engine half* of the dispatch gate into it, so the PostgreSQL collectors' own `AppliesTo` gates were still invisible to it, and none of the eight `get_pg_*` reads asked the question at all. A SQL Server target asked `get_pg_xmin_horizon` got `no_holder` - "nothing is holding back the xmin horizon", a confident all-clear about a mechanism that engine does not have - and `get_pg_autovacuum_health` reported `no_pending_maintenance`, which is a fabricated clean bill of health rather than a weak one. `CollectorEngineCapability.TargetsWithEngineKind` now sweeps every target shape an engine kind permits, fixing `IsAurora` (nothing turns a stock PostgreSQL server into an Aurora one) and varying `PostgresMajorVersion`, `PostgresVersionNum` and `IsInRecovery`, all three of which an upgrade or a writer connection moves. So `pg_wait_stats` and `pg_statement_stats`, which read `aurora_stat_system_waits()` and `aurora_stat_statements()`, are permanent gaps on stock PostgreSQL and say so - while `pg_stat_io`'s PG16 floor and the writer-only autovacuum read keep the `unavailable` vocabulary that correctly sends an operator to look. The message names the real reason: a foreign DIALECT is stopped by the dispatch gate's engine half, a same-dialect gap is the collector's own gate, and one sentence for both would have told a PostgreSQL operator that a PostgreSQL collector is not written for PostgreSQL. Shipped with the guard [#2532] made its prerequisite: `EveryFactAPostgresGateReads_IsVariedBySweepOrFixedByKind` decodes the IL of every PostgreSQL `AppliesTo` and fails the build if the sweep leaves a fact it reads at its CLR default - the twin of [#2518]'s edition-axis guard, and the reason a sweep can be trusted not to over-claim silently. +- **The Overview lanes skewed apart when one lane's numbers got long** ([#2533]) - the five lanes are five SEPARATE plots, and `SyncXAxes` only ever gave them the same X-axis LIMITS. Nothing gave them the same pixel geometry: each plot sizes its own left gutter from its own Y tick labels, so on a server whose wait time runs into six figures the Wait ms/sec lane began its data area ~27 px right of the I/O latency lane reading `0.4`, and the lanes stopped lining up on the time axis. Every value stayed correct throughout - which is why the tooltips still read true, and what identified this as layout rather than data. **Not a regression**, and the reporter's "I think this worked before" is right for his machine: `MinimumSize`, `Layout.Fixed` and `PixelPadding` have never appeared in any SKU on any branch, so nothing was removed - alignment was always incidental, holding only while the five lanes happened to need labels of similar width, and five- or six-digit wait ms/sec is what breaks the coincidence. The same build therefore looks right on one server and skewed on another. The lanes now share ONE gutter, sized to the widest lane's own natural gutter and re-derived on every refresh - a frozen constant would have to cover six-digit wait times and would then steal that width from every server whose lanes read in single digits, and it ratchets back down when the big values go away. The mechanism is `Axes.Left.MinimumSize` rather than a fixed layout, because a fixed layout pins all four sides and the bottom lane legitimately needs more bottom room than the other four - it alone draws the time labels. This is the second alignment defect in these lanes and the first one shipped with nothing that could have caught it, so the fix comes with headless render tests: ScottPlot draws without a window, so CI now builds lanes with deliberately mismatched magnitudes, renders them, and compares their data rectangles - shown failing at 20.02 px on the unfixed control before it passed. **Darling viewer only.** Lite's Overview control is the same five separate plots and carries the same latent defect, but nobody has reported hitting it there; the aligner lives in the shared `PerformanceMonitor.Ui`, so picking Lite up later is one call site. Reported by @ghauan. +- **Lite's Overview lanes carried the same latent gutter misalignment as [#2533], now fixed there too** ([#2535]) - Lite's `CorrelatedTimelineLanesControl` is a third copy of the same five separate `WpfPlot` lanes, with the same axis-limits-only `SyncXAxes` and the same missing pixel-geometry step; confirmed present by reading the shipped source rather than assumed from the shared XAML. `SyncXAxes` now calls the shared `PerformanceMonitor.Ui.LaneAxisAligner.AlignLeftGutters` between setting X limits and refreshing the lanes, identical to the [#2533] fix - no second implementation, since the aligner lives in the shared `PerformanceMonitor.Ui` project Lite already references. A source-parse wiring pin (`Lite.Tests/LaneAxisAlignerWiringTests.cs`) asserts the control calls the aligner inside `SyncXAxes` and before the refresh loop, bounded by brace matching so Lite's `AddGhostLine` helper - which sits a few members later and also calls `Refresh()` - cannot satisfy it by accident; confirmed red (via a throwaway harness running the same parse logic) against the pre-fix source before the call site was added. The behavioral coverage - the gutter is independent of plot width, non-decreasing in plot height - is not duplicated here, since those are properties of the shared helper already pinned by `Darling.Tests/LaneAxisAlignerTests.cs`. Not run against a live Lite build: WPF targets `net10.0-windows` and only runs in CI on Windows. +- **A SQL Server read aimed at a PostgreSQL target says so, instead of telling you to check collection** ([#2530]) - [#2511] taught the miss vocabulary the engine-EDITION axis, so a read served by a collector gated off on Azure SQL Database answers `not_collected` naming the engine. A PostgreSQL target got `unavailable` - "never collected, check that collection is running" - which is true and points at the wrong cause: those collectors do not apply to that engine at all. That was the honest answer while nothing in the store distinguished a PostgreSQL target from an unconnected one; now something does. The answer is derived the same way the edition axis is, from the collectors' own dispatch gate rather than from a list, so flipping a definition's target engine moves it. **Unknown targets keep making no claim** - a pre-v82 row, a server that has not connected since the rung landed, and a token a newer build wrote all stay silent, because the distinction added is "known to be PostgreSQL" and never "not known to be SQL Server". It runs both ways: a `get_pg_*` read aimed at a known SQL Server target now says the same thing in reverse. Limit worth knowing: only the reads [#2511] wired to the capability helper can answer on either axis, so a read like `get_database_config` - whose collector runs on every SQL Server, so the edition axis gave it nothing to say - still reports `unavailable` on a PostgreSQL target. +- **A fresh store never seeded, so the control plane silently overrode `darling.json`** ([#2524]) - the seed INSERT named sixteen columns and supplied fifteen: a bound parameter was never written into the VALUES list, so every value after it shifted one column left and PostgreSQL refused the whole statement with `42601`. The failure was caught and logged and the service carried on with the file config, which is why it shipped - the damage came next, because the control plane is authoritative after first contact and was answering from a row that had never been written, where every toggle is false. So `web.enabled` and `mcp.enabled` were **true** in `darling.json`, **false** in the store, the store won as designed, and neither endpoint ever listened. That was the entire first-run experience of a container deployment, and the only clue was one line saying the config would seed on a later start - it would not, because the statement was malformed rather than racing. Found by running the shipped nightly image against a live target rather than by reading it, since no test drove the seed against an empty store. +- **The `tempdb Space` alert measured distance to the next autogrow, not distance to the ceiling** ([#2515]) - `TempDbSpaceInfo.UsedPercent` was `total_reserved / (total_reserved + unallocated)`, and both halves come from `tempdb.sys.dm_db_file_space_usage`, which reports the data files **as currently allocated**. On a pre-sized on-prem box that reads like real headroom only because such a tempdb has already grown to its cap. On **Azure SQL Database** it is not: the platform creates tempdb small and grows it toward the tier limit, so on `GP_S_Gen5_2` - four data files of 16 MB, **62.44 MB allocated**, each with `max_size` 2,097,152 pages for a **65,536 MB** ROWS ceiling - one ordinary ~57 MB `#temp` table read **95.7% full** against the allocation and **0.09%** against the cap. At the shipped 80% default that is a page on the first busy minute of every Azure target, which is what made enabling Azure tempdb collection ([#2512]) unsafe to ship. **The ceiling is discoverable, so this was a denominator bug rather than a policy call**: `tempdb_stats` now collects `max_size_mb` = `SUM(max_size)` over tempdb's **ROWS** files (`dm_db_file_space_usage` reports DATA allocation, so folding the log's cap in would understate usage everywhere), and both SKUs divide by it. Where any data file is unlimited (`max_size` -1) there is no ceiling to measure against and the current allocation stays the denominator - today's behaviour exactly, so **unlimited-growth on-prem and RDS targets do not move**. **What DOES move, deliberately:** any target whose tempdb has a fixed `max_size` it has not yet grown into now reports a **LOWER** percentage than before - 800 MB reserved inside a 1,000 MB allocation that may grow to 4,000 MB reads 20%, not 80% - and climbs back to the alert as the files actually approach the cap. That is a correction, not a regression, but it will move live numbers on existing servers. Deliberately **not** a size floor, the other candidate: a floor suppresses the alert on a genuinely full LARGE tempdb at the moment it starts growing, still fires at the autogrow boundary once cleared, and any value silently redefines "small" for every on-prem and RDS target already depending on today's behaviour. The `analyze_server` `TEMPDB_USAGE` fact carried its own copy of the same denominator and was fixed with it, so the pager and the analysis pass cannot describe one server two ways; the alert detail gained a **Max Size** field naming which denominator produced the percentage (a cap in MB, `Unlimited`, or `Unknown` for a snapshot taken before the ceiling was collected). `FactScorer.ScoreTempDbVersionStore`'s 1/2/5 GB bars are unchanged and needed no change - they are absolute MB and a denominator cannot move their reachability. Darling schema **v81**, Lite schema **v56**; the column is nullable with no backfill, and a historical row's NULL reads as "no ceiling measured", which reports the percentage it always did. +- **Nineteen Darling reads (eighteen on Lite) told an Azure SQL Database user to go and start a capture that cannot exist there** ([#2511]) - `sys.dm_xe_sessions` does not exist on Azure SQL Database (verified live on `EngineEdition` 5, General Purpose and Hyperscale), so the `system_health` ring buffer can never be read on that engine. `SystemHealthEventsCollector` knows this and gates itself off, correctly and permanently. **The reads did not**: all nine `get_health_parser_*` tools answered an Azure caller with *"check that collection is running for this server and that its system_health session is started"* - collection **was** running, and the session cannot be started on that engine. Thirteen collectors gate off on Azure SQL Database and no read anywhere branched on engine edition, so the same confident, specific, wrong instruction was waiting behind `get_server_config`, `get_trace_flags`, `get_default_trace_events`, `get_cpu_scheduler_pressure`, `get_memory_pressure_events`, `get_tempdb_trend`, `get_running_jobs` and the config-change reads too. Every one of those now answers **`not_collected`** - the miss vocabulary's existing word for *"the input names something this server does not collect"* - naming the engine, the collector and the capture path, and saying plainly that the gap is permanent. `unavailable` is kept for a server whose engine **does** support the path and simply has no rows, which is the case that is still worth chasing. **The capability answer is derived from the collectors' own `AppliesTo` gates**, never from a list: a collector is reported as a permanent engine gap only when *no* target of that engine edition runs it, swept across every version, msdb-access and RDS combination - so a gate that reads a **fixable** fact (msdb access, a version floor) is deliberately *not* reported as an engine gap, and a gate that opens up silently stops the claim rather than leaving a stale one standing. +- **Azure SQL Database gets tempdb monitoring back, and the reason it lost it was checkably false** ([#2512]) - `TempDbStatsCollector.AppliesTo` was `!target.IsAzureSqlDb`, so the entire tier - General Purpose and Hyperscale alike - collected **no tempdb data at all**: no `get_tempdb_trend`, no tempdb space alerting, nothing in the viewer's tempdb surfaces, on a platform where the customer cannot go and look at the box themselves. The doc comment justified that by saying the first result set's three-part `tempdb.sys.dm_db_file_space_usage` reference could not be served there and the collector "could only ever fail". Running the collector's SQL verbatim against live Azure SQL Database says otherwise: on `GP_S_Gen5_2` (EngineEdition 5, 12.0.2000.8) it returns user 5.44 MB / internal 1.81 MB / version store 0.00 MB / unallocated 54.19 MB plus one session over threshold, on `HS_S_Gen5_2` user 1.88 MB / unallocated 60.69 MB, and `SELECT COUNT(*) FROM tempdb.sys.dm_db_file_space_usage` returns 4 on both. **Both** result sets bind, so the whole row this collector promises is available - the second one (`sys.dm_db_session_space_usage`) was never in doubt. The numbers also *describe* something rather than merely returning: allocating ~57 MB into a `#temp` table moved `user_mb` 1.88 -> 59.75 and `unallocated_mb` 60.69 -> 2.69, with the session view attributing 59.25 MB to the session that did it. And it is worth **more** here than on a box, because the tempdb ceiling on Azure SQL Database is governed by the SERVICE TIER: you cannot add files and cannot grow past the cap, so "tempdb is filling" means "change service objective", not "go look at the disk". **Managed Instance is unchanged** - it was never gated, it has a real tempdb, and the pin asserts both directions so this cannot be re-narrowed on the strength of a stale comment. What the #2150 field report was actually about is PERMISSIONS, and that is now handled where it belongs: error 262 ("permission denied in database 'tempdb'") was classified `Unclassified`, which means **ERROR every cycle forever** - so a login that cannot read the DMV now degrades to PERMISSIONS with the same kind of explanatory Azure sentence error 300 already gets, rather than polluting collection health. The number set behind that classification lived in **six** hand-maintained places and had already drifted three different ways (916 was in `SqlServerTargetProvider.Classify` and the two failed-jobs msdb reads but in neither collector catch; 8189 was in the collector catches but in neither failed-jobs read; Lite's XE-session catch had neither); all six now read one shared `SqlServerPermissionErrors.IsPermissionDenied`, which is also the only way to pin them, since `SqlException` cannot be constructed in a test. **One thing to know before turning tempdb alerting on for Azure targets**: `unallocated_mb` is free space in the tempdb files *as currently allocated*, and Azure SQL Database creates them small (4 files, ~62 MB total on both 2-vCore tiers measured) and autogrows toward the tier cap - so the `reserved / (reserved + unallocated)` ratio the `tempdb Space` alert fires on measures distance to the next autogrow, not distance to the ceiling. The measurement above takes that ratio from 3% to 96% with a single `#temp` table, against an 80% default. The absolute MB and the trend are the honest signal on this tier; giving the ratio a size floor (the shape `ScorePlanCacheBloat` already uses for the same reason) is filed as [#2515] rather than retuning a threshold every existing on-prem target depends on. Applies to Lite and Darling, which share the collector definition, the classifier and the alert engine +- **Lite's .NET runtime prerequisites were documented against the wrong artifact** ([#2489]) - the README hung "requires .NET 10 Desktop Runtime" on the **`Setup.exe`** download line. That is the one artifact with no prerequisite at all: it is packed by `vpk` from a `--self-contained -r win-x64` publish, and its `runtimeconfig.json` carries `includedFrameworks`, not `frameworks`. Meanwhile the portable **ZIP** - a framework-dependent publish whose `runtimeconfig.json` names **three** shared frameworks - had its prerequisites written down nowhere, and `Lite/README.md` had no prerequisites section at all. So a tester who read the one runtime sentence we published installed a runtime they did not need, while a tester who unzipped onto a stock Windows Server got the .NET host's bare `You must install .NET to run this application` with nothing of ours on it. Each artifact now carries its OWN prerequisites, in the root README and in a new **Prerequisites** section in `Lite/README.md`. **ASP.NET Core Runtime 10 is required unconditionally** for the ZIP, and that is the part nobody expects: `ModelContextProtocol.AspNetCore` brings the framework reference in transitively, so it is named at build time and leaving MCP switched off does not remove it - and because the host reports only the FIRST framework it cannot find, installing just the Desktop Runtime buys a second identical failure. Lite cannot gate on this the way [#2481] did for Darling, because Lite has no install script and the host error precedes every line of our code; what ships instead is `READ-ME-FIRST.txt` beside `PerformanceMonitorLite.exe`, the one surface of ours a stranded reader can still reach. `LiteRuntimePrerequisiteDocsTests` keeps all of it honest by DERIVING the claims - the required runtimes from the built `runtimeconfig.json`, and which artifact is self-contained from the `dotnet publish` lines in `build.yml` and `nightly.yml` - so dropping the MCP package, making the ZIP self-contained, or moving to a new .NET major turns the docs red instead of stale. `README.md` and `Lite/README.md` are named in build.yml's `lite` path filter for the same reason: markdown is carved out of every area filter and a docs-only PR skips .NET setup entirely, so the guard would otherwise be unrunnable on precisely the change it exists to catch. + ## [3.5.0] - 2026-08-19 ### Added @@ -45,7 +164,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **The database-state heal tests no longer assume a best-effort maintenance cycle always runs** ([#2266] item 2, root-caused from source) - `RebaselinedByHandDuringAnOutage_HealsOnceTheDatabaseRecovers` failed on a PR whose diff could not reach it and passed on a re-run of the same commit, reporting `ExpectedState = SUSPECT` alongside `StateDesc = ONLINE`. That pair is only reachable one way: `GetDatabaseStateDeviationsAsync` performs its seeding, the [#2189] heal, the [#2203] forget and the prune inside a block that opens the write connection with a **5-second** lock acquisition and, on `TimeoutException`, skips the entire maintenance block while still running the deviation read - which is deliberate and documented, because skipping is the only lossless option when archival holds the lock. That write lock is **static, shared by the whole process**, and xunit runs test classes in parallel, so another class can hold it long enough for a cycle to skip its maintenance. The test was therefore asserting that the heal lands in ONE cycle, which the design does not promise; it now sweeps until the expectation settles, bounded, which is the actual contract. It cannot mask a regression: a genuinely broken heal never settles, every cycle runs, and the caller's own assertion fails on the final result with its own message exactly as before. Ruled out first: the fixture mints a unique temp directory per instance so no two classes share a store file, and the server id is used by no other class - the shared resource is the LOCK, not the data. **Not fixed here, and worth knowing**: that `catch (TimeoutException)` logs nothing, so in production a sustained contention window means baselines quietly stop being seeded and healed with no evidence anywhere, which matters because [#2189] exists precisely because an unhealed baseline inverts the alert permanently. Item 1 of that issue (the sub-second scale-test comparison) is deliberately untouched: every candidate fix needs the jitter distribution, and guessing a threshold is how an intermittent test stops looking broken without becoming correct. -- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-pos-use2-apex-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it. +- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-sql-use2-alpha-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it. - **The Query Store tick and the Query Store backfill no longer run against one server at the same time** ([#2165]) - the per-tick `query_store` collection and the [#2058] first-contact backfill were independent loops with no per-server coordination at all, and both do heavy Query Store text extraction. Dogfood evidence from a 4-core multi-tenant box mid-consolidation: a 64 MB backfill slice for a freshly restored database ran concurrently with the tick's collection of a SIBLING database - a 12:50:58 backfill ship overlapping a 12:51:09 tick completion - so roughly 128 MB of extraction was in flight at once on the box least able to afford it. That overlap is not bad luck: a big catalog arriving is exactly what triggers BOTH the backfill and budget-bound tick passes, so the two loops collide precisely when the server is already drowning. A per-server gate now excludes them, in both apps - Darling's `DarlingWorker` keyed by server id, Lite's `RemoteCollectorService` keyed by server - sharing one `QueryStoreServerGate` primitive beside `AbandonableStep` (that one bounds how long a step may hold a loop; this one bounds what may run beside it). **Nothing ever waits, deliberately.** Both sides try-acquire with a zero timeout and SKIP on failure, because these are shared fleet loops: an in-flight slice runs to a 180-300 second abandonment deadline, so a blocking acquire would let one slow server stall collection for the entire fleet - the [#2148] wedge arriving through a lock instead of a hang. **Skipping is safe for this collector specifically** because its window is a watermark ([#1960]): the next pass resumes from the same boundary, so a skipped pass defers rows rather than dropping them - which is also why the gate must not be reused for a collector whose window is wall-clock derived. The "tick wins, backfill defers" bias is realized by CADENCE rather than preemption (stopping a statement already running on the monitored server would mean killing it): the tick retries on its own ~1-minute interval against the backfill's 5, so it recovers five times faster from a collision, and a slice is byte-budgeted so it is short in the healthy case. The backfill takes the gate OUTSIDE its `AbandonableStep`, so an abandoned-but-still-wedged slice keeps the gate closed - the statement is genuinely still running on the server and the tick must keep yielding to it. Built on an interlocked flag rather than a `SemaphoreSlim`: a gate that never waits needs none of what a semaphore provides, and would otherwise own one undisposed kernel object per monitored server forever. Leases are idempotent on dispose, because a stray second `Dispose()` would otherwise clear a flag the OTHER loop had since taken and let both run at once - the exact condition being prevented, reached from the wrong direction. @@ -62,8 +181,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **The viewer's store-schema probe is pinned to its reader's arity** - `StoreSchemaProbe_ColumnCount_MatchesTheMapArity`. The probe SQL's columns are read positionally, so adding a sentinel without a parameter made a fully-migrated store report one rung short, and adding a parameter without a column threw `IndexOutOfRange` at connect time. Both were runtime-only against a live store, and the existing arity test builds its arguments from the signature so it agreed with itself either way. - **Darling monitors PostgreSQL and Amazon Aurora PostgreSQL** ([#2213]) - a monitored server can now declare `"engine": "postgres"` and is collected by seven PostgreSQL collectors instead of the T-SQL ones, into the same store, on the same naive-UTC contract and the same `server_id` identity. A mixed fleet is one store, one viewer, one MCP endpoint; nothing is partitioned by engine. Every collector definition declares the engine it targets and is never dispatched against the other one, so a PostgreSQL target is never sent T-SQL and a SQL Server target never sees `pg_stat_statements`. What gets collected: cumulative wait events and per-query-shape execution statistics (both from Aurora's own functions, because core PostgreSQL has no cumulative wait counters in any version), `pg_stat_io` attributed to a (backend type, object, context) triple rather than to a file, per-table autovacuum state stored beside each table's OWN computed trigger threshold, and three outage predictors that have no SQL Server counterpart at all - transaction-ID and MultiXact freeze headroom, what is holding the vacuum horizon back attributed to the specific holder, and replication-slot retained WAL with whether it is still growing. Those three ALERT, graded against the target's own settings rather than a constant, because each names a condition that stops the server outright and each is silent until it is nearly too late. Permissions are one `GRANT pg_monitor` and nothing is created on the monitored instance - no Extended Events sessions to provision, no server setting to bootstrap. `--test-connection` reports what a PostgreSQL target actually is (version, writer or reader, Aurora or not) and how many of the seven collectors will really run against it, which is the difference between "this is configured" and "this will collect": an Aurora writer clears all seven, a reader clears six, a self-hosted PostgreSQL 15 reader clears three. There is a step-by-step runbook with a proof point at every step in `docs/postgres-first-target-runbook.md`. -- **The store can keep plan XML readable over plain SQL: `plan_xml_compression` ([#2171], asked for by @argpna)** - V54 moved plan XML into gzip bytes (`query_plan_dim.query_plan_gz`), which every app surface and MCP tool decompresses client-side - and which nothing reading the store DIRECTLY over SQL can decompress at all, because PostgreSQL exposes no inflate: the field workaround was `plpython3u` plus a hand-rolled gunzip UDF, an untrusted-language extension in a monitoring store just to read your own data. The store now carries a service setting, `plan_xml_compression`, default `gzip` (today's behavior, unchanged): set it to `none` and the dimension writer stores plans as PLAIN TEXT in `query_plan_xml` instead - lz4 TOAST does the compressing (~8.9x measured, vs gzip's 14.0x), and Grafana-class consumers read the column bare, no extension, no UDF, no one-shot container. Flipping it affects NEW rows only, in either direction: the readers' text-first-else-gz resolution already covers every mix of eras and modes, the dimension stays content-addressed so dedup is codec-independent, and nothing rewrites existing rows. The setting rides `config_service` (V62) like its sibling knobs, hot-reloads within one poll, and anything that is not exactly `none` reads as `gzip` - a hand-edited row fails toward the shipped default, and a CHECK constraint enforces the same set DB-side. `--recompress-plan-dim` now refuses to run against a store set to `none`, naming why: it would convert exactly the rows the live writer keeps producing as text, the two fighting forever - set `gzip` back first if that is what you want. The default stays `gzip` because the cost is real at fleet scale (the measured 52-server plan dimension would grow roughly 100 GB to 160 GB); a store whose plans are read by people rather than only by the apps is exactly who should flip it. - +- **The store can keep plan XML readable over plain SQL: `plan_xml_compression` ([#2171], asked for by @argpna)** - V54 moved plan XML into gzip bytes (`query_plan_dim.query_plan_gz`), which every app surface and MCP tool decompresses client-side - and which nothing reading the store DIRECTLY over SQL can decompress at all, because PostgreSQL exposes no inflate: the field workaround was `plpython3u` plus a hand-rolled gunzip UDF, an untrusted-language extension in a monitoring store just to read your own data. The store now carries a service setting, `plan_xml_compression`, default `gzip` (today's behavior, unchanged): set it to `none` and the dimension writer stores plans as PLAIN TEXT in `query_plan_xml` instead - lz4 TOAST does the compressing (~8.9x measured, vs gzip's 14.0x), and Grafana-class consumers read the column bare, no extension, no UDF, no one-shot container. Flipping it affects NEW rows only, in either direction: the readers' text-first-else-gz resolution already covers every mix of eras and modes, the dimension stays content-addressed so dedup is codec-independent, and nothing rewrites existing rows. The setting rides `config_service` (V62) like its sibling knobs, hot-reloads within one poll, and anything that is not exactly `none` reads as `gzip` - a hand-edited row fails toward the shipped default, and a CHECK constraint enforces the same set DB-side. `--recompress-plan-dim` now refuses to run against a store set to `none`, naming why: it would convert exactly the rows the live writer keeps producing as text, the two fighting forever - set `gzip` back first if that is what you want. The default stays `gzip` because the cost is real at fleet scale (the measured 52-server plan dimension would grow roughly 100 GB to 160 GB); a store whose plans are read by people rather than only by the apps is exactly who should flip it. + - **Alerts carry a monotonic occurrence total per incident, not just a rolling-window count** ([#2216], reported by @gotqn) - the `Occurrences` fact on a grouped incident counts events inside the collectors' one-hour read window, so it rises as events arrive and falls as they age out. A consumer that only sees throttled deliveries - one per #1154 per-fingerprint cooldown - cannot recover the true total from a series of those readings, because a 3 followed by a 3 is indistinguishable from nothing-happened and three-happened-while-three-aged-out. Blocking and deadlock incidents now also carry **Total Occurrences**, accumulated per dedup fingerprint over the life of the incident, plus **Incident Since** so a total that restarts can be told from one that continues. Both ride every surface that iterates the alert details (Teams facts, Slack fields, email, the in-app dialog) and persist in the history row's context JSON. The existing `Occurrences` fact is deliberately unchanged: downstream automation keys on fact NAMES, so redefining that one from a gauge to a total would have broken every current consumer silently - the two answer different questions and both are emitted. Keyed per FINGERPRINT rather than per (server, metric), because a deadlock over one object set and a deadlock over another are separate incidents whose totals must not pool. Persisted in both stores (Darling V61 `config.incident_occurrences`, Lite v54 `config_incident_occurrences`) rather than as columns on the existing watermark row, for two independent reasons: the key is wrong, and Lite writes that row with `INSERT OR REPLACE` over a partial column list, which would reset a counter living there to zero on every fired alert - precisely when it is read. **The honest bound**: this is a lower bound on occurrences, exact whenever the read window outlives the gap between OBSERVATIONS - the loss is one occurrence per event that ages out of the window between two of them, because a retirement and an arrival cancel in the gauge. The accumulation therefore runs on every SWEEP, not on every delivery: with a one-hour window and a sweep measured in seconds nothing can arrive and age out inside one interval, so it is exact in practice, whereas observing only at delivery time undercounts a long incident by roughly the number of events the window retired while a cooldown was suppressing delivery. What no window gauge can recover is an occurrence that both arrives AND ages out between two observations - that needs an event-identity watermark at the collector, which is a different feature. A row left behind by a host that died mid-incident is ignored once it is older than the read window, so a stranded counter can never undercount the NEXT incident on the same fingerprint under a stale start time. - **A failing forced plan now alerts** ([#2157]) - when Query Store cannot reproduce a forced plan, the query silently runs on whatever the optimizer picks instead, and nothing in the product witnessed it: the only trace was a counter climbing inside Query Store. The new **Forced Plan Failing** alert fires per plan when `force_failure_count` RISES between collections, carrying the database, query and plan ids, whether the force was MANUAL or automatic plan correction, the engine's own failure reason, and how many failures are new. It is a rise and never a level, because the counter is cumulative AND travels with a restored database - alerting on the level would page forever about failures that happened on hardware you may no longer own. A counter that drops (an unforce/re-force cycle) is silence, and a plan seen for the first time waits one cycle, since 'new' is not knowable from a single observation. Severity is Warning for every rise: a failing force is not an outage, it is somebody's mitigation quietly not working. @@ -80,7 +199,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **The compose `statement_timeout` is a store setting instead of a hardcoded 15s** ([#2357]) - it is the hard backstop a composed query can never exceed, and 15 seconds is a judgement about store size and disk speed that the product cannot make for someone else's deployment. V78 adds `config_service.compose_statement_timeout_seconds`, defaulting to 15 so nothing changes for anyone who does not touch it, clamped to [5,600] on both read and use because a zero would remove the ceiling the whole design leans on. It could not be fixed by raising the constant: the value lives in role PROVISIONING DDL, so an existing install already has the old one baked in - but that DDL re-runs on every managed start, so a change reaches a running store on its next restart with no new machinery. Reported by @carlei1978, who hit it on a 30-day `get_query_store_top`. - **MCP tool results serialize compact instead of pretty-printed** ([#2350]) - the only consumer of a tool result is a language model, and indentation buys a model nothing. One property on the shared `McpHelpers.JsonOptions` in Common, so both SKUs move together, plus the two readers that carry their own options for the `/api/*` twins. Saving is payload-shaped - 23% of the bytes on a 15-field record array, 36% on a narrow one - and the TOKEN saving is smaller than the byte saving, because BPE tokenizers pack runs of spaces efficiently. The config files people hand-edit (`servers.json`, profiles, schedules) keep indenting, and a test pins that boundary in both directions. - **The test suites run under Microsoft.Testing.Platform instead of VSTest** ([#2347]) - the .NET 10 SDK dropped VSTest-mode support for MTP-based frameworks, so xunit.v3 4.0.0 could not be taken at all (blocking two Dependabot bumps). All four suites now build as their own executables and CI invokes them with `dotnet run`; `xunit.runner.visualstudio` and `Microsoft.NET.Test.Sdk` are gone entirely, since the first exists only to bridge xunit to VSTest and the second is the VSTest host. Deliberately done on the CURRENT xunit 3.2.2 so the runner migration and the version bump are separately reviewable. -- **The query_store per-database log split now names the plan-XML and text fetches** ([#2312] investigation) - the per-item sql: stopwatch wraps the whole read, which since the separate fetches landed includes two more queries against the Query Store catalogs after the payload drain - so on a closed-only cycle that shipped ZERO rows, a 298-second bill (measured, ayr-01) had nowhere visible to live: drain silently absorbed it. Cycles that ran a separate fetch now log `... + plan_fetch:Nms + text_fetch:Nms`, the fetch phases come out of drain in the one shipped DrainMsFrom arithmetic (pinned like the #2164 split it extends), and every collector line without a separate fetch is byte-identical to before. Darling-only by construction - Lite runs no separate fetches and its zeros mean exactly that. +- **The query_store per-database log split now names the plan-XML and text fetches** ([#2312] investigation) - the per-item sql: stopwatch wraps the whole read, which since the separate fetches landed includes two more queries against the Query Store catalogs after the payload drain - so on a closed-only cycle that shipped ZERO rows, a 298-second bill (measured, beta-01) had nowhere visible to live: drain silently absorbed it. Cycles that ran a separate fetch now log `... + plan_fetch:Nms + text_fetch:Nms`, the fetch phases come out of drain in the one shipped DrainMsFrom arithmetic (pinned like the #2164 split it extends), and every collector line without a separate fetch is byte-identical to before. Darling-only by construction - Lite runs no separate fetches and its zeros mean exactly that. - **Query Store statement text is now resolved from `collect.query_store_text`, and the separate fetch is ON** ([#2150]) - the flip of `FetchQueryTextSeparately` and the conversion of every reader that projects `query_text`, in one change, because they cannot land apart: flipping the flag nulls the payload's inline `query_sql_text`, so any reader still reading that column would show BLANK text for newly collected rows while looking perfectly healthy - no error, no empty result, just a grid with the statement missing. Six blocks across five files now resolve the side table first and fall back to the inline column: the stored-plan resolver behind Get Actual Plan, the MCP `get_query_store_top` read, the Viewer's Query Store grid, its current-vs-baseline comparison, its regressions grid, and the PLAN_REGRESSION drill-down. **The fallback is permanent, not a migration step**: rows collected before this carry their text inline and nothing backfills them, so removing the fallback later would blank all existing history - the two arms are the two populations, not an old way and a new way. **Where the resolution goes was chosen to keep the filters honest.** Three of these queries also FILTER on the text (the [#1565] `WAITFOR` self-exclusion, and the resolver's `IS NOT NULL`), and testing the raw column there would have excluded every post-cutover row - the whole set the change exists to serve - so each one resolves the text ONCE, inside the existing lateral or a derived table, and the filter tests the resolved value under its original name. That also keeps the diffs to one block each rather than a projection edit plus a filter edit plus a repeated `COALESCE` in the `WHERE`. The comparison read needed `query_id` projected through its dedup CTEs (it groups by `query_hash`, but text is stored per `query_id`) and both its arms converted, since the final projection coalesces current over baseline and converting one arm would leave a GONE row with nothing to fall back to. **Verified against the live 52-server store rather than by reading the SQL**, which is the only instrument that can see any of this - there is no CI job that executes these Postgres strings. All six shipped query bodies were extracted from source, planned with `EXPLAIN (GENERIC_PLAN)` (six plans, every one containing `query_store_text` nodes), then run against real data twice: with the side table EMPTY, each converted query returned a byte-identical md5 to its pre-change form over 330 rows (1 / 50 / 50 / 179 / 50 / 1), proving the fallback arm and proving no keyed join fans out; then with the post-cutover condition INDUCED by shadowing the table with distinctive rows, all 331 rows resolved from the side table at identical counts, proving the arm that will serve every row once this ships. `FetchRowsAsync` deliberately stays OFF - it hands rows straight to its caller and writes nothing, so nulling the inline column there would lose the text outright instead of relocating it - and Lite is untouched, its DuckDB store having no side table to resolve from. - **Lite's Query Store prune names the state keys it retired, like Darling's** ([#2205]) - Darling's prune uses `DELETE ... RETURNING` and logs which keys went; Lite's DuckDB twin summed the affected-row counts into a number. Correctness was already identical - same anti-join, same freshness guard, pinned by the [#2195] tests - so this is purely forensic, and it matters on this path specifically because the symptom of a *wrong* delete here is a silent refetch, which leaves nothing else behind to diagnose it with. A count of three cannot be told apart from a mistaken prune of the same size. @@ -157,7 +276,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **The Job History tab speaks display names** ([#2126], asked by ghauan) - both the Server filter dropdown and the Server column showed the raw collected server name while every other tab shows the operator's alias, so a fleet navigated by aliases turned into a memory quiz on exactly the tab an operator visits during an incident. Both readers (job history and the Agent status header) now resolve through the servers registry - the alias when one exists, the raw name otherwise - so the filter, the column, the per-column filter popup, and the CSV export all speak the same names as the rest of the viewer, and the Agent roll-up sorts by them. Lite's Job History tab had the same gap through a different mechanism (review catch): Lite's display-name concept lives at the CONFIG layer, not in DuckDB (the stored servers.display_name column is unpopulated by design), so the shell now passes a server_id-to-alias snapshot into the tab Overview-style and rows swap in the alias on every refresh - a server no longer in config keeps its raw collected name, the durable-record case. - **The long-query completion XE session actually gets created now** ([#2129], from ghauan's field report on #2061 - they enabled the collector on two servers and the Long Queries tab stayed empty forever) - the session DDL SET a customizable attribute `collect_object_name` on `sqlserver.rpc_completed`, and no such attribute exists on that event on ANY version (it belongs to `sp_statement_completed`) - `object_name` is one of rpc_completed's DEFAULT data fields, collected with no SET at all. So the CREATE failed on every server, the session never existed, and the reconcile's follow-up START surfaced as the confusing second error ('Cannot alter the event session... does not exist'). Never caught in dogfood because the collector ships OFF by design, and the DDL test pin asserted the wrong claim, so CI enforced the bug. The SET is gone (the reader already shreds the default field generically - no reader or table change), and the pin now asserts the attribute is ABSENT, with the story attached. Anyone who flipped the collector on before this fix: it starts working on the next reconcile tick after upgrading, no re-toggle needed. - **`--collapse-legacy-slices` narrows its slice instead of dying when a day does not fit the statement timeout** ([#2105] round three, ghauan once more - with the decompression rail lifted, the run made it ~15 minutes in and died at the NEW wall: a day-wide stage aggregation on a store carrying 60k split intervals blows through the 15-minute per-statement timeout, and the operator got the same bare stream exception) - the verb's fixed day-per-slice loop now runs the same adaptive schedule the Query Store backfill worker shipped this week (`AdaptiveSpan`, 24h base): a failed slice halves the window and retries the SAME start (announced with a [RETRY] line naming the error, so narrowing reads as progress rather than a hang), a completed slice resets to full width, and only a slice that fails at the ~22-minute floor gives up to the existing idempotent re-run message. Healthy stores still repair in a handful of day-wide slices - the narrowing costs nothing until a slice actually fails. -- **Query Store collection no longer has a fixed cost that big catalogs cannot pay** ([#2133], the actual root cause under the whole catch-up saga - #2102's death spiral, #2111's yield, and #2125's adaptive shrink were all mitigating it) - the collector joined its slice aggregate straight into the query_store_plan/query/text catalog TVFs, handing the optimizer nothing but fixed-guess cardinalities, and the plan it picked re-materialized a TVF per probe: on an 82k-plan catalog that was a fixed 30-second-plus cost that NO catch-up window width could reduce - which is exactly why the fleet's big databases (echo, oak, Surge, spruce, insa...) pinned at the 15-minute shrink floor and never converged while their smaller neighbors on the same servers stayed current. Bisected live: the aggregate alone ran in 81 ms and each TVF scanned bare in ~300 ms, yet aggregate-JOIN-plan could not finish in 30 s, hinted or not. The payload now STAGES the aggregate in a temp table and joins FROM it - real row counts instead of guesses, each TVF scanned exactly once, sp_QuickieStore's architecture for the same reason - and the old LOOP JOIN hint is gone for good (looping from the temp into the TVFs is the same per-probe re-materialization by another name). The interval pre-filter also resolves ids from the tiny interval catalog now instead of scanning runtime_stats itself (20 ms vs 426 ms, same id set). Measured end to end on the wedged field store: the full 55-column batch with plan capture completed a one-hour backlog in 21.3 s where the old shape never finished inside 60; the staged core is 524 ms. Same batch = one result set, TOP WITH TIES / derived-watermark / byte-budget semantics unchanged, both SKUs, both engine arms, live and backfill. +- **Query Store collection no longer has a fixed cost that big catalogs cannot pay** ([#2133], the actual root cause under the whole catch-up saga - #2102's death spiral, #2111's yield, and #2125's adaptive shrink were all mitigating it) - the collector joined its slice aggregate straight into the query_store_plan/query/text catalog TVFs, handing the optimizer nothing but fixed-guess cardinalities, and the plan it picked re-materialized a TVF per probe: on an 82k-plan catalog that was a fixed 30-second-plus cost that NO catch-up window width could reduce - which is exactly why the fleet's big databases (echo, oak, Surge, spruce, AppDatabaseOne...) pinned at the 15-minute shrink floor and never converged while their smaller neighbors on the same servers stayed current. Bisected live: the aggregate alone ran in 81 ms and each TVF scanned bare in ~300 ms, yet aggregate-JOIN-plan could not finish in 30 s, hinted or not. The payload now STAGES the aggregate in a temp table and joins FROM it - real row counts instead of guesses, each TVF scanned exactly once, sp_QuickieStore's architecture for the same reason - and the old LOOP JOIN hint is gone for good (looping from the temp into the TVFs is the same per-probe re-materialization by another name). The interval pre-filter also resolves ids from the tiny interval catalog now instead of scanning runtime_stats itself (20 ms vs 426 ms, same id set). Measured end to end on the wedged field store: the full 55-column batch with plan capture completed a one-hour backlog in 21.3 s where the old shape never finished inside 60; the staged core is 524 ms. Same batch = one result set, TOP WITH TIES / derived-watermark / byte-budget semantics unchanged, both SKUs, both engine arms, live and backfill. - **`--collapse-legacy-slices` runs ~40% fewer window scans and narrates its progress** (#2105 operator feedback - ghauan's successful run 'took some time', all of it against a silent console) - each slice bracketed its work with two window-wide COUNT(*) scans purely to compute the removed-row figure, and on the catch-up-bloated stores this verb exists for, one such scan measures ~12 seconds (hash-aggregate spill to temp disk plus a backward index scan over the uncompressed hot chunk) - paid twice more per slice for bookkeeping. The removed count now derives from the DELETE and INSERT statements' own affected-row counts (deleted minus reinserted IS the net removal, same number by construction), and every completed slice prints an [OK] line with its window, removal count, and percent-of-span - so a long backlog reads as visible progress instead of a hang. ## [3.4.0] - 2026-08-06 @@ -183,9 +302,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **Version Store (PVS) pressure alert in both apps** ([#1984], the alerting follow-up [#1951] deferred) - a new alert (default on) fires when an ADR database's persistent version store reaches a share of the database's own data files - 40% by default, warning meaningfully before the "close to 50% of the database size" that Microsoft's troubleshooting guide calls large. Percent-of-database rather than absolute size because a shipped absolute guess is workload-specific and would page half a fleet (the `ag_redo_queue_alert_kb` precedent shipped OFF at 0 for exactly that reason); the ratio is also what both FinOps grids already compute as PVS % of Database, so the alert and the grid can never tell different stories. A second knob, the **GB floor** (default 1 GB), is an AND qualifier rather than the volume alert's either-breach-fires OR: a 10 MB database at 60% is six megabytes, and nobody should be paged for six megabytes - 0 removes the floor, and a percent of 0 disables the check outright since percent is the alert's only trigger. One alert per server names the worst (highest-share) database with up to five breaching databases in the context, each carrying its PVS size, data-file denominator, aborted-transaction count, whether the aborted-version cleaner is mid-run (Microsoft's start-time-without-end-time shape), and the aborted/active transaction-id lag - presented as the gap itself, never a verdict, for the same reason the grids refused to invent a threshold Microsoft does not document. No severity tier either, for the same reason. **Re-fires only on a fresh or WORSENING breach** (5 percentage points over the last-alerted level): measured on a live rig, PVS space stays allocated even after the pinning transaction clears and the cleaner runs to completion, so a plain per-cooldown level check would re-notify for hours after the incident ended - the same standing-condition treatment the volume alert got, with the direction flipped because a version store worsens by rising. Fires through the shared cross-app alert path, so Lite and Darling evaluate identically from their stored `pvs_stats` (hourly), with the same cooldown, mute, alert-history, tray, and email plumbing, and a resolved notice when every version store drops back under. Darling's knobs ride the store control plane (V48 migration; the Viewer gains the Settings row, its schema gate a V48 rung, and `get_alert_settings`/`update_alert_settings` a `pvs` group); Lite's live in `settings.json` (`alert_pvs_*`) with its Settings row. Validated end-to-end against a Docker SQL2022 rig carrying real ADR pressure - a pinned-cleanup workload grown to 78% PVS-of-database - through the real collector, store, read-adapter, and engine path. -- **Lite's data now lives OUTSIDE the install directory - and on any version older than this one, upgrading with `Setup.exe` destroys it** ([#1832]) - Lite kept everything in `%LOCALAPPDATA%\PerformanceMonitorLite`, which is also the installer's own install root. Re-running `Setup.exe` over an existing install renames that folder aside and deletes it, so an installer-based upgrade took the DuckDB store, the Parquet archive, the logs, and `settings.json` with it, in one move, before any code from the new build ran. **In-app updates were always safe** - Help > About replaces the `current\` subfolder in place and never touches the data sitting beside it - which is precisely why the loss looked random instead of reproducible. Two things always survived and still do: the monitored-server list (`%ProgramData%\PerformanceMonitorLite\config\servers.json`, machine-wide) and every password and webhook URL in Windows Credential Manager. +- **Lite's data now lives OUTSIDE the install directory - and on any version older than this one, upgrading with `Setup.exe` destroys it** ([#1832]) - Lite kept everything in `%LOCALAPPDATA%\PerformanceMonitorLite`, which is also the installer's own install root. Re-running `Setup.exe` over an existing install renames that folder aside and deletes it, so an installer-based upgrade took the DuckDB store, the Parquet archive, the logs, and `settings.json` with it, in one move, before any code from the new build ran. **In-app updates were always safe** - the About button in the left sidebar replaces the `current\` subfolder in place and never touches the data sitting beside it - which is precisely why the loss looked random instead of reproducible. Two things always survived and still do: the monitored-server list (`%ProgramData%\PerformanceMonitorLite\config\servers.json`, machine-wide) and every password and webhook URL in Windows Credential Manager. - **The one-time caveat: reaching THIS version via `Setup.exe` still loses your data.** The installer deletes the old directory before the new build starts, so no fix shipping inside the new build can prevent it - the code that would migrate your data has not run yet when the data is deleted. Upgrade **in place** instead and nothing is lost: **Help > About** downloads and applies the update, or extract the portable ZIP over your existing copy. From this version forward `Setup.exe` is safe, because the data is no longer where the installer writes. + **The one-time caveat: reaching THIS version via `Setup.exe` still loses your data.** The installer deletes the old directory before the new build starts, so no fix shipping inside the new build can prevent it - the code that would migrate your data has not run yet when the data is deleted. Upgrade **in place** instead and nothing is lost: the **About** button in the left sidebar downloads and applies the update, or extract the portable ZIP over your existing copy. From this version forward `Setup.exe` is safe, because the data is no longer where the installer writes. Data now lives in the sibling `%LOCALAPPDATA%\PerformanceMonitorLite-Data`, and the first start MOVES `config\`, `archive\`, `logs\`, `monitor.duckdb` (plus its WAL) and `alert_state.json` there from the old location automatically - a same-volume rename, so a multi-GB store transfers instantly. Nothing already in the new location is overwritten, nothing anywhere is deleted, and a `DATA-MOVED.txt` signpost is left behind pointing at the new path. The move is per-artifact rather than all-or-nothing, so a run interrupted by a locked store file finishes itself on the next start rather than stranding the store in a folder the next installer run would delete. The updater's own `Update.exe`, `current\` and `packages\` share that old folder and are deliberately left where they are. @@ -2839,6 +2958,35 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 [#2312]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2312 [#2306]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2306 [#2302]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2302 +[#2501]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2501 +[#2499]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2499 +[#2495]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2495 +[#2506]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2506 +[#2511]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2511 +[#2512]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2512 +[#2515]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2515 +[#2520]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2520 +[#2525]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2525 +[#2530]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2530 +[#2539]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2539 +[#2540]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2540 +[#2508]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2508 +[#2541]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2541 +[#2542]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2542 +[#2532]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2532 +[#2518]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2518 +[#2533]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2533 +[#2535]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2535 +[#2480]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2480 +[#2489]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2489 +[#2481]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2481 +[#2437]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2437 +[#1563]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/1563 +[#2475]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2475 +[#2422]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2422 +[#2460]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2460 +[#2459]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/2459 +[#2446]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2446 [#2296]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2296 [#2299]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2299 [#2294]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2294 @@ -2847,6 +2995,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 [#2233]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2233 [#2234]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2234 [#2150]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2150 +[#2484]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2484 +[#2485]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2485 +[#2524]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2524 [#2201]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2201 [#2205]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2205 [#2195]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2195 +[#2548]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2548 +[#2554]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2554 +[#2546]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2546 +[#2552]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2552 +[#2529]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2529 +[#2569]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2569 +[#2572]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2572 +[#2574]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2574 +[#2579]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2579 +[#2559]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2559 +[#2562]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2562 +[#2550]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2550 +[#2578]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2578 +[#2564]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2564 +[#2653]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2653 +[#2655]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2655 +[#2658]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2658 +[#2659]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2659 +[#2661]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2661 +[#2663]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2663 +[#2665]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/2665 diff --git a/Darling/Darling.Tests/AgAlertPolicyTests.cs b/Darling/Darling.Tests/AgAlertPolicyTests.cs index e56f61c73..8445b5dce 100644 --- a/Darling/Darling.Tests/AgAlertPolicyTests.cs +++ b/Darling/Darling.Tests/AgAlertPolicyTests.cs @@ -6,6 +6,7 @@ * Licensed under the MIT License. See LICENSE file in the project root for full license information. */ +using System; using PerformanceMonitor.Common; using Xunit; @@ -81,6 +82,102 @@ public void DecideConnection_AlreadyDisconnectedAtFirstSighting_StaysSilent_ButI Assert.Equal(AgConnectionDecision.Reconnected, AgAlertPolicy.DecideConnection("DISCONNECTED", "CONNECTED")); } + /* ---------------- #2426: the disconnect re-fire ---------------- */ + + private static readonly DateTime Noon = new(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc); + + [Theory] + [InlineData("CONNECTED", "DISCONNECTED", AgConnectionDecision.Disconnected)] + [InlineData("DISCONNECTED", "CONNECTED", AgConnectionDecision.Reconnected)] + [InlineData("DISCONNECTED", "DISCONNECTED", AgConnectionDecision.None)] + [InlineData("CONNECTED", "CONNECTED", AgConnectionDecision.None)] + [InlineData(null, "DISCONNECTED", AgConnectionDecision.None)] + [InlineData("CONNECTED", null, AgConnectionDecision.None)] + public void DecideConnection_RefireOff_IsTheEdgeOnlyOverloadExactly( + string? previous, string? current, AgConnectionDecision expected) + { + /* The shipped default, and the whole matrix rather than one case: null, zero and a negative all + mean OFF, and off has to be byte-for-byte the two-argument behavior or an upgrade would start + re-alerting on a knob nobody set. */ + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, null, null, Noon)); + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, TimeSpan.Zero, null, Noon)); + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, TimeSpan.FromMinutes(-5), null, Noon)); + } + + [Fact] + public void DecideConnection_StillDisconnected_WaitsOutTheWindow_ThenSaysItAgain() + { + var refire = TimeSpan.FromMinutes(10); + + /* Announced at noon: inside the window there is nothing new to say. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddMinutes(9))); + + /* At the boundary, and still hours later — the point of the knob is that a week-long outage does + not read like a blip. */ + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddMinutes(10))); + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddHours(8))); + } + + [Fact] + public void DecideConnection_ARealEdgeOutranksARefire() + { + var refire = TimeSpan.FromMinutes(10); + + /* The edge that OPENS the outage announces as Disconnected even though the window is trivially due + on it, or the caller would have two reasons to announce the same sweep. */ + Assert.Equal( + AgConnectionDecision.Disconnected, + AgAlertPolicy.DecideConnection("CONNECTED", "DISCONNECTED", refire, null, Noon)); + + /* And a recovery is a recovery whatever the clock says. */ + Assert.Equal( + AgConnectionDecision.Reconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "CONNECTED", refire, Noon.AddHours(-9), Noon)); + } + + [Fact] + public void DecideConnection_NoStampIsDueNow_SoARestartMidOutageStillReAnnounces() + { + var refire = TimeSpan.FromMinutes(10); + + /* Both apps hold this edge state in memory, so a restart during a week-long outage sees a replica + already DISCONNECTED with no record of it ever having been announced. Rule 1's silent baseline + would make that silence permanent, which is the exact failure the knob exists to prevent — so + with re-fire ON, and only then, a first sighting announces. */ + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection(null, "DISCONNECTED", refire, null, Noon)); + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, null, Noon)); + + /* With it off, rule 1 governs unchanged. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection(null, "DISCONNECTED", null, null, Noon)); + } + + [Fact] + public void DecideConnection_ARefireStillNeedsAnExactDisconnected() + { + var refire = TimeSpan.FromMinutes(10); + + /* Same rule the edge follows, and it matters more here: a re-fire pages repeatedly, so a state + string the product never learned to interpret must not become a standing page. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("SOMETHING_NEW", "SOMETHING_NEW", refire, null, Noon)); + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("CONNECTED", null, refire, null, Noon)); + } + /* ---------------- suspension ---------------- */ [Theory] diff --git a/Darling/Darling.Tests/AlertEngineTests.cs b/Darling/Darling.Tests/AlertEngineTests.cs index 8ad3b34e2..f944c60bf 100644 --- a/Darling/Darling.Tests/AlertEngineTests.cs +++ b/Darling/Darling.Tests/AlertEngineTests.cs @@ -1052,6 +1052,12 @@ public async Task TempDb_FiresAtThreshold_AndResolutionCarriesTheCurrentPercent( h.Settings.TempDbSpaceEnabled = true; var engine = h.Build(); + /* #2515: 1000 MB total and NO MaxSizeMb, which is deliberate and stays that way. The fixture predates + the ceiling, so it describes a tempdb whose cap was never measured — and that is exactly the case + where the denominator remains the allocation. It does NOT imply a 1000 MB cap; if it did, this pin + would have to move, and the fact that it does not is the guarantee that no existing on-prem or RDS + target with an unlimited (or uncollected) tempdb sees its number change. The capped case gets its + own test below rather than being folded in here. */ h.Adapter.TempDb = new TempDbSpaceInfo { TotalReservedMb = 800, UnallocatedMb = 200 }; /* 80% used */ await engine.EvaluateServerAsync(Harness.Snapshot()); var fired = Assert.Single(h.Deliverer.Outcomes); @@ -1066,6 +1072,36 @@ public async Task TempDb_FiresAtThreshold_AndResolutionCarriesTheCurrentPercent( Assert.Equal("SRV-A: tempdb usage back to 20%", resolution.Message); /* :461,:464 */ } + /// + /// #2515, through the ENGINE rather than the arithmetic: the Azure shape from the issue must not fire, and + /// the same allocation without a ceiling must. Same 59.75 MB of reserved tempdb in both, same 80% default — + /// the only difference is whether the collector could see how far the files are allowed to grow. + /// + /// This is the assertion the whole change exists for. proves + /// the arithmetic and the store round-trip; this proves the alert engine's decision follows it, which is + /// what actually pages someone. + /// + [Fact] + public async Task TempDb_TheAzureCeiling_SuppressesTheAlert_ThatTheAllocationWouldFire() + { + var h = new Harness(); + h.Settings.TempDbSpaceEnabled = true; + Assert.Equal(80, h.Settings.TempDbSpaceThresholdPercent); + var engine = h.Build(); + + /* GP_S_Gen5_2 with one ~57 MB #temp table: 62.44 MB allocated, 65,536 MB of headroom behind it. */ + h.Adapter.TempDb = new TempDbSpaceInfo { TotalReservedMb = 59.75, UnallocatedMb = 2.69, MaxSizeMb = 65_536 }; + await engine.EvaluateServerAsync(Harness.Snapshot()); + Assert.Empty(h.Deliverer.Outcomes); + + /* The identical snapshot with the ceiling unmeasured is the pre-#2515 reading, and it pages. */ + h.Adapter.TempDb = new TempDbSpaceInfo { TotalReservedMb = 59.75, UnallocatedMb = 2.69 }; + await engine.EvaluateServerAsync(Harness.Snapshot()); + var fired = Assert.Single(h.Deliverer.Outcomes); + Assert.Equal("tempdb Space", fired.MetricName); + Assert.Equal("96% used (60 MB)", fired.CurrentValue); + } + /* ---------------- low disk ---------------- */ [Fact] diff --git a/Darling/Darling.Tests/AnalysisAsOfAnchorTests.cs b/Darling/Darling.Tests/AnalysisAsOfAnchorTests.cs new file mode 100644 index 000000000..8dc0c60b8 --- /dev/null +++ b/Darling/Darling.Tests/AnalysisAsOfAnchorTests.cs @@ -0,0 +1,376 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Text.Json; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Analysis; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Analysis; +using PerformanceMonitor.Darling.Service.Mcp; +using PerformanceMonitor.Darling.Storage; +using Xunit; + +namespace Darling.Tests; + +/// +/// #2506: the ANALYSIS family's window anchor. #2495 anchored the read surface and deliberately left +/// these four out, because their window is built inside the analysis engine rather than in the tool and +/// because analyze_server PERSISTS what it finds. +/// +/// These are the ungated halves: the engine's own rule about when a pass may write, and the shape +/// of the refusal. The live half — seed a past window, analyze it anchored, analyze it unanchored, and +/// prove the two disagree — is . +/// +public sealed class AnalysisAsOfAnchorTests +{ + /// + /// The whole persistence decision, stated as one derivable rule rather than as a flag a caller has to + /// remember to set. A settable flag would let "anchored AND persist" be expressed, and there is no + /// caller that should be able to express it. + /// + [Fact] + public void AnAnchoredPass_DoesNotPersist_AndAnUnanchoredOneStillDoes() + { + Assert.True(new AnalysisContext().PersistFindings); + + Assert.True(new AnalysisContext + { + TimeRangeStart = new DateTime(2026, 8, 19, 1, 0, 0, DateTimeKind.Utc), + TimeRangeEnd = new DateTime(2026, 8, 19, 5, 0, 0, DateTimeKind.Utc) + }.PersistFindings); + + Assert.False(new AnalysisContext + { + TimeRangeStart = new DateTime(2026, 8, 19, 1, 0, 0, DateTimeKind.Utc), + TimeRangeEnd = new DateTime(2026, 8, 19, 5, 0, 0, DateTimeKind.Utc), + AsOfUtc = new DateTime(2026, 8, 19, 5, 0, 0, DateTimeKind.Utc) + }.PersistFindings); + } + + /// + /// The rule is on the ANCHOR, not on how old the window happens to be. An anchor a few seconds in the + /// past — what an agent sends when it computes "now" from its own clock — is still an anchor and still + /// exploratory, because the alternative is a persistence decision that depends on how long the pass + /// took to start, which is not a decision anybody could reason about. + /// + [Fact] + public void AnAnchorAtEssentiallyNow_IsStillAnAnchor() + { + Assert.False(new AnalysisContext { AsOfUtc = DateTime.UtcNow }.PersistFindings); + } + + /// + /// The window builder honours the anchor: hours_back stays the LENGTH and the anchor becomes the + /// END. Asserted on the engine's own entry point rather than on a tool, because that is the seam #2495 + /// could not reach and the whole reason this was a separate issue. + /// + [Fact] + public void TheEngineEntryPoint_BuildsTheWindowEndingAtTheAnchor() + { + var anchor = new DateTime(2026, 8, 19, 5, 0, 0, DateTimeKind.Utc); + + /* Built the same way DarlingAnalysisService.AnalyzeAsync builds it — the arithmetic is the + contract, and it is one line in both SKUs. */ + var context = new AnalysisContext + { + TimeRangeEnd = anchor, + TimeRangeStart = anchor.AddHours(-4), + AsOfUtc = anchor + }; + + Assert.Equal(new DateTime(2026, 8, 19, 1, 0, 0, DateTimeKind.Utc), context.TimeRangeStart); + Assert.Equal(4 * 60 * 60 * 1000d, context.PeriodDurationMs); + Assert.False(context.PersistFindings); + } +} + +/// +/// Gated (DARLING_TEST_PG) proof of the point: an incident planted in a PAST window is analyzed when the +/// call is anchored there and is invisible on the default anchor — and the anchored run leaves the +/// findings table exactly as it found it, while the SAME window analyzed unanchored writes to it. +/// +[Collection("live-postgres")] +public sealed class AnalysisAsOfAnchorLivePostgresTests +{ + private const string ServerName = "darling-analysis-asof-e2e"; + private const string WaitType = "AN2506_TEST_WAIT"; + private const string HistoricStoryHash = "an2506-historic-hash"; + private const string AheadStoryHash = "an2506-ahead-of-now-hash"; + private static readonly int ServerId = ServerIdHelper.GetDeterministicHashCode(ServerName); + + private static string? ConnectionString => Environment.GetEnvironmentVariable("DARLING_TEST_PG"); + + [Fact] + public async Task TheAnalysisFamilyAnswersAboutAPastWindow_AndTheAnchoredRunPersistsNothing() + { + var cs = ConnectionString; + Assert.SkipWhen(string.IsNullOrEmpty(cs), "Set DARLING_TEST_PG to a Postgres connection string to run the live analysis as_of test."); + + var ct = TestContext.Current.CancellationToken; + using var connection = new NpgsqlConnection(cs); + await connection.OpenAsync(ct); + await PgMigrations.MigrateAsync(connection, ct); + await DeleteRowsAsync(connection, ct); + + await using var postgres = NpgsqlDataSource.Create(cs!); + + var bodySucceeded = false; + try + { + await RegisterServerAsync(connection, ct); + var service = new DarlingAnalysisService(postgres); + + /* ── the fixture, DarlingAnalysisPipelineTests' recipe verbatim ── + >24h of wait_stats history in ONE hour-of-day x day-of-week bucket (a past Monday 10:00 + UTC, 8-14 days back), then three heavy collections on the FOLLOWING Monday at the same + hour. The span carries the pipeline's 24h data-span gate on real MIN/MAX arithmetic, and + the heavy window rates at ~666 ms/sec against a 250 ms/sec fallback bar, which is one + ANOMALY_WAIT_PROFILE fact — enough to root a story. */ + var day = DateTime.UtcNow.Date.AddDays(-8); + while (day.DayOfWeek != DayOfWeek.Monday) day = day.AddDays(-1); + var historyStart = Naive(day.AddHours(10)); + + for (var i = 0; i < 12; i++) + { + await PlantWaitAsync(connection, i + 1, historyStart.AddMinutes(5 * i), (i == 5 || i == 6) ? 0L : 60_000L, ct); + } + + var incidentStart = historyStart.AddDays(7); + await PlantWaitAsync(connection, 100, incidentStart.AddMinutes(5), 200_000L, ct); + await PlantWaitAsync(connection, 101, incidentStart.AddMinutes(10), 200_000L, ct); + await PlantWaitAsync(connection, 102, incidentStart.AddMinutes(15), 200_000L, ct); + + /* The anchor: one hour AFTER the incident's first collection, so `hours_back = 1` names + exactly the window the pipeline fixture was built for. Deliberately not "the incident plus + a bit either side": the baseline is keyed on the window's START hour, and a window starting + at 09:30 would be scored against a different hour-of-day bucket than the one the history + fills — a real behaviour of this engine, and one an anchored test should not stumble into + by accident. */ + var incidentEnd = incidentStart.AddHours(1); + var anchor = Utc(incidentEnd).ToString("o"); + + /* ── 1. get_analysis_facts: the incident's window scores facts; the default window has none. + Both branches are reachable — the planted rows are >= 7 days old, so no legal default + window (the widest on this tool is its own hours_back) can reach them. */ + var factsAnchored = await DarlingMcpTools.GetAnalysisFacts(service, postgres, ServerName, 1, as_of: anchor); + Assert.Contains(WaitType, factsAnchored, StringComparison.Ordinal); + + var factsNow = await DarlingMcpTools.GetAnalysisFacts(service, postgres, ServerName); + Assert.Equal("unavailable", StatusOf(factsNow)); + + /* ── 2. analyze_server anchored: the anomaly is found, and the result says out loud that it + was not written down. */ + var analysisAnchored = await DarlingMcpTools.AnalyzeServer(service, postgres, ServerName, 1, anchor); + using (var doc = JsonDocument.Parse(analysisAnchored)) + { + Assert.Equal("findings", doc.RootElement.GetProperty("status").GetString()); + Assert.False(doc.RootElement.GetProperty("persisted").GetBoolean()); + Assert.Contains( + "NOT written to the store", + doc.RootElement.GetProperty("persistence_note").GetString()!, + StringComparison.Ordinal); + } + + Assert.Contains("ANOMALY_WAIT_PROFILE", analysisAnchored, StringComparison.Ordinal); + + /* ── 3. …and it really did not write. This is the assertion the whole persistence argument + rests on: a note in the payload that the row was withheld is worth nothing if the row + is there anyway. */ + Assert.Equal(0L, await CountFindingsAsync(connection, ct)); + + /* ── 4. The SAME window, unanchored, DOES persist — so step 3 is the anchor's doing and not + a fixture that could never produce a finding in the first place. Without this the + previous assertion passes for a server that simply has nothing to say. */ + var persistedFindings = await service.AnalyzeAsync(new AnalysisContext + { + ServerId = ServerId, + ServerName = ServerName, + TimeRangeStart = incidentStart, + TimeRangeEnd = incidentEnd, + ServerUtcOffset = TimeSpan.Zero + }); + + Assert.NotEmpty(persistedFindings); + Assert.True(await CountFindingsAsync(connection, ct) > 0); + + /* ── 5. get_analysis_findings windows on ANALYSIS TIME. The rows step 4 just wrote are + stamped NOW, and a row planted as if a scheduled pass had run 30 hours ago is outside + every default window (the tool's default is 24). Anchored at that pass, the read + returns it and NOT the rows from now — which is what proves the upper bound is doing + work rather than the read simply starting earlier. */ + var historicRun = Naive(DateTime.UtcNow.AddHours(-30)); + await PlantFindingAsync(connection, historicRun, ct); + + var findingsNow = await DarlingMcpTools.GetAnalysisFindings(service, postgres, ServerName); + + /* The default read must actually be RETURNING something — step 4's rows — or the + DoesNotContain below would pass on a read that came back empty for an unrelated reason, + which is an assertion that cannot fail and therefore proves nothing. */ + Assert.True( + JsonDocument.Parse(findingsNow).RootElement.GetProperty("finding_count").GetInt32() >= 1, + "the default findings read returned nothing, so the exclusion below would prove nothing"); + Assert.DoesNotContain(HistoricStoryHash, findingsNow, StringComparison.Ordinal); + + var findingsAnchored = await DarlingMcpTools.GetAnalysisFindings( + service, postgres, ServerName, 4, false, Utc(historicRun.AddMinutes(30)).ToString("o")); + Assert.Contains(HistoricStoryHash, findingsAnchored, StringComparison.Ordinal); + + /* Exactly one: the anchored window's UPPER bound is what keeps step 4's now-stamped rows out. + A read that only moved its START earlier would return them too, and would look like a + working anchor right up until someone counted. */ + Assert.Equal(1, JsonDocument.Parse(findingsAnchored).RootElement.GetProperty("finding_count").GetInt32()); + + /* ── 6. compare_analysis hangs BOTH windows off the anchor — the peak-vs-a-week-ago + comparison the tool exists for, which was unreachable while both windows were pinned + to now. 168 and not 169: baseline_hours_back is a span like any other and + ValidateHoursBack caps it at MaxHoursBack, so the baseline lands an hour past the + history rows rather than on them. That does not weaken the claim — the claim is that + both bounds MOVED WITH THE ANCHOR, and it is asserted on the rendered instants rather + than on a field path, so it holds for the data-bearing shape and for the "neither + window had facts" status envelope alike. */ + var compared = await DarlingMcpTools.CompareAnalysis(service, postgres, ServerName, 1, 168, anchor); + Assert.Contains(Utc(incidentEnd).ToString("o"), compared, StringComparison.Ordinal); + Assert.Contains(Utc(incidentEnd.AddHours(-168)).ToString("o"), compared, StringComparison.Ordinal); + + /* The control: unanchored, the same two instants are nowhere in the payload, because both + windows hang off now. Without this the assertions above would pass on a tool that echoed + the anchor into the payload and windowed on the clock. */ + var comparedNow = await DarlingMcpTools.CompareAnalysis(service, postgres, ServerName, 1, 168); + Assert.DoesNotContain(Utc(incidentEnd).ToString("o"), comparedNow, StringComparison.Ordinal); + + /* ── 7. The UNANCHORED findings read keeps its half-open window, which is not a detail: + bounding it at "now" broke a live test on the first attempt, and the mechanism behind + that is real. analysis_time is stamped by the WRITER and filtered by the READER, so a + default read bounded at the reader's clock drops a run written a moment earlier the + day those two clocks stop being the same one. A row stamped ahead of now is the cheap, + deterministic stand-in for that skew. */ + await PlantFindingAsync(connection, Naive(DateTime.UtcNow.AddMinutes(30)), ct, AheadStoryHash); + Assert.Contains( + AheadStoryHash, + await DarlingMcpTools.GetAnalysisFindings(service, postgres, ServerName), + StringComparison.Ordinal); + + /* …and the ANCHORED read still excludes it, so the bound is present exactly where it is + needed and absent exactly where it would do harm. */ + Assert.DoesNotContain( + AheadStoryHash, + await DarlingMcpTools.GetAnalysisFindings( + service, postgres, ServerName, 4, false, Utc(historicRun.AddMinutes(30)).ToString("o")), + StringComparison.Ordinal); + + /* ── 8. Refusals reach the caller as the tool's own message, on the persisting tool too — a + bad anchor must never fall back to "now" and then run a real, persisting analysis. */ + Assert.StartsWith( + "Invalid as_of", + await DarlingMcpTools.AnalyzeServer(service, postgres, ServerName, 1, "last tuesday"), + StringComparison.Ordinal); + Assert.Contains( + "future", + await DarlingMcpTools.AnalyzeServer(service, postgres, ServerName, 1, DateTime.UtcNow.AddDays(1).ToString("o")), + StringComparison.Ordinal); + + bodySucceeded = true; + } + finally + { + await LiveStoreCleanup.RunAsync(cs!, bodySucceeded, async (cleanup, cleanupCt) => + await DeleteRowsAsync(cleanup, cleanupCt)); + } + } + + private static string StatusOf(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("status").GetString()!; + } + + /// Naive-UTC, the Kind every timestamp column in this store is bound with. + private static DateTime Naive(DateTime value) => DateTime.SpecifyKind(value, DateTimeKind.Unspecified); + + /// The same instant re-labelled UTC, so ToString("o") renders the Z form as_of parses. + private static DateTime Utc(DateTime value) => DateTime.SpecifyKind(value, DateTimeKind.Utc); + + private static async Task CountFindingsAsync(NpgsqlConnection connection, CancellationToken ct) + { + using var command = new NpgsqlCommand("SELECT COUNT(*) FROM analysis_findings WHERE server_id = $1", connection); + command.Parameters.AddWithValue(ServerId); + return Convert.ToInt64(await command.ExecuteScalarAsync(ct) ?? 0L); + } + + private static async Task RegisterServerAsync(NpgsqlConnection connection, CancellationToken ct) + { + using var command = new NpgsqlCommand(@" +INSERT INTO servers (server_id, server_name, display_name, is_enabled, sql_major_version, created_date, modified_date) +VALUES ($1, $2, $3, TRUE, 15, $4, $4) +ON CONFLICT (server_id) DO UPDATE SET is_enabled = TRUE, sql_major_version = 15;", connection); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(Naive(DateTime.UtcNow)); + await command.ExecuteNonQueryAsync(ct); + } + + private static async Task PlantWaitAsync( + NpgsqlConnection connection, long collectionId, DateTime at, long waitMs, CancellationToken ct) + { + using var command = new NpgsqlCommand(@" +INSERT INTO wait_stats + (collection_id, collection_time, server_id, server_name, wait_type, delta_waiting_tasks, delta_wait_time_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7)", connection); + command.Parameters.AddWithValue(collectionId); + command.Parameters.AddWithValue(at); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(WaitType); + command.Parameters.AddWithValue(50L); + command.Parameters.AddWithValue(waitMs); + await command.ExecuteNonQueryAsync(ct); + } + + /// + /// One finding row as a scheduled pass would have left it 30 hours ago. Planted rather than produced, + /// because producing it would mean an analysis whose analysis_time is stamped NOW — the very + /// thing the anchored run refuses to do. + /// + private static async Task PlantFindingAsync( + NpgsqlConnection connection, DateTime analysisTime, CancellationToken ct, string? storyPathHash = null) + { + using var command = new NpgsqlCommand(@" +INSERT INTO analysis_findings + (finding_id, analysis_time, server_id, server_name, database_name, + time_range_start, time_range_end, severity, confidence, category, + story_path, story_path_hash, story_text, + root_fact_key, root_fact_value, leaf_fact_key, leaf_fact_value, fact_count, + incident_id, remediation_action_json, drill_down_json) +VALUES ($1, $2, $3, $4, NULL, $5, $6, 0.9, 0.8, 'waits', + 'AN2506_HISTORIC', $7, 'planted for the #2506 anchored findings read', + 'AN2506_HISTORIC', 1, NULL, NULL, 1, NULL, NULL, NULL)", connection); + command.Parameters.AddWithValue(CollectionIdGenerator.Next()); + command.Parameters.AddWithValue(analysisTime); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(analysisTime.AddHours(-4)); + command.Parameters.AddWithValue(analysisTime); + command.Parameters.AddWithValue(storyPathHash ?? HistoricStoryHash); + await command.ExecuteNonQueryAsync(ct); + } + + private static async Task DeleteRowsAsync(NpgsqlConnection connection, CancellationToken ct) + { + using var cleanup = new NpgsqlCommand( + $"DELETE FROM wait_stats WHERE server_id = {ServerId}; " + + $"DELETE FROM analysis_findings WHERE server_id = {ServerId}; " + + $"DELETE FROM analysis_muted WHERE server_id = {ServerId}; " + + $"DELETE FROM servers WHERE server_id = {ServerId};", connection); + await cleanup.ExecuteNonQueryAsync(ct); + } +} diff --git a/Darling/Darling.Tests/AnalysisPassTokenThreadingTests.cs b/Darling/Darling.Tests/AnalysisPassTokenThreadingTests.cs new file mode 100644 index 000000000..713a88bce --- /dev/null +++ b/Darling/Darling.Tests/AnalysisPassTokenThreadingTests.cs @@ -0,0 +1,368 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Text.RegularExpressions; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Analysis; +using PerformanceMonitor.Darling.Analysis; +using Xunit; + +namespace Darling.Tests; + +/// +/// The analysis pass's token has to reach the reads, not just the gaps between them (#2443). +/// +/// #2438 armed a per-pass budget and #2430's classifier told a timeout from a stop, so an +/// overrunning server became loudly stuck instead of silently skipped forever. What neither did was +/// hand the token to the store layer: 168 of this project's async store calls took the no-argument +/// overload, so abandonment happened BETWEEN stages. A pass gave up while its wedged query kept a +/// connection and a session alive until the query finished on its own — on a healthy fleet +/// invisible, and on the pathological case these fixes exist for, the monitoring tool holding a +/// session it had already stopped waiting for. +/// +/// The threading itself is mechanical; this file is the part that lasts. A sweep this size +/// only stays swept if something counts it, and the failure mode it guards against is specific: +/// adding CancellationToken cancellationToken = default to signatures and stopping there, +/// which leaves every call site passing while LOOKING threaded. +/// So the pin is on the CALL, not the signature — a no-argument ExecuteReaderAsync() is what +/// goes red, and no default can hide it. +/// +public sealed class AnalysisPassTokenThreadingTests +{ + /// + /// The no-argument overloads. Each has a token-taking sibling, so the empty parentheses are the + /// whole tell — there is no shape of correct code in this project that needs them. + /// + private static readonly Regex s_untokenedStoreCall = new( + @"\.(?:ExecuteReaderAsync|ExecuteNonQueryAsync|ExecuteScalarAsync|ReadAsync|OpenAsync|OpenConnectionAsync)\(\s*\)", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + /// + /// A member declaration, used only to attribute a call to the method that makes it. Deliberately + /// crude — it needs to name the enclosing method, not parse C#. The [^=;()]* is what keeps + /// field initializers and expression-bodied properties out. + /// + private static readonly Regex s_memberDeclaration = new( + @"^\s*(?:public|private|internal|protected)[^=;()]*\s(?[A-Za-z_][A-Za-z0-9_]*)\s*\(", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + private const string ExemptionMarker = "#2443 exempt"; + + /// + /// Every method allowed to make an untokened store call, and the reason it is allowed to. This is + /// an enumeration and not a wildcard on purpose: the exemption list this repo spent a day removing + /// was a wildcard, and the way that one grew was that nobody had to name a new entry. + /// + /// Two kinds live here and they are not the same kind. The read-back surfaces are OFF the + /// pass — the viewer, the MCP and the retention sweep have no per-pass budget and no wedged + /// analysis to abandon, so there is no token to thread and inventing one would be pretending. + /// InsertFindingAsync is the other kind: it is ON the pass and the token is deliberately + /// withheld, because a finding set cut in half is a third outcome that reads like neither of the + /// two the classifier names (see the method's own note, and + /// ). + /// + private static readonly Dictionary s_exempt = new(StringComparer.Ordinal) + { + ["GetRecentFindingsAsync"] = "read-back: the MCP + viewer findings read, no pass to abandon", + ["GetLatestFindingsAsync"] = "read-back: the viewer's Recommendations tab, no pass to abandon", + ["GetMutedStoriesAsync"] = "read-back: the viewer's mute registry read, no pass to abandon", + ["MuteStoryAsync"] = "off-pass write: the MCP/viewer mute verb, no pass to abandon", + ["UnmuteStoryAsync"] = "off-pass write: the viewer's unmute verb, no pass to abandon", + ["CleanupOldFindingsAsync"] = "off-pass write: the worker's retention sweep, its own lifetime", + ["InsertFindingAsync"] = "on-pass write, token withheld: a half-written finding set must not exist" + }; + + /// + /// The sweep, and the thing that keeps it swept. Every store call in the analysis project must + /// take the pass's token, and the ones that do not must be exempt BY NAME with a marker in their + /// own doc comment — so the reason travels with the code rather than living only here. + /// + /// Scanned over the whole project directory rather than a file list, because a file list is + /// a thing a new partial can be added outside of. A seventh PgFactCollector partial is + /// covered the day it exists. + /// + [Fact] + public void NoStoreCallOnTheAnalysisPassRunsWithoutThePassToken() + { + var offenders = new List(); + + foreach (var (file, lines) in AnalysisSources()) + { + var enclosing = string.Empty; + + for (var i = 0; i < lines.Length; i++) + { + var declaration = s_memberDeclaration.Match(lines[i]); + if (declaration.Success) + { + enclosing = declaration.Groups["name"].Value; + } + + if (!s_untokenedStoreCall.IsMatch(lines[i])) + { + continue; + } + + if (s_exempt.ContainsKey(enclosing) && DocBlockAbove(lines, IndexOfDeclaration(lines, i)).Contains(ExemptionMarker, StringComparison.Ordinal)) + { + continue; + } + + offenders.Add($"{file}:{i + 1} in {enclosing}(): {lines[i].Trim()}"); + } + } + + Assert.True(offenders.Count == 0, + "These store calls run without the pass's cancellation token, so the pass can be abandoned " + + "around them but never inside one. Pass context.CancellationToken (or the method's own token " + + $"parameter). If the call genuinely must complete, add its method to {nameof(s_exempt)} and put a " + + $"'{ExemptionMarker}' note in its doc comment saying why:\n" + + string.Join("\n", offenders)); + } + + /// + /// The other half of the agreement: an exemption that no longer corresponds to an untokened call is + /// a claim the code stopped making, and it must not be allowed to sit there authorizing the next + /// one. Every entry must still be earning its place, and every marker in the source must belong to + /// an entry. + /// + [Fact] + public void EveryExemptionIsStillEarningItsPlace() + { + var untokenedBy = new Dictionary(StringComparer.Ordinal); + var markedMethods = new HashSet(StringComparer.Ordinal); + + foreach (var (_, lines) in AnalysisSources()) + { + var enclosing = string.Empty; + var enclosingAt = -1; + + for (var i = 0; i < lines.Length; i++) + { + var declaration = s_memberDeclaration.Match(lines[i]); + if (declaration.Success) + { + enclosing = declaration.Groups["name"].Value; + enclosingAt = i; + if (DocBlockAbove(lines, enclosingAt).Contains(ExemptionMarker, StringComparison.Ordinal)) + { + markedMethods.Add(enclosing); + } + } + + if (s_untokenedStoreCall.IsMatch(lines[i]) && enclosingAt >= 0) + { + untokenedBy[enclosing] = untokenedBy.TryGetValue(enclosing, out var n) ? n + 1 : 1; + } + } + } + + Assert.Equal(s_exempt.Keys.OrderBy(k => k, StringComparer.Ordinal), markedMethods.OrderBy(k => k, StringComparer.Ordinal)); + Assert.Equal(s_exempt.Keys.OrderBy(k => k, StringComparer.Ordinal), untokenedBy.Keys.OrderBy(k => k, StringComparer.Ordinal)); + + /* The counts are stated so a NEW untokened call inside an already-exempt method cannot ride in + on the exemption. 15 read-back calls across six methods, and the one finding INSERT. */ + Assert.Equal(16, untokenedBy.Values.Sum()); + Assert.Equal(1, untokenedBy["InsertFindingAsync"]); + } + + /// + /// The fact collector is where the token would have been silently useless. Its per-query catches were + /// bare catch { } — deliberately, so a missing table degrades to "no facts" — which meant an + /// armed token produced thirty-three swallowed cancellations and a pass that carried on collecting + /// under a token that had already fired. Every catch there now lets an abandonment through, which is + /// what turns the threaded token into an exit rather than thirty-three wasted connection opens. + /// + [Fact] + public void EveryCatchOnTheFactCollectorLetsAnAbandonmentThrough() + { + var bare = new List(); + var opensCatch = new Regex(@"^\s*catch\b", RegexOptions.Compiled | RegexOptions.CultureInvariant); + var classified = new Regex( + @"^\s*catch\s*\(\s*Exception\s+ex\s*\)\s*when\s*\(\s*!AnalysisShutdown\.IsExpectedAbandon\(", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + foreach (var (file, lines) in AnalysisSources()) + { + if (!Path.GetFileName(file).StartsWith("PgFactCollector.", StringComparison.Ordinal)) + { + continue; + } + + for (var i = 0; i < lines.Length; i++) + { + if (opensCatch.IsMatch(lines[i]) && !classified.IsMatch(lines[i])) + { + bare.Add($"{file}:{i + 1}: {lines[i].Trim()}"); + } + } + } + + Assert.True(bare.Count == 0, + "A fact-collector catch that does not classify swallows the abandonment the token was armed for, " + + "and the pass keeps collecting under a fired token:\n" + string.Join("\n", bare)); + + /* 27 collect methods guard their read this way; the number is stated so deleting one is a decision. + The other four collect methods have no catch at all and never had one — their reads propagate + straight to the pass, which is the same outcome by a shorter route. */ + Assert.Equal(27, AnalysisSources() + .Where(s => Path.GetFileName(s.File).StartsWith("PgFactCollector.", StringComparison.Ordinal)) + .SelectMany(s => s.Lines) + .Count(line => classified.IsMatch(line))); + } + + /// + /// The behavioural half, and the reason a source pin alone is not enough here. Reverted, this test + /// fails with an NpgsqlException rather than an , which + /// is the whole defect in one line: handed a token that had ALREADY fired, + /// still dialled the store. The collectors that carry + /// a catch then swallowed the failure and the pass carried on to the next one, so a cancelled pass + /// spent thirty-three connection opens finding out what the token had said before the first. + /// + /// No store is needed to run this, precisely because the token is already cancelled — a pass + /// that observes it never reaches the network. + /// + [Fact] + public async Task AFiredTokenStopsFactCollectionAtTheFirstReadInsteadOfRunningEveryCollector() + { + await using var dataSource = NpgsqlDataSource.Create( + "Host=127.0.0.1;Port=1;Username=nobody;Password=nobody;Database=nowhere;Timeout=1;Command Timeout=1"); + + var collector = new PgFactCollector(dataSource); + var context = new AnalysisContext + { + ServerId = 1, + ServerName = "wedged", + TimeRangeStart = DateTime.UtcNow.AddHours(-4), + TimeRangeEnd = DateTime.UtcNow, + CancellationToken = new CancellationToken(canceled: true) + }; + + await Assert.ThrowsAnyAsync( + () => collector.CollectFactsAsync(context)); + } + + /// + /// The decision this issue asked to be made explicitly rather than assumed: what a cancelled PERSIST + /// means. Every row shares one analysis_time and the latest-findings read keys on + /// MAX(analysis_time), so a batch cut in half does not read as truncated; it reads as a + /// complete analysis that found fewer problems, and the server looks HEALTHIER for having been + /// abandoned. That is a third outcome, distinct from the timeout and the shutdown the classifier + /// names, and it must not exist. The connection open is therefore the last abandonment point: before + /// the first row, or not at all. + /// + /// #2448 closed the other half of that. #2443 could only reason about the CANCELLATION path, + /// and the same truncated set was still reachable from an ordinary store fault mid-batch, where no + /// amount of token discipline reaches. The batch is now one transaction, so it is all-or-nothing + /// against a fault as well — which is why the between-rows assertion below is still right, but no + /// longer the only thing standing between a store fault and a set that understates a server. Both + /// halves are pinned here because they are one property. + /// + [Fact] + public void TheFindingInsertIsAbandonedBeforeItStartsOrNotAtAll() + { + var store = ReadSource(Path.Combine(AnalysisDirectory(), "PgFindingStore.cs")); + + /* The abandonment point: the batch's connection open observes the pass token. */ + Assert.Contains( + "await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken);", + store, StringComparison.Ordinal); + + /* And nothing after it does. A ThrowIfCancellationRequested between rows would MAKE the partial + set rather than prevent it, which is why there is none. */ + var insertBatch = Between(store, "public async Task> InsertFindingsAsync(", "public async Task> SaveFindingsAsync("); + Assert.DoesNotContain("ThrowIfCancellationRequested", insertBatch, StringComparison.Ordinal); + + /* #2448: the batch is one transaction, and the rows are enlisted in it. Both halves are pinned + because either alone is silently useless — a transaction the rows do not join commits nothing + of theirs, and enlisting in a transaction nobody commits writes nothing at all. */ + Assert.Contains("await connection.BeginTransactionAsync();", insertBatch, StringComparison.Ordinal); + Assert.Contains("await transaction.CommitAsync();", insertBatch, StringComparison.Ordinal); + Assert.Contains("new NpgsqlCommand(InsertFindingSql, connection, transaction)", store, StringComparison.Ordinal); + + /* And the row write must not swallow. A per-row catch inside a transaction is worse than useless: + PostgreSQL refuses every later statement once one fails (25P02), so it buys N-k ERROR lines for + one event and a CommitAsync that returns normally having written nothing. */ + var rowWrite = Between(store, "private async Task InsertFindingAsync(", "private static AnalysisFinding ReadFinding("); + Assert.DoesNotContain("catch", rowWrite, StringComparison.Ordinal); + + /* The row write states the decision where someone changing it will read it. */ + Assert.Contains(ExemptionMarker, Between(store, "/// Inserts one finding on an already-open connection", "private async Task InsertFindingAsync("), StringComparison.Ordinal); + } + + private static string Between(string source, string start, string end) + { + var from = source.IndexOf(start, StringComparison.Ordinal); + Assert.True(from >= 0, $"anchor not found: {start}"); + var to = source.IndexOf(end, from, StringComparison.Ordinal); + Assert.True(to > from, $"anchor not found after {start}: {end}"); + return source[from..to]; + } + + /// The declaration line a call at belongs to (scanning up). + private static int IndexOfDeclaration(string[] lines, int callLine) + { + for (var i = callLine; i >= 0; i--) + { + if (s_memberDeclaration.IsMatch(lines[i])) + { + return i; + } + } + + return 0; + } + + /// The contiguous /// block immediately above a declaration, as one string. + private static string DocBlockAbove(string[] lines, int declarationLine) + { + var doc = new List(); + for (var i = declarationLine - 1; i >= 0 && lines[i].TrimStart().StartsWith("///", StringComparison.Ordinal); i--) + { + doc.Add(lines[i]); + } + + return string.Join("\n", doc); + } + + private static IEnumerable<(string File, string[] Lines)> AnalysisSources() + { + var files = Directory.GetFiles(AnalysisDirectory(), "*.cs", SearchOption.TopDirectoryOnly); + Assert.NotEmpty(files); + + foreach (var file in files.OrderBy(f => f, StringComparer.Ordinal)) + { + /* The working copy is CRLF; split on the LF so a line never carries a stray CR. */ + yield return (Path.GetFileName(file), ReadSource(file).Replace("\r\n", "\n", StringComparison.Ordinal).Split('\n')); + } + } + + private static string AnalysisDirectory() => + Path.Combine(RepoRoot(), "Darling", "PerformanceMonitor.Darling.Analysis"); + + private static string RepoRoot([CallerFilePath] string thisFile = "") + { + var dir = Path.GetDirectoryName(thisFile)!; + while (dir is not null && !File.Exists(Path.Combine(dir, "PerformanceMonitor.sln")) && !Directory.Exists(Path.Combine(dir, ".git"))) + { + dir = Path.GetDirectoryName(dir); + } + + Assert.NotNull(dir); + return dir!; + } + + private static string ReadSource(string path) => File.ReadAllText(path); +} diff --git a/Darling/Darling.Tests/AnalysisShutdownResidueTests.cs b/Darling/Darling.Tests/AnalysisShutdownResidueTests.cs index 5e5d11932..219076d1b 100644 --- a/Darling/Darling.Tests/AnalysisShutdownResidueTests.cs +++ b/Darling/Darling.Tests/AnalysisShutdownResidueTests.cs @@ -49,13 +49,13 @@ public sealed class AnalysisShutdownResidueTests [Fact] public void EveryShutdownShapeIsAbandonedOnceTheTokenFires() { - Assert.True(AnalysisShutdown.IsShutdownAbandon(new OperationCanceledException(), s_fired)); - Assert.True(AnalysisShutdown.IsShutdownAbandon(new ObjectDisposedException("NpgsqlDataSource"), s_fired)); - Assert.True(AnalysisShutdown.IsShutdownAbandon( + Assert.True(AnalysisShutdown.IsExpectedAbandon(new OperationCanceledException(), s_fired)); + Assert.True(AnalysisShutdown.IsExpectedAbandon(new ObjectDisposedException("NpgsqlDataSource"), s_fired)); + Assert.True(AnalysisShutdown.IsExpectedAbandon( new NpgsqlException("wrapper", new ObjectDisposedException("NpgsqlDataSource")), s_fired)); - Assert.True(AnalysisShutdown.IsShutdownAbandon(SqlState("57P01"), s_fired)); - Assert.True(AnalysisShutdown.IsShutdownAbandon(SqlState("57P02"), s_fired)); - Assert.True(AnalysisShutdown.IsShutdownAbandon(SqlState("57P03"), s_fired)); + Assert.True(AnalysisShutdown.IsExpectedAbandon(SqlState("57P01"), s_fired)); + Assert.True(AnalysisShutdown.IsExpectedAbandon(SqlState("57P02"), s_fired)); + Assert.True(AnalysisShutdown.IsExpectedAbandon(SqlState("57P03"), s_fired)); } /// @@ -66,9 +66,9 @@ public void EveryShutdownShapeIsAbandonedOnceTheTokenFires() [Fact] public void TheSameShapesStayErrorsWhileTheServiceIsRunning() { - Assert.False(AnalysisShutdown.IsShutdownAbandon(new OperationCanceledException(), CancellationToken.None)); - Assert.False(AnalysisShutdown.IsShutdownAbandon(new ObjectDisposedException("NpgsqlDataSource"), CancellationToken.None)); - Assert.False(AnalysisShutdown.IsShutdownAbandon(SqlState("57P01"), CancellationToken.None)); + Assert.False(AnalysisShutdown.IsExpectedAbandon(new OperationCanceledException(), CancellationToken.None)); + Assert.False(AnalysisShutdown.IsExpectedAbandon(new ObjectDisposedException("NpgsqlDataSource"), CancellationToken.None)); + Assert.False(AnalysisShutdown.IsExpectedAbandon(SqlState("57P01"), CancellationToken.None)); } /// @@ -78,20 +78,67 @@ public void TheSameShapesStayErrorsWhileTheServiceIsRunning() [Fact] public void ATimeoutIsNeverRelabelledAsShutdown() { - Assert.False(AnalysisShutdown.IsShutdownAbandon(new TimeoutException("deadline"), s_fired)); - Assert.False(AnalysisShutdown.IsShutdownAbandon( + Assert.False(AnalysisShutdown.IsExpectedAbandon(new TimeoutException("deadline"), s_fired)); + Assert.False(AnalysisShutdown.IsExpectedAbandon( new NpgsqlException("Exception while reading from stream", new TimeoutException()), s_fired)); - Assert.False(AnalysisShutdown.IsShutdownAbandon(SqlState("57014"), s_fired)); + Assert.False(AnalysisShutdown.IsExpectedAbandon(SqlState("57014"), s_fired)); + } + + /// + /// #2430. Once the pass token is a BUDGET linked from the stopping token, "the token fired" stops + /// meaning "we are stopping" — and this is the pin that keeps the two apart. Arming the token + /// without this split is what would have made every ordinary overrun on a healthy service report + /// itself as "abandoned at shutdown", at Information, on exactly the signal someone would use to + /// decide the budget needs raising. + /// + [Fact] + public void ABudgetExpiryIsATimeoutAndNotAShutdown() + { + Assert.Equal( + AnalysisAbandonKind.Timeout, + AnalysisShutdown.Classify(new OperationCanceledException(), CancellationToken.None, s_fired)); + + /* And a stop is still a stop, including when it lands on a pass that had already overrun — + reporting that as a timeout would invent an incident out of a clean Stop-Service. */ + Assert.Equal( + AnalysisAbandonKind.Shutdown, + AnalysisShutdown.Classify(new OperationCanceledException(), s_fired, s_fired)); + Assert.Equal( + AnalysisAbandonKind.Shutdown, + AnalysisShutdown.Classify(new ObjectDisposedException("NpgsqlDataSource"), s_fired, s_fired)); + } + + /// + /// The timeout arm is narrower than the shutdown arm on purpose, and this is why. A disposed data + /// source and a 57P0x mean the STORE went away, which a budget expiring on a running service does + /// not cause — so during the window after any pass overruns, those must keep the ERROR #2299 gave + /// them rather than being relabelled as something we asked for. Widening this arm to the whole + /// residue set would erase that bug's only evidence for every server that ever times out. + /// + [Fact] + public void AStoreThatVanishesIsNeverExcusedByTheBudget() + { + Assert.Equal( + AnalysisAbandonKind.None, + AnalysisShutdown.Classify(new ObjectDisposedException("NpgsqlDataSource"), CancellationToken.None, s_fired)); + Assert.Equal( + AnalysisAbandonKind.None, + AnalysisShutdown.Classify(SqlState("57P01"), CancellationToken.None, s_fired)); + + /* Nothing at all fired: a fault is a fault. */ + Assert.Equal( + AnalysisAbandonKind.None, + AnalysisShutdown.Classify(new OperationCanceledException(), CancellationToken.None, CancellationToken.None)); } /// Ordinary faults during a stop are still faults — structural shapes only, never "anything goes". [Fact] public void OrdinaryFaultsAreNeverShutdownResidueEvenMidStop() { - Assert.False(AnalysisShutdown.IsShutdownAbandon(SqlState("42703"), s_fired)); - Assert.False(AnalysisShutdown.IsShutdownAbandon( + Assert.False(AnalysisShutdown.IsExpectedAbandon(SqlState("42703"), s_fired)); + Assert.False(AnalysisShutdown.IsExpectedAbandon( new NpgsqlException("Exception while reading from stream", new IOException("reset")), s_fired)); - Assert.False(AnalysisShutdown.IsShutdownAbandon(new InvalidOperationException("something else"), s_fired)); + Assert.False(AnalysisShutdown.IsExpectedAbandon(new InvalidOperationException("something else"), s_fired)); } /// @@ -103,7 +150,7 @@ public void OrdinaryFaultsAreNeverShutdownResidueEvenMidStop() [Fact] public void EveryErrorLoggingCatchOnTheAnalysisPassClassifiesShutdown() { - const string contextFilter = "when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken))"; + const string contextFilter = "when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken))"; /* The detector: NO bare catch is permitted at all — its nine identical per-detector catches were the bulk of the burst, and a tenth detector must arrive classified. Every `catch (Exception ex)` @@ -113,12 +160,12 @@ public void EveryErrorLoggingCatchOnTheAnalysisPassClassifiesShutdown() Assert.Equal( Count(detector, "catch (Exception ex)"), Count(detector, "catch (Exception ex) " + contextFilter) - + Count(detector, "catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken))")); + + Count(detector, "catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken))")); Assert.True(Count(detector, contextFilter) >= 9, "a detector catch lost its shutdown classification"); /* The baseline provider: its single catch produced five of the seven lines. */ var baseline = ReadSource(Path.Combine("Darling", "PerformanceMonitor.Darling.Analysis", "PgBaselineProvider.cs")); - Assert.Equal(1, Count(baseline, "when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken))")); + Assert.Equal(1, Count(baseline, "when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken))")); /* The finding store: only its two PASS methods run under the worker's token; its read-back surfaces serve other lifetimes and are deliberately untouched. */ @@ -133,15 +180,15 @@ surfaces serve other lifetimes and are deliberately untouched. */ /* The service: the ONE Information line a stop is allowed to cost, and the data-span probe must not convert shutdown residue into a bogus "0 hours of history" skip. */ var service = ReadSource(Path.Combine("Darling", "PerformanceMonitor.Darling.Analysis", "DarlingAnalysisService.cs")); - Assert.Equal(1, Count(service, "when (AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken))")); - Assert.Equal(1, Count(service, "when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken))")); + Assert.Equal(1, Count(service, "AnalysisShutdown.Classify(ex, context.ShutdownToken, context.CancellationToken)")); + Assert.Equal(1, Count(service, "when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken))")); Assert.Contains("Analysis abandoned at shutdown", service, StringComparison.Ordinal); /* The worker: the pass must RECEIVE the stopping token (an uncancellable pass makes every filter above unreachable), and the stop path must hold the sweep open for the unwind grace with the already-fired token deliberately not forwarded. */ var worker = ReadSource(Path.Combine("Darling", "PerformanceMonitor.Darling.Service", "DarlingWorker.cs")); - Assert.Contains("AnalyzeAsync(serverId, storageName, hoursBack: 4, stoppingToken)", worker, StringComparison.Ordinal); + Assert.Contains("serverId, storageName, hoursBack: 4, cts.Token, stoppingToken", worker, StringComparison.Ordinal); Assert.Contains("WaitAsync(s_analysisShutdownGrace, CancellationToken.None)", worker, StringComparison.Ordinal); } diff --git a/Darling/Darling.Tests/AsOfWindowAnchorTests.cs b/Darling/Darling.Tests/AsOfWindowAnchorTests.cs new file mode 100644 index 000000000..3464d6cc7 --- /dev/null +++ b/Darling/Darling.Tests/AsOfWindowAnchorTests.cs @@ -0,0 +1,659 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Text.Json; +using System.Text.RegularExpressions; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Service; +using PerformanceMonitor.Darling.Service.Mcp; +using PerformanceMonitor.Darling.Storage; +using Xunit; + +namespace Darling.Tests; + +/// +/// #2495: the window ANCHOR. Every windowed read took hours_back measured from now, so nothing +/// could ask about a past incident — and widening hours_back until the incident falls inside is +/// NOT the same question, because for an aggregate read a wider window is a different answer rather than +/// the same answer with more rows. +/// +/// These are the ungated halves: the shared resolver's contract, and the two surfaces agreeing about +/// which reads carry the parameter. The live half — seed a past window, read it anchored, read it +/// unanchored, prove the two disagree — is . +/// +public sealed class AsOfWindowAnchorTests +{ + /* ── the resolver's contract ── */ + + [Fact] + public void NoAnchor_ResolvesToNow_WhichIsThePreChangeBehaviour() + { + var before = DateTime.UtcNow; + Assert.Null(McpHelpers.ResolveAsOf(null, out var end)); + var after = DateTime.UtcNow; + + Assert.InRange(end, before, after); + Assert.Equal(DateTimeKind.Utc, end.Kind); + } + + [Fact] + public void BlankAnchor_IsTreatedAsAbsent_NotAsAParseFailure() + { + var before = DateTime.UtcNow; + Assert.Null(McpHelpers.ResolveAsOf(" ", out var end)); + Assert.True(end >= before); + } + + /// + /// The four accepted spellings all land on ONE instant. The offset-less form matters most: it is read + /// as UTC, never as the SERVICE HOST's local time — the store is UTC throughout and the caller is an + /// agent on some other machine, so a local-time reading would silently shift the window by the host's + /// offset with nothing in the result to show for it. + /// + [Theory] + [InlineData("2026-08-18T14:30:00Z")] + [InlineData("2026-08-18T14:30:00")] + [InlineData("2026-08-18T16:30:00+02:00")] + [InlineData("2026-08-18T09:30:00-05:00")] + [InlineData("2026-08-18T14:30Z")] + [InlineData("2026-08-18T14:30:00.000Z")] + [InlineData(" 2026-08-18T14:30:00Z ")] + public void EveryAcceptedSpelling_ResolvesToTheSameUtcInstant(string asOf) + { + Assert.Null(McpHelpers.ResolveAsOf(asOf, out var end)); + Assert.Equal(new DateTime(2026, 8, 18, 14, 30, 0, DateTimeKind.Utc), end); + Assert.Equal(DateTimeKind.Utc, end.Kind); + } + + [Fact] + public void DateOnly_IsMidnightUtc() + { + Assert.Null(McpHelpers.ResolveAsOf("2026-08-18", out var end)); + Assert.Equal(new DateTime(2026, 8, 18, 0, 0, 0, DateTimeKind.Utc), end); + } + + /// + /// An anchor we cannot use is REFUSED, following . Silently falling + /// back to now is the one outcome this parameter exists to prevent: a read answering "the last 4 hours" + /// when it was asked for "the 4 hours ending Tuesday 03:00" is indistinguishable from a correct answer. + /// + /// The slash forms are the interesting half, and the reason the parser is an ISO-8601 ALLOWLIST + /// rather than a general date parse (review catch): a plain DateTime.TryParse under the invariant + /// culture accepts 01/02/2026 as M/d/yyyy, so a caller who meant 1 February gets a window + /// around 2 January and nothing anywhere says so. That is the same defect one step removed — an answer to + /// a question nobody asked — so it is refused rather than guessed at. + /// + [Theory] + [InlineData("last tuesday")] + [InlineData("2026-13-45")] + [InlineData("4 hours ago")] + [InlineData("08/18/2026")] + [InlineData("01/02/2026")] + [InlineData("2026/08/18")] + [InlineData("18 Aug 2026")] + [InlineData("2026-08-18 14:30:00")] + public void AnUnusableAnchor_IsRefused_NotSilentlyTreatedAsNow(string asOf) + { + var error = McpHelpers.ResolveAsOf(asOf, out _); + Assert.NotNull(error); + Assert.Contains("Invalid as_of", error, StringComparison.Ordinal); + Assert.Contains(asOf, error, StringComparison.Ordinal); + } + + [Fact] + public void AFutureAnchor_IsRefused_BecauseTheStoreCannotHoldDataNotYetCollected() + { + var error = McpHelpers.ResolveAsOf(DateTime.UtcNow.AddHours(2).ToString("o"), out _); + Assert.NotNull(error); + Assert.Contains("future", error, StringComparison.Ordinal); + } + + /// + /// The future refusal has a clock-skew allowance, and it is not a grace period for asking about the + /// future: an agent that computes "now" from its own clock and sends it must not be refused because + /// that clock runs a minute fast, on a stored read whose newest row is minutes old anyway. + /// + [Fact] + public void AClientClockRunningSlightlyFast_IsStillAccepted() + { + var justInside = DateTime.UtcNow + McpHelpers.AsOfFutureTolerance - TimeSpan.FromMinutes(1); + Assert.Null(McpHelpers.ResolveAsOf(justInside.ToString("o"), out _)); + } + + /// + /// There is deliberately NO lower bound. An anchor older than anything the store holds is a legitimate + /// question whose honest answer is the read's own empty / unavailable status — and the + /// caller knows the anchor they sent, so that status is unambiguous. A hardcoded floor would have to + /// guess at retention, which is per-deployment, per-server and per-collector. + /// + [Fact] + public void AnAnchorOlderThanAnyRetention_IsAccepted_AndLeftToTheReadsOwnMissVocabulary() + { + Assert.Null(McpHelpers.ResolveAsOf("1999-01-01T00:00:00Z", out var end)); + Assert.Equal(1999, end.Year); + } + + /* ── the two knobs together ── */ + + [Fact] + public void ValidateWindow_ReportsTheSpanBeforeTheAnchor_SoTheOldMessageIsUnchanged() + { + /* A caller who sent both wrong is told about hours_back first, exactly as before as_of existed. */ + var zeroHours = McpHelpers.ValidateWindow(0, "not a date", out _); + Assert.NotNull(zeroHours); + Assert.Contains("hours_back", zeroHours, StringComparison.Ordinal); + + var tooManyHours = McpHelpers.ValidateWindow(McpHelpers.MaxHoursBack + 1, "not a date", out _); + Assert.NotNull(tooManyHours); + Assert.Contains("exceeds maximum", tooManyHours, StringComparison.Ordinal); + } + + [Fact] + public void ValidateWindow_WithNoAnchor_IsExactlyTheOldValidateHoursBack() + { + foreach (var hours in new[] { -1, 0, 1, 24, McpHelpers.MaxHoursBack, McpHelpers.MaxHoursBack + 1 }) + { + Assert.Equal(McpHelpers.ValidateHoursBack(hours), McpHelpers.ValidateWindow(hours, null, out _)); + } + } + + [Fact] + public void TheWindow_IsHoursBackEndingAtTheAnchor() + { + Assert.Null(McpHelpers.ValidateWindow(3, "2026-08-18T15:30:00Z", out var end)); + + Assert.Equal(new DateTime(2026, 8, 18, 15, 30, 0, DateTimeKind.Utc), end); + Assert.Equal(new DateTime(2026, 8, 18, 12, 30, 0, DateTimeKind.Utc), end.AddHours(-3)); + } + + /* ── the two surfaces carry the same convention ── */ + + /// + /// A read reachable over MCP but not over /api/read/{name} is the failure mode the catalog exists + /// to prevent: the descriptor is the ONLY input truth for the web surface, so a parameter missing from it + /// is simply unreachable there, and an unknown query key is ignored rather than rejected — the panel + /// quietly gets the read's default window and nothing anywhere says so. + /// + [Fact] + public void EveryDispatchedReadWhoseToolTakesAnAnchor_AdvertisesItInTheCatalog() + { + var missing = DarlingWebEndpoints.BuildReadDispatch().Keys + .Where(ToolTakesAnAnchor) + .Where(name => !DarlingWebEndpoints.CatalogDescriptors[name].Params.Any(p => p.Name == "as_of")) + .OrderBy(n => n, StringComparer.Ordinal) + .ToArray(); + + Assert.True( + missing.Length == 0, + "these reads take as_of over MCP but do not advertise it in the catalog, so it cannot be sent " + + "over /api/read: " + string.Join(", ", missing)); + } + + /// The other direction: nothing advertises a parameter its tool would ignore. + [Fact] + public void NoCatalogEntry_AdvertisesAnAnchorItsToolDoesNotTake() + { + var extra = DarlingWebEndpoints.BuildReadDispatch().Keys + .Where(name => DarlingWebEndpoints.CatalogDescriptors[name].Params.Any(p => p.Name == "as_of")) + .Where(name => !ToolTakesAnAnchor(name)) + .OrderBy(n => n, StringComparer.Ordinal) + .ToArray(); + + Assert.True(extra.Length == 0, "catalog advertises as_of for reads whose tool ignores it: " + string.Join(", ", extra)); + } + + /// + /// The convention itself: a read that windows takes the anchor. The exclusions are named rather than + /// discovered, so removing one from the list is a deliberate act — and adding a new windowed read + /// without an anchor fails here rather than shipping half the convention. + /// + [Fact] + public void EveryWindowedRead_TakesTheAnchor_ExceptTheNamedExclusions() + { + /* The analysis family came OFF this list in #2506 — the anchor now reaches the engine, not just + the tool, and analyze_server refuses to persist when it is anchored rather than declining the + anchor. What is left are the two permanent kinds. get_pvs_stats and get_fleet_overview mix a + latest-snapshot measurement with a windowed one: anchoring only the windowed half would return a + result whose two halves describe different instants, which is worse than not offering it. + get_store_metrics windows in days over the store's own growth series. */ + var excluded = new[] + { + "get_pvs_stats", "get_fleet_overview", "get_store_metrics", + }; + + var unanchored = DarlingWebEndpoints.BuildReadDispatch().Keys + .Where(name => DarlingWebEndpoints.CatalogDescriptors[name].Params.Any(p => p.Name == "hours")) + .Where(name => !DarlingWebEndpoints.CatalogDescriptors[name].Params.Any(p => p.Name == "as_of")) + .Where(name => !excluded.Contains(name, StringComparer.Ordinal)) + .OrderBy(n => n, StringComparer.Ordinal) + .ToArray(); + + Assert.True( + unanchored.Length == 0, + "these reads window but cannot be anchored, and are not on the named exclusion list: " + + string.Join(", ", unanchored)); + } + + /// + /// Every tool that ADVERTISES as_of actually USES the instant it resolved. + /// + /// Why this exists. The convention's failure mode is not a compile error and not a wrong + /// number — it is a tool that takes the parameter, validates it, refuses a bad one correctly, and then + /// computes its window from anyway. The caller gets "now" labelled as their + /// window, the validation succeeding is what makes them believe it, and nothing in the result says + /// otherwise. That is #2495's own defect, one level up. + /// + /// It is not hypothetical: rebasing this work onto a moving dev produced eight of + /// them at once, across both SKUs, every one of which compiled and passed every other test here. Review + /// caught them; this catches the next eight. Both SKUs are scanned from ONE test because the category is + /// not per-SKU — the two copies drifted together. + /// + /// A source scan rather than a behavioural assertion, for the reason + /// gives about the JS it pins: the alternative is ~110 live round-trips + /// to prove a property that is visible in the text, and a property nobody checks is the one that breaks. + /// It reads the SHIPPED files, so it cannot drift into agreeing with a stale copy of itself. + /// + [Theory] + [InlineData("Darling/PerformanceMonitor.Darling.Service/Mcp")] + [InlineData("Lite/Mcp")] + public void EveryToolThatTakesTheAnchor_ActuallyUsesIt(string mcpDirectory) + { + var offenders = new List(); + var examined = 0; + var anchorable = AnchorableServiceMethods(); + var anchorableAnalysis = AnchorableAnalysisServiceMethods(); + + foreach (var file in RepoFilesIn(mcpDirectory)) + { + var source = File.ReadAllText(file).Replace("\r\n", "\n", StringComparison.Ordinal); + var marks = Regex.Matches(source, @"McpServerTool\(Name = ""([a-z_0-9]+)"""); + + for (var i = 0; i < marks.Count; i++) + { + var end = i + 1 < marks.Count ? marks[i + 1].Index : source.Length; + var declaration = source.IndexOf("public static", marks[i].Index, StringComparison.Ordinal); + if (declaration < 0 || declaration > end) + { + continue; + } + + var signatureEnd = CloseParenAfter(source, source.IndexOf('(', declaration)); + if (signatureEnd < 0) + { + continue; + } + + if (!source[declaration..signatureEnd].Contains("as_of", StringComparison.Ordinal)) + { + continue; + } + + examined++; + var body = source[signatureEnd..end]; + + /* The anchor reaches the window either as the resolved local, or by being forwarded whole + to a shared collector (the system_health family resolves it one level down). */ + var reaches = body.Contains("windowEnd", StringComparison.Ordinal) + || body.Contains("anchorEnd", StringComparison.Ordinal) + || Regex.IsMatch(body, @"\w+Async\([^;]*\bas_of\b"); + + /* + An anchored tool's ONLY source of "now" is the anchor it resolved, so it must not name + DateTime.UtcNow at all. Written as an absolute rather than as "does not ASSIGN UtcNow" + on evidence: the assignment form passed while get_file_io_trend was demonstrably broken, + because the body still contained the word windowEnd from its own `out var`. A check that + green-lights the defect it was written for is worse than none. + */ + var namesNow = body.Contains("DateTime.UtcNow", StringComparison.Ordinal); + + /* + The other half, and the one the reaches-check cannot see: Lite's tools resolve the anchor + and then hand it to LocalDataService. A read that CAN take asOfUtc and is not given it is + a tool that validated an anchor, believed itself anchored, and queried the present. Five + of those shipped past a green suite before this line existed. + */ + var dropped = Regex.Matches(body, @"dataService\.(\w+)\(") + .Where(m => anchorable.Contains(m.Groups[1].Value) && !CallPasses(body, m.Index, "asOfUtc")) + .Select(m => m.Groups[1].Value) + .Distinct() + .ToArray(); + + /* + #2506: the same category, one seam over. The analysis family's window is built inside + AnalysisService / DarlingAnalysisService rather than in the tool, so its tools resolve + the anchor and then hand it to a SERVICE — and a tool that resolves an anchor and calls + AnalyzeAsync without it satisfies every other check on this list (the body still names + windowEnd, from its own `out var`) while answering as of now. That is precisely the + shape #2495's review found eight of, so it gets its own arm rather than a comment. + */ + var droppedAnalysis = Regex.Matches(body, @"analysisService\.(\w+)\(") + .Where(m => anchorableAnalysis.Contains(m.Groups[1].Value) && !CallPasses(body, m.Index, "asOfUtc")) + .Select(m => m.Groups[1].Value) + .Distinct() + .ToArray(); + + if (!reaches || namesNow || dropped.Length > 0 || droppedAnalysis.Length > 0) + { + offenders.Add($"{Path.GetFileName(file)}:{marks[i].Groups[1].Value}" + + (reaches ? "" : " (never uses the resolved anchor)") + + (namesNow ? " (names DateTime.UtcNow)" : "") + + (dropped.Length > 0 ? $" (anchor not passed to {string.Join(", ", dropped)})" : "") + + (droppedAnalysis.Length > 0 ? $" (anchor not passed to analysisService.{string.Join(", ", droppedAnalysis)})" : "")); + } + } + } + + /* A scan that parsed nothing passes for free, which is the worst outcome a check like this can have: + it converts an open question into confidence. Both SKUs anchor dozens of reads, so a handful means + the signature extraction broke rather than that the surface shrank. */ + Assert.True( + examined >= 40, + $"only {examined} anchored tools were found under {mcpDirectory} — the scan is broken, not the surface"); + + Assert.True( + offenders.Count == 0, + "these tools advertise as_of and then answer as of NOW, which is worse than not offering it — " + + "the validation succeeds, so the caller believes the window moved: " + string.Join("; ", offenders)); + } + + /// + /// The LocalDataService reads that ACCEPT the anchor, read off the shipped source rather than + /// listed here — a transcribed list would go stale in exactly the direction that makes the pin pass. + /// + private static HashSet AnchorableServiceMethods() + { + var methods = new HashSet(StringComparer.Ordinal); + foreach (var file in RepoFilesIn("Lite/Services").Where(f => Path.GetFileName(f).StartsWith("LocalDataService", StringComparison.Ordinal))) + { + var source = File.ReadAllText(file).Replace("\r\n", "\n", StringComparison.Ordinal); + foreach (Match m in Regex.Matches(source, @"public (?:async )?[\w<>?,\[\]\(\) ]+ (\w+)\(([^{]*?)\)\s*(?:=>|\n\s*\{)")) + { + if (m.Groups[2].Value.Contains("asOfUtc", StringComparison.Ordinal)) + { + methods.Add(m.Groups[1].Value); + } + } + } + + Assert.True(methods.Count >= 40, $"only {methods.Count} anchor-taking service methods found — the scan is broken"); + return methods; + } + + /// + /// The ANALYSIS-service entry points that accept the anchor (#2506) — the same trick as + /// , read off both SKUs' shipped sources. + /// + /// Both SKUs are scanned into ONE set because the two services are deliberate twins with the + /// same member names, and the tools that call them are twins too. A per-SKU set would let one side + /// drop the anchor while the other kept it and still report nothing, which is the drift the + /// cross-SKU shape of this whole test exists to catch. + /// + private static HashSet AnchorableAnalysisServiceMethods() + { + var methods = new HashSet(StringComparer.Ordinal); + foreach (var directory in new[] { "Lite/Analysis", "Darling/PerformanceMonitor.Darling.Analysis" }) + { + foreach (var file in RepoFilesIn(directory).Where(f => Path.GetFileName(f).EndsWith("AnalysisService.cs", StringComparison.Ordinal))) + { + var source = File.ReadAllText(file).Replace("\r\n", "\n", StringComparison.Ordinal); + foreach (Match m in Regex.Matches(source, @"public (?:async )?[\w<>?,\[\]\(\) ]+ (\w+)\(([^{]*?)\)\s*(?:=>|\n\s*\{)")) + { + if (m.Groups[2].Value.Contains("asOfUtc", StringComparison.Ordinal)) + { + methods.Add(m.Groups[1].Value); + } + } + } + } + + /* AnalyzeAsync, CollectAndScoreFactsAsync, GetRecentFindingsAsync — the three the analysis tools + reach the window through. Asserted so a regex that stopped matching cannot empty the set and + make the arm above pass by finding nothing to check. */ + Assert.True(methods.Count >= 3, $"only {methods.Count} anchor-taking analysis-service methods found — the scan is broken"); + return methods; + } + + /// Whether one call expression passes an argument whose text contains . + private static bool CallPasses(string body, int callStart, string argument) + { + var open = body.IndexOf('(', callStart); + var close = CloseParenAfter(body, open); + return close > open && body[open..close].Contains(argument, StringComparison.Ordinal); + } + + private static int CloseParenAfter(string source, int open) + { + var depth = 0; + for (var i = open; i < source.Length; i++) + { + if (source[i] == '(') depth++; + else if (source[i] == ')' && --depth == 0) return i; + } + + return -1; + } + + private static string[] RepoFilesIn(string relativeDirectory, [CallerFilePath] string thisFile = "") + { + for (var dir = new DirectoryInfo(Path.GetDirectoryName(thisFile)!); dir is not null; dir = dir.Parent) + { + var candidate = Path.Combine(dir.FullName, relativeDirectory); + if (Directory.Exists(candidate)) + { + var files = Directory.GetFiles(candidate, "*.cs"); + Assert.NotEmpty(files); + return files; + } + } + + throw new DirectoryNotFoundException($"Could not locate {relativeDirectory} walking up from {thisFile}"); + } + + /// + /// Whether the [McpServerTool] method behind a read name takes the anchor. + /// + /// Both pins above are of the form "no read is in set A but not set B", so a lookup that quietly + /// returned false for a read it could not FIND would empty the failing set and make them pass for the + /// wrong reason. The method is therefore asserted to exist rather than defaulted, and a partially + /// loadable assembly is reported rather than silently enumerated as the types that happened to load. + /// + private static bool ToolTakesAnAnchor(string readName) + { + Type[] types; + try + { + types = typeof(DarlingMcpDataTools).Assembly.GetTypes(); + } + catch (System.Reflection.ReflectionTypeLoadException ex) + { + Assert.Fail( + "the service assembly did not fully load, so a missing tool would look like a passing pin: " + + string.Join("; ", ex.LoaderExceptions.Where(e => e is not null).Select(e => e!.Message).Distinct())); + return false; + } + + var method = types + .SelectMany(t => t.GetMethods(System.Reflection.BindingFlags.Public | System.Reflection.BindingFlags.Static)) + .FirstOrDefault(m => m.GetCustomAttributes(typeof(ModelContextProtocol.Server.McpServerToolAttribute), false) + .Cast() + .Any(a => a.Name == readName)); + + Assert.True(method is not null, $"no [McpServerTool] method is named '{readName}', which the read dispatch serves"); + return method!.GetParameters().Any(p => p.Name == "as_of"); + } +} + +/// +/// Gated (DARLING_TEST_PG) proof of the whole point: rows seeded in a PAST window come back when the read +/// is anchored there and do NOT come back on the default anchor — and widening hours_back instead is +/// a demonstrably different answer, not the same one with more rows. +/// +[Collection("live-postgres")] +public sealed class AsOfWindowAnchorLivePostgresTests +{ + private const string ServerName = "darling-asof-anchor-e2e"; + private static readonly int ServerId = ServerIdHelper.GetDeterministicHashCode(ServerName); + + /* The incident: 30 hours ago, i.e. OUTSIDE every default window on the surface (the widest default is + 24). "Now" carries a different wait type so the two windows are told apart by content, not by count. */ + private const string IncidentWait = "PAGEIOLATCH_SH"; + private const string RecentWait = "CXPACKET"; + + private static string? ConnectionString => Environment.GetEnvironmentVariable("DARLING_TEST_PG"); + + [Fact] + public async Task AnAnchoredRead_SeesThePastWindow_AndTheDefaultAnchorDoesNot() + { + var cs = ConnectionString; + Assert.SkipWhen(string.IsNullOrEmpty(cs), "Set DARLING_TEST_PG to a Postgres connection string to run the live as_of test."); + + var ct = TestContext.Current.CancellationToken; + using var connection = new NpgsqlConnection(cs); + await connection.OpenAsync(ct); + await PgMigrations.MigrateAsync(connection, ct); + await DeleteRowsAsync(connection, ct); + + await using var postgres = NpgsqlDataSource.Create(cs!); + + var bodySucceeded = false; + try + { + await RegisterServerAsync(connection, ct); + + var now = Truncate(DateTime.UtcNow); + var incident = now.AddHours(-30); + var anchor = incident.AddMinutes(30).ToString("o"); /* the incident sits inside a 4-hour window ending here */ + + await PlantWaitAsync(connection, incident, IncidentWait, 900_000L, ct); + await PlantWaitAsync(connection, incident.AddMinutes(5), IncidentWait, 800_000L, ct); + await PlantWaitAsync(connection, now.AddMinutes(-10), RecentWait, 1_000L, ct); + + /* 1. The default anchor is unchanged: the last 24 hours, which the incident is not in. */ + var live = await DarlingMcpDataTools.GetWaitStats(postgres, ServerName); + Assert.Contains(RecentWait, live, StringComparison.Ordinal); + Assert.DoesNotContain(IncidentWait, live, StringComparison.Ordinal); + + /* 2. Anchored at the incident, the SAME read returns the incident and nothing since. This is + the capability the issue says was unreachable. */ + var anchored = await DarlingMcpDataTools.GetWaitStats(postgres, ServerName, 4, 20, anchor); + Assert.Contains(IncidentWait, anchored, StringComparison.Ordinal); + Assert.DoesNotContain(RecentWait, anchored, StringComparison.Ordinal); + + /* 3. And widening hours_back is NOT the same question. A 48-hour window reaches the incident, + but it reaches everything since too — the aggregate now describes both, so the top-N and + the totals are answers to a question nobody asked. */ + var widened = await DarlingMcpDataTools.GetWaitStats(postgres, ServerName, 48); + Assert.Contains(IncidentWait, widened, StringComparison.Ordinal); + Assert.Contains(RecentWait, widened, StringComparison.Ordinal); + Assert.NotEqual(WaitTypesIn(anchored).Count, WaitTypesIn(widened).Count); + + /* 4. The collection log — the read the gap was noticed on — moves with the anchor too, and its + two-branch miss still tells a quiet window from a server that never collected. */ + await PlantCollectionLogAsync(connection, incident, ct); + var logAnchored = await DarlingMcpDataTools.GetCollectionLog(postgres, ServerName, 4, 200, anchor); + Assert.Contains("wait_stats", logAnchored, StringComparison.Ordinal); + Assert.Equal("empty", StatusOf(await DarlingMcpDataTools.GetCollectionLog(postgres, ServerName))); + + /* 5. Refusals reach the caller as the tool's own message, not as a silently-different answer. */ + var future = await DarlingMcpDataTools.GetWaitStats(postgres, ServerName, 4, 20, DateTime.UtcNow.AddDays(1).ToString("o")); + Assert.Contains("future", future, StringComparison.Ordinal); + Assert.StartsWith("Invalid as_of", await DarlingMcpDataTools.GetWaitStats(postgres, ServerName, 4, 20, "last tuesday"), StringComparison.Ordinal); + + /* 6. An anchor older than anything the store holds is the read's honest empty, not a refusal. */ + Assert.Equal( + "unavailable", + StatusOf(await DarlingMcpDataTools.GetWaitStats(postgres, ServerName, 4, 20, "1999-01-01T00:00:00Z"))); + + bodySucceeded = true; + } + finally + { + await LiveStoreCleanup.RunAsync(cs!, bodySucceeded, async (cleanup, cleanupCt) => + await DeleteRowsAsync(cleanup, cleanupCt)); + } + } + + private static System.Collections.Generic.List WaitTypesIn(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("waits").EnumerateArray() + .Select(w => w.GetProperty("wait_type").GetString()!) + .ToList(); + } + + private static string StatusOf(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("status").GetString()!; + } + + private static DateTime Truncate(DateTime value) => + new(value.Year, value.Month, value.Day, value.Hour, value.Minute, value.Second, DateTimeKind.Utc); + + private static DateTime Naive(DateTime value) => DateTime.SpecifyKind(value, DateTimeKind.Unspecified); + + private static async Task RegisterServerAsync(NpgsqlConnection connection, CancellationToken ct) + { + using var command = new NpgsqlCommand(@" +INSERT INTO servers (server_id, server_name, display_name, is_enabled, sql_major_version, created_date, modified_date) +VALUES ($1, $2, $3, TRUE, 15, $4, $4) +ON CONFLICT (server_id) DO UPDATE SET is_enabled = TRUE, sql_major_version = 15;", connection); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(Naive(DateTime.UtcNow)); + await command.ExecuteNonQueryAsync(ct); + } + + private static async Task PlantWaitAsync(NpgsqlConnection connection, DateTime at, string waitType, long deltaMs, CancellationToken ct) + { + using var command = new NpgsqlCommand(@" +INSERT INTO wait_stats (collection_id, collection_time, server_id, server_name, wait_type, delta_wait_time_ms, delta_signal_wait_time_ms, delta_waiting_tasks) +VALUES ($1,$2,$3,$4,$5,$6,$7,$8)", connection); + command.Parameters.AddWithValue(CollectionIdGenerator.Next()); + command.Parameters.AddWithValue(Naive(at)); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue(waitType); + command.Parameters.AddWithValue(deltaMs); + command.Parameters.AddWithValue(deltaMs / 10); + command.Parameters.AddWithValue(50L); + await command.ExecuteNonQueryAsync(ct); + } + + private static async Task PlantCollectionLogAsync(NpgsqlConnection connection, DateTime at, CancellationToken ct) + { + using var command = new NpgsqlCommand(@" +INSERT INTO collection_log (log_id, server_id, server_name, collector_name, collection_time, duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) +VALUES ($1, $2, $3, $4, $5, 120, 'SUCCESS', NULL, 7, 90, 30)", connection); + command.Parameters.AddWithValue(CollectionIdGenerator.Next()); + command.Parameters.AddWithValue(ServerId); + command.Parameters.AddWithValue(ServerName); + command.Parameters.AddWithValue("wait_stats"); + command.Parameters.AddWithValue(Naive(at)); + await command.ExecuteNonQueryAsync(ct); + } + + private static async Task DeleteRowsAsync(NpgsqlConnection connection, CancellationToken ct) + { + var sql = string.Join(" ", new[] { "wait_stats", "collection_log" } + .Select(table => $"DELETE FROM {table} WHERE server_id = {ServerId};")); + sql += $" DELETE FROM servers WHERE server_id = {ServerId};"; + using var command = new NpgsqlCommand(sql, connection); + await command.ExecuteNonQueryAsync(ct); + } +} diff --git a/Darling/Darling.Tests/CollectionLogFanoutRollupStoreTests.cs b/Darling/Darling.Tests/CollectionLogFanoutRollupStoreTests.cs new file mode 100644 index 000000000..9f54c7e26 --- /dev/null +++ b/Darling/Darling.Tests/CollectionLogFanoutRollupStoreTests.cs @@ -0,0 +1,323 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Darling.Service.Mcp; +using PerformanceMonitor.Darling.Storage; +using PerformanceMonitor.Darling.Viewer; +using Xunit; + +namespace Darling.Tests; + +/// +/// The V80 rung (#2472) — the per-database fan-out rollup on collection_log. No longer the top of the +/// ladder: V81 (#2515) took that, and the assertions that belong to whichever rung is newest moved with it to +/// . What stays here is everything that is true of THIS rung forever. +/// +/// What it is for. Five collectors run once per DATABASE and the run writes ONE row whose +/// duration_ms is the sum across all of them. So 8 databases at 10.1s and one at 62s beside seven at +/// 2.7s are both 80,900 ms, and they want opposite fixes — bounded parallelism for the first, a per-database +/// override for the second. #2468 could not be decided because nothing recorded which it was. +/// +/// Why not the tail statistics. #2460 gave every collector p95_duration_ms and +/// max_duration_ms, which is a real improvement and does not help here: +/// is the arithmetic. Both aggregate over RUNS, +/// and each of those runs is one blended row. +/// +public class CollectionLogFanoutRollupStoreTests +{ + [Fact] + public void TheRungIsRegisteredInADenseLadder() + { + var versions = PgMigrations.Scripts.Select(s => s.Version).ToList(); + + Assert.Equal("collection-log-fanout-rollup", PgMigrations.Scripts.Single(s => s.Version == 80).Name); + + /* Demoted at V81 (#2515): "80 is the maximum" was true of the LADDER while this was its top rung, + not of this rung, and leaving it here is how a demotion turns into a red build on the next one. + The density and ordering checks below stay — those are properties of the whole ladder that every + rung's test may assert. */ + Assert.Equal(StorageVersion.SchemaVersion, versions.Max()); + + Assert.Equal(versions.Distinct().OrderBy(v => v), versions); + var above = versions.Where(v => v > 45).OrderBy(v => v).ToList(); + Assert.Equal(Enumerable.Range(above[0], above.Count), above); + } + + /// + /// Three nullable columns and a view refresh. Nullable with no DEFAULT is the whole reason this rung is + /// safe to run on the largest table in the store: it is a catalog-only change in PostgreSQL and stays + /// instant on a compressed hypertable, where adding a column WITH a default is the shape TimescaleDB has + /// historically refused. + /// + [Fact] + public void TheRungAddsTheColumns_Idempotently_AndWithoutADefault() + { + var sql = PgMigrations.Scripts.Single(s => s.Version == 80).Sql; + + Assert.Contains("ALTER TABLE collect.collection_log", sql, StringComparison.Ordinal); + Assert.Contains("ADD COLUMN IF NOT EXISTS fanout_item_count integer", sql, StringComparison.Ordinal); + Assert.Contains("ADD COLUMN IF NOT EXISTS slowest_item text", sql, StringComparison.Ordinal); + Assert.Contains("ADD COLUMN IF NOT EXISTS slowest_item_ms integer", sql, StringComparison.Ordinal); + + /* No DEFAULT on any of the three, and no backfill. A row written before this rung genuinely does not + know its fan-out; NULL says so, where 0 would read as "fanned out over nothing". */ + Assert.DoesNotContain("DEFAULT", sql, StringComparison.Ordinal); + Assert.DoesNotContain("UPDATE ", sql, StringComparison.Ordinal); + } + + /// + /// The view refresh, which is the half of this rung that is easy to forget and impossible to notice. + /// Postgres FREEZES a view's SELECT * column list at CREATE, so without this the passthrough every + /// read goes through would keep serving eleven columns forever and the new ones would be invisible to a + /// store that was upgraded rather than created. V14 exists because that already happened once. + /// + [Fact] + public void TheRungRefreshesThePassthroughView() + { + var sql = PgMigrations.Scripts.Single(s => s.Version == 80).Sql; + + Assert.Contains( + "CREATE OR REPLACE VIEW collect.v_collection_log AS SELECT * FROM collect.collection_log;", + sql, StringComparison.Ordinal); + + /* And the reads really do go through the view rather than the table, which is what makes the refresh + load-bearing rather than tidy. */ + Assert.Contains("FROM v_collection_log", DarlingDataReader.CollectionHealthSql, StringComparison.Ordinal); + } + + [Fact] + public void TheProbeAsksForTheColumn_AndTheReaderStillFetchesItsOrdinal() + { + Assert.Contains( + "table_name = 'collection_log' AND column_name = 'slowest_item_ms'", + ViewerDataService.StoreSchemaProbeSql, StringComparison.Ordinal); + + /* Demoted at V81 (#2515): this used to read `mapParameters - 1`, which is the NEWEST sentinel's + ordinal, so once a rung was appended it silently started testing that rung's wiring instead of + this one's. Pinned at 55 — this rung's own ordinal — which cannot slide. The "and no more than + that" half is the top rung's to assert, and it moved to TempDbCeilingStoreTests with it. */ + Assert.Contains("reader.GetBoolean(55)", ReadViewerSource(), StringComparison.Ordinal); + } + + [Fact] + public void TheProbeMapsAStoreAtExactly80To80() + { + Assert.Equal(StorageVersion.SchemaVersion, ViewerDataService.RequiredStoreSchemaVersion); + + var method = typeof(ViewerDataService) + .GetMethod("MapProbedSchemaVersion", System.Reflection.BindingFlags.NonPublic | System.Reflection.BindingFlags.Static)!; + var arity = method.GetParameters().Length; + + /* 55 positional sentinels, then this rung's own, then FALSE for anything a later rung appends. The + leading count is FIXED at this rung's ordinal deliberately — see the note on V79's twin: deriving + it from arity reads identically while this is the top rung, then slides one place right per new + rung, and the assertion keeps passing while quietly testing a newer arm. V81 has since been + appended and this test needed no edit, which is the fixed count doing exactly what it is for. */ + var all = Enumerable.Repeat(true, 55).Cast().ToArray(); + object[] Args(bool ownFlag) => all + .Concat(new object[] { ownFlag }) + .Concat(Enumerable.Repeat((object)false, arity - 56)) + .ToArray(); + + Assert.Equal(80, (int)method.Invoke(null, Args(true))!); + Assert.Equal(79, (int)method.Invoke(null, Args(false))!); + } + + /// + /// THE ARITHMETIC — the reason this rung exists rather than a note saying the existing columns are + /// enough. Two fan-outs that any run-level statistic reports identically, separated by the rollup. + /// + /// Written as a test rather than as prose in the issue because "max and p95 cannot answer this" is + /// a claim that would otherwise have to be re-derived by hand every time someone proposes deleting these + /// three columns. + /// + [Fact] + public void TheTailStatistics_CannotSeparateTheTwoShapes() + { + /* #2472 writes these as 8 x 10.1s and 1 x 62s + 7 x 2.7s, which round to the same 80,900 ms but do + not actually sum to it (8 x 10,100 is 80,800). The figures below are the issue's shapes made to + balance EXACTLY, because "these two are indistinguishable" is the claim under test and it cannot + be demonstrated with two numbers that differ by 100. */ + var even = new[] { 10_100, 10_100, 10_100, 10_100, 10_100, 10_100, 10_100, 10_100 }; + var dominated = new[] { 61_900, 2_700, 2_700, 2_700, 2_700, 2_700, 2_700, 2_700 }; + + /* The same run, as collection_log records it: one row, one duration. Every run-level aggregate over + a window of such runs — AVG, MAX, PERCENTILE_DISC — therefore sees the same number for both. */ + Assert.Equal(80_800, even.Sum()); + Assert.Equal(80_800, dominated.Sum()); + Assert.Equal(even.Sum(), dominated.Sum()); + + var evenHealth = Health(even); + var dominatedHealth = Health(dominated); + + /* And now they are told apart, by the one number the rollup adds. */ + Assert.Equal(1.0, evenHealth.FanoutDominance!.Value, 2); + Assert.Equal(6.13, dominatedHealth.FanoutDominance!.Value, 2); + + /* The threshold the remedies turn on (#2468): near 1.0 the cost is the fan-out's WIDTH and bounded + parallelism is the lever; at 2.0 and above one database dominates and only a per-database override + or a stagger reaches it. */ + Assert.True(evenHealth.FanoutDominance < 2.0); + Assert.True(dominatedHealth.FanoutDominance >= 2.0); + } + + /// + /// A collector that does not fan out reports nothing rather than a zero. NULL is the honest answer for + /// "this run had no fan-out", which is not the same claim as "its fan-out was free" — the sentinel + /// discipline the PostgreSQL collectors already hold to. + /// + [Fact] + public void ACollectorThatDoesNotFanOut_ReportsNothingRatherThanZero() + { + var plain = new CollectorHealth { AvgDurationMs = 42 }; + + Assert.Null(plain.FanoutItems); + Assert.Null(plain.SlowestItem); + Assert.Null(plain.FanoutDominance); + + /* And a fan-out whose run somehow recorded no duration is null too: a ratio against nothing is a + wrong answer, not a smaller one. */ + var zeroRun = new CollectorHealth + { + FanoutItems = 8, + SlowestItem = "db", + SlowestItemMs = 10, + SlowestRunDurationMs = 0, + }; + Assert.Null(zeroRun.FanoutDominance); + } + + /// + /// The accumulator both SKUs and both fan-out shapes feed. Empty batches count: their read time is in the + /// blended total, so leaving them out would inflate the dominance of whichever database had rows — which + /// is precisely the wrong database to send an operator after. + /// + [Fact] + public void TheAccumulator_CountsEveryItemAndKeepsTheDearest() + { + var acc = new FanoutCostAccumulator(); + Assert.Null(acc.Result); + + acc.Observe("alpha", 2_700); + acc.Observe("bravo", 62_000); + acc.Observe("charlie", 0); + + var result = acc.Result!.Value; + Assert.Equal(3, result.ItemCount); + Assert.Equal("bravo", result.SlowestItem); + Assert.Equal(62_000, result.SlowestItemMs); + + /* A single free item still becomes the slowest one — the floor is -1, not 0, so a fan-out that + really did run never reports a null item. */ + var free = new FanoutCostAccumulator(); + free.Observe("only", 0); + Assert.Equal("only", free.Result!.Value.SlowestItem); + Assert.Equal(1, free.Result!.Value.ItemCount); + + /* Ties keep the first item seen, so the answer does not wobble between equally-priced databases from + one cycle to the next. An operator chasing a name needs it to be the same name tomorrow. + Adversarially named so an alphabetical tie-break would pick the other one. */ + var tie = new FanoutCostAccumulator(); + tie.Observe("zulu", 500); + tie.Observe("alpha", 500); + Assert.Equal("zulu", tie.Result!.Value.SlowestItem); + } + + /// + /// The tool that serves this names the enumeration-driven collectors it applies to, and the list is + /// DERIVED from the collector sources rather than typed out here — the first draft of this test hand-typed + /// five names, which is the same failure mode as the prose it was meant to guard. + /// + /// The issue said "query_store plus the two snapshot ones" and a comment in the runner said the + /// same; two more had joined since (query_store_health #2319, plan_correction #1952). And + /// the description's FIRST version then made the opposite mistake, claiming five collectors fan out full + /// stop — there are two mechanisms, and RunsPerDatabase puts eight more on a per-database + /// connection loop on Azure SQL DB plus pg_autovacuum_stats always. Both are asserted, because + /// the accumulator feeds from both and a caller told only half would look past a collector that has a + /// fanout block. + /// + /// Pinned on the DESCRIPTION because that is what a caller reads before deciding whether the block + /// applies to the collector in front of them. + /// + [Fact] + public void TheToolDescription_NamesEveryCollectorThatFansOut() + { + var description = ReadRepoFile(System.IO.Path.Combine( + "Darling", "PerformanceMonitor.Darling.Service", "Mcp", "DarlingMcpDataTools.cs")); + + /* Derived, not remembered: every collector source that overrides BuildEnumerationQuery drives the + enumeration fan-out, so a sixth one starting to enumerate fails here instead of quietly falling + out of the description. */ + var collectorsDir = System.IO.Path.Combine(RepoRoot(), "PerformanceMonitor.Collectors"); + var enumerating = System.IO.Directory + .EnumerateFiles(collectorsDir, "*Collector.cs", System.IO.SearchOption.TopDirectoryOnly) + .Where(f => System.IO.File.ReadAllText(f).Contains("override CollectorQuery? BuildEnumerationQuery", StringComparison.Ordinal)) + .Select(f => System.IO.Path.GetFileNameWithoutExtension(f)) + .ToList(); + + Assert.Equal(5, enumerating.Count); + + foreach (var typeName in enumerating) + { + /* CollectorName is the snake_case name the description writes; derive it from the type name so + this needs no second hand-typed mapping either. */ + var snake = string.Concat(typeName[..^"Collector".Length] + .Select((c, i) => char.IsUpper(c) && i > 0 ? "_" + char.ToLowerInvariant(c) : char.ToLowerInvariant(c).ToString())); + + Assert.Contains(snake, description, StringComparison.Ordinal); + } + + /* The OTHER mechanism, which the description's first version left out entirely. */ + Assert.Contains("per-database connection loop when the target is Azure SQL DB", description, StringComparison.Ordinal); + Assert.Contains("pg_autovacuum_stats", description, StringComparison.Ordinal); + + /* And both SKUs say the same thing, because the payload is field-identical and a caller should not + have to learn which product it is talking to. */ + var lite = ReadRepoFile(System.IO.Path.Combine("Lite", "Mcp", "McpHealthTools.cs")); + + /* The formula itself, with no surrounding punctuation: the first version of this assertion included + a word the description writes in backticks and matched neither file. */ + Assert.Contains("slowest_ms * items / run_ms", lite, StringComparison.Ordinal); + Assert.Contains("slowest_ms * items / run_ms", description, StringComparison.Ordinal); + } + + // ── helpers ────────────────────────────────────────────────────────────────────────────────────── + + /// One collector-health row as the read would build it from a fan-out of the given per-database + /// costs — the slowest item, the width, and the blended run duration collection_log actually stores. + private static CollectorHealth Health(int[] perDatabaseMs) => new() + { + FanoutItems = perDatabaseMs.Length, + SlowestItem = "worst", + SlowestItemMs = perDatabaseMs.Max(), + SlowestRunDurationMs = perDatabaseMs.Sum(), + }; + + private static string ReadViewerSource() => + ReadRepoFile(System.IO.Path.Combine("Darling", "PerformanceMonitor.Darling.Viewer", "ViewerDataService.cs")); + + private static string ReadRepoFile(string relative) => + System.IO.File.ReadAllText(System.IO.Path.Combine(RepoRoot(), relative)); + + private static string RepoRoot([System.Runtime.CompilerServices.CallerFilePath] string thisFile = "") + { + for (var dir = new System.IO.DirectoryInfo(System.IO.Path.GetDirectoryName(thisFile)!); dir is not null; dir = dir.Parent) + { + if (System.IO.Directory.Exists(System.IO.Path.Combine(dir.FullName, "PerformanceMonitor.Common"))) + { + return dir.FullName; + } + } + + throw new System.IO.DirectoryNotFoundException($"Could not locate the repo root walking up from {thisFile}"); + } + +} diff --git a/Darling/Darling.Tests/CollectorEngineCapabilityDerivationTests.cs b/Darling/Darling.Tests/CollectorEngineCapabilityDerivationTests.cs new file mode 100644 index 000000000..6ecf80bfd --- /dev/null +++ b/Darling/Darling.Tests/CollectorEngineCapabilityDerivationTests.cs @@ -0,0 +1,1241 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Reflection; +using System.Reflection.Emit; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Darling.Tests; + +/// +/// #2518: proof that +/// FOLLOWS a gate that moves, rather than merely agreeing with the gates as they ship today. +/// +/// Why the existing tests are not this. Every assertion in +/// CollectorEngineCapabilityTests is a statement about the shipped collectors — system_health is a +/// gap on Azure SQL Database, job_history is not, box editions have none. All of them are true, none of them +/// can distinguish a derivation from a hard-coded set of gaps that happens to match today's gates, and a +/// hard-coded set is exactly the failure the derivation exists to prevent. Nothing in the suite moves a gate +/// and watches the answer move with it, so nothing in the suite would notice if the answer stopped being +/// derived at all. +/// +/// Why a synthetic collector rather than a real one. The obvious demonstration is to point at a +/// collector whose gate changed — and #2512/#2516 supplied one, TempDbStatsCollector, while #2511 was +/// still open. Pinning the derivation to the one gate that is in flight is pinning the wrong thing: the pin +/// then fails whenever that lane lands, says "update the pin", and gets updated rather than read. A gate this +/// test owns moves on demand, moves in both directions, and drags no real collector into the assertion. The +/// collectors as shipped stay the subject of the OTHER file, which is where a claim about them belongs. +/// +/// Both directions, on ONE object. The gate is mutated between calls on the same instance, so +/// the two answers differ with literally nothing else changed — not the collector name, not the object +/// identity, not the edition. A pair of differently-constructed fakes would leave "it keyed off something +/// about the instance" open; this does not. +/// +public sealed class CollectorEngineCapabilityMovingGateTests +{ + private const int Enterprise = 3; + private const int AzureSqlDb = CollectorEngineCapability.AzureSqlDatabaseEngineEdition; + private const int AzureMi = CollectorEngineCapability.AzureManagedInstanceEngineEdition; + + /// + /// Every SERVERPROPERTY('EngineEdition') value Microsoft documents, as a contiguous range rather + /// than a curated list — this is the DOMAIN of the question, so covering it densely costs nothing and a + /// value that gets defined later is already included. + /// + private static readonly int[] AllEngineEditions = Enumerable.Range(1, 12).ToArray(); + + /// + /// A collector definition whose gate the test owns and can move between calls. Everything else is the + /// bare minimum demands — none of it participates in the capability + /// question, and giving it plausible-looking values would only invite the reader to wonder whether it + /// does. + /// The name is deliberately not a catalog name. That is load-bearing: + /// uses the + /// same object through both overloads and gets different answers, which is what shows the definition + /// overload is evaluating the gate rather than resolving a name. + /// + private sealed class SyntheticCollector : ICollectorSchemaInfo + { + /// The gate under test. Settable, not init, so one instance can be moved. + public required Func Gate { get; set; } + + /// The engine half of the dispatch gate, and the discriminator the KIND axis derives from + /// (#2530). Settable for the same reason is: the engine-kind assertions move it + /// between calls on ONE instance, so the two answers differ with nothing else changed. + public CollectorTargetEngine TargetEngine { get; set; } = CollectorTargetEngine.SqlServer; + + public string Name => "synthetic_gate_2518"; + + public string TargetTable => "synthetic_gate_2518"; + + public bool IncludesCollectionId => true; + + public string PrefixIdColumnName => "collection_id"; + + public string PrefixTimeColumnName => "collection_time"; + + public IReadOnlyList PayloadColumns => Array.Empty(); + + public bool YieldsOnLockTimeout => false; + + public IReadOnlyList StateKeys => Array.Empty(); + + public bool AppliesTo(CollectorTargetInfo target) => Gate(target); + } + + /// + /// The assertion the issue was filed for. One collector, one object, three gates: closed against Azure + /// SQL Database, open everywhere, then closed against everything BUT Azure SQL Database. The capability + /// answer tracks all three. + /// + /// The third position matters as much as the first two. A "derivation" that had quietly become + /// "Azure SQL Database is where the gaps are" would pass the first two moves and fail here, because here + /// the gap is on Enterprise and Managed Instance and Azure SQL Database is the edition that collects. + /// + [Fact] + public void TheAnswerMoves_WhenTheGateMoves_InBothDirections() + { + var collector = new SyntheticCollector { Gate = target => !target.IsAzureSqlDb }; + + /* Gate closed against Azure SQL Database: a permanent gap there, and nowhere else. */ + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureSqlDb)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, Enterprise)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureMi)); + + /* The gate OPENS — the #2516 direction, the one a hand-kept list fails silently in, because a list + goes on claiming a gap that the gate stopped producing and nothing says so. */ + collector.Gate = _ => true; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureSqlDb)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, Enterprise)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureMi)); + + /* And closes again the OTHER way round, so the answer cannot be "Azure SQL Database is special". */ + collector.Gate = target => target.IsAzureSqlDb; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureSqlDb)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, Enterprise)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureMi)); + } + + /// + /// The by-definition overload asks the GATE; the by-name overload asks the CATALOG and then the gate it + /// finds. Handed a definition the catalog has never heard of, they must therefore disagree — the + /// definition form reports the gap its gate produces, the name form makes no claim at all because there + /// is no gate behind that name to derive one from. + /// + /// This is the assertion that stops the new overload being a rename. If it delegated back to the + /// name, or resolved the definition through the catalog, both calls below would return true and + /// the moving-gate test above would be exercising nothing. + /// + [Fact] + public void TheByDefinitionOverload_AsksTheGate_WhereTheByNameOverloadCanOnlyAskTheCatalog() + { + var collector = new SyntheticCollector { Gate = _ => false }; + + Assert.Null(CollectorCatalog.Find(collector.Name)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureSqlDb)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector.Name, AzureSqlDb)); + } + + /// + /// For every collector the catalog DOES know, the two overloads are the same function. The by-name form + /// is the one both MCP trees call, so a divergence would mean the tests below prove a property of code + /// nothing in production reaches. + /// + [Fact] + public void TheTwoOverloads_AgreeOnEveryShippedCollectorAtEveryEdition() + { + Assert.NotEmpty(CollectorCatalog.All); + + foreach (var definition in CollectorCatalog.All) + { + foreach (var edition in AllEngineEditions) + { + Assert.Equal( + CollectorEngineCapability.IsCollectedOnEngineEdition(definition, edition), + CollectorEngineCapability.IsCollectedOnEngineEdition(definition.Name, edition)); + } + } + } + + /// + /// A gate that closes on a FIXABLE fact is not an engine gap, demonstrated on gates this test controls + /// rather than on job_history's. Same claim as AFixableGate_IsNotReportedAsAnEngineGap, + /// but it stays true when the Agent collectors' gates change, and it covers the gate shapes no shipped + /// collector happens to have today. + /// + /// The last case is the one worth reading: a gate that reads BOTH a fixable fact and the edition + /// is still an engine gap on that edition, because no combination of the fixable facts rescues it. That + /// is the distinction the whole helper exists to draw, and neither half of it alone would show it. + /// + [Fact] + public void AGateOnAFixableFact_IsNotAnEngineGap_ButTheSameGateAndAnEditionStillIs() + { + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.HasMsdbAccess }, Enterprise)); + + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => !target.IsAwsRds }, Enterprise)); + + /* Both fixable facts at once: the sweep has to carry the COMBINATION, not just each value somewhere. */ + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.HasMsdbAccess && target.IsAwsRds }, Enterprise)); + + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => !target.HasMsdbAccess && !target.IsAwsRds }, Enterprise)); + + /* Fixable AND edition-bound: permanent on Azure SQL Database, fixable everywhere else. */ + var mixed = new SyntheticCollector { Gate = target => target.HasMsdbAccess && !target.IsAzureSqlDb }; + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(mixed, AzureSqlDb)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(mixed, Enterprise)); + } + + /// + /// A version gate is answered across the real majors, not by whichever single representative value the + /// sweep happened to pick. Every major the sweep carries must be reachable ON ITS OWN, or a gate written + /// as a RANGE — supported on 15 and 16, dropped on 17 — would be reported as a permanent engine gap on + /// hardware it runs on perfectly well. + /// The floor and ceiling cases bracket it: a gate that needs the newest engine, and one that only + /// tolerates the oldest, are both answered "collected". + /// + [Fact] + public void AVersionGate_IsAnsweredAcrossTheRealMajors_NotByOneRepresentativeValue() + { + foreach (var major in new[] { 11, 12, 13, 14, 15, 16, 17 }) + { + var only = major; + Assert.True( + CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.SqlMajorVersion == only }, Enterprise), + $"a gate that passes only on SQL major {only} was reported as a permanent engine gap, so the " + + "sweep does not carry that major — every real major has to be reachable on its own"); + } + + /* A range, the shape the sweep's own doc comment says it exists for. */ + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.SqlMajorVersion is 15 or 16 }, Enterprise)); + + /* A floor above every real major, and the "unknown, assume newest" value. */ + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.SqlMajorVersion >= 99 }, Enterprise)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + new SyntheticCollector { Gate = target => target.SqlMajorVersion == 0 }, Enterprise)); + } + + /// + /// The non-vacuity floor, both ways and on every edition: a gate shut everywhere is a gap everywhere, a + /// gate open everywhere is a gap nowhere. If either of these ever stopped holding, every other assertion + /// in this file would be passing for a reason unrelated to the gate. + /// Unknown (0) is excluded from the loop and asserted separately, because it is the one edition + /// where a shut gate must STILL make no claim — "we have not probed this server" is not "this will never + /// work", and that branch sits above the sweep where no gate can reach it. + /// + [Fact] + public void AGateShutEverywhere_IsAGapOnEveryEdition_AndAnOpenGateOnNone() + { + var shut = new SyntheticCollector { Gate = _ => false }; + var open = new SyntheticCollector { Gate = _ => true }; + + foreach (var edition in AllEngineEditions) + { + Assert.False( + CollectorEngineCapability.IsCollectedOnEngineEdition(shut, edition), + $"a gate that refuses every target was not reported as a gap on engine edition {edition}"); + Assert.True( + CollectorEngineCapability.IsCollectedOnEngineEdition(open, edition), + $"a gate that accepts every target was reported as a gap on engine edition {edition}"); + } + + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition( + shut, CollectorEngineCapability.UnknownEngineEdition)); + } + + /// + /// A definition written in another engine's dialect is never a gap on a SQL Server edition, however its + /// own gate is set — the question does not apply to it. + /// The Gate = _ => true case is the one that isolates the short-circuit. The sweep asks + /// , whose engine half + /// rejects every SQL Server shape, so WITHOUT the short-circuit an always-open PostgreSQL collector comes + /// back as a permanent gap on all nine editions — a confident claim about an engine it was never meant to + /// run on. With _ => false the two paths agree by accident and prove nothing. + /// + [Fact] + public void AForeignEngineDefinition_IsNeverAGap_HoweverItsOwnGateIsSet() + { + foreach (var gate in new Func[] { _ => true, _ => false }) + { + var postgres = new SyntheticCollector + { + TargetEngine = CollectorTargetEngine.PostgreSql, + Gate = gate, + }; + + foreach (var edition in AllEngineEditions) + { + Assert.True( + CollectorEngineCapability.IsCollectedOnEngineEdition(postgres, edition), + $"a PostgreSQL definition was reported as a permanent gap on engine edition {edition}"); + } + } + } + + /* ───────── the engine-KIND axis (#2530) ───────── */ + + /// + /// The kind axis is DERIVED too, and this is what says so: one object, one engine kind, the + /// definition's own moved between calls, and the answer + /// moves with it. Nothing else changes — not the name, not the object identity, not the token. + /// + /// Without this the kind axis could have been "the eight pg_ collectors are the PostgreSQL ones", + /// which is true today, would pass every assertion anyone would write about the shipped catalog, and + /// would go stale in the direction that keeps passing — exactly the failure #2518 exists to prevent + /// one axis over. + /// + [Fact] + public void TheKindAnswerMoves_WhenTheDefinitionsEngineMoves_InBothDirections() + { + var collector = new SyntheticCollector { Gate = _ => true, TargetEngine = CollectorTargetEngine.SqlServer }; + + /* A SQL Server dialect asked about a PostgreSQL target: a permanent gap, on both PG tokens. */ + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.Postgres)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.AuroraPostgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.SqlServer)); + + /* The SAME object, one field moved. Every answer inverts. */ + collector.TargetEngine = CollectorTargetEngine.PostgreSql; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.Postgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.AuroraPostgres)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.SqlServer)); + + /* And back, so the answer cannot be "PostgreSQL is the special one". */ + collector.TargetEngine = CollectorTargetEngine.SqlServer; + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.Postgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.SqlServer)); + } + + /// + /// A gate SHUT everywhere is a kind gap on both PostgreSQL tokens; a gate on a fact the sweep varies is + /// a gap on neither. This is the assertion #2532 replaced: until then the kind axis asked only the + /// ENGINE half of the dispatch gate, and a definition's own AppliesTo could not make it claim + /// anything at all. + /// + /// Why that narrowing was right then and wrong now. The discipline it protected — never + /// report a FIXABLE gate as permanent — is unchanged and is the whole subject of this file. What it + /// lacked was a sweep: with no way to separate "excluded on every target of this kind" from "excluded on + /// the one target shape somebody happened to construct", the only safe answer was to decline. Now + /// varies the PostgreSQL version floors and + /// the recovery state and fixes only , so the fixable gates + /// answer TRUE on their own merits rather than by the axis refusing to look — which is what the second + /// half of this test is for, and what + /// EveryFactAPostgresGateReads_IsVariedBySweepOrFixedByKind keeps true as gates are added. + /// + [Fact] + public void AShutAppliesToGate_IsAnEngineKindGap_ButAFixableOneIsNot() + { + var postgres = new SyntheticCollector { Gate = _ => false, TargetEngine = CollectorTargetEngine.PostgreSql }; + + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.Postgres)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.AuroraPostgres)); + + /* ...and the engine half still speaks, so a PostgreSQL definition is a gap on SQL Server for the + dialect reason rather than for this one. */ + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.SqlServer)); + + /* The half that keeps the narrowing's point: a gate reading a fact an operator can move is not a + permanent gap on EITHER token. An upgrade crosses the floor; a writer connection leaves recovery. + A sweep that stopped varying either of these would report both of them as "never will". */ + foreach (var fixableGate in new Func[] + { + target => target.PostgresMajorVersion >= 16, + target => target.PostgresVersionNum >= 170005, + target => !target.IsInRecovery, + target => target.IsInRecovery, + }) + { + postgres.Gate = fixableGate; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.Postgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.AuroraPostgres)); + } + + /* And the one fact the KIND fixes: Aurora-ness. Nothing an operator does turns a stock PostgreSQL + server into an Aurora one, so this is the only PostgreSQL gate shape that produces a permanent + gap — and it produces it on exactly one of the two tokens, in each direction. */ + postgres.Gate = target => target.IsAurora; + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.Postgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.AuroraPostgres)); + + postgres.Gate = target => !target.IsAurora; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.Postgres)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(postgres, MonitoredEngineKind.AuroraPostgres)); + } + + /// + /// The SQL Server side of the same question, and the reason the two axes do not answer each other's: + /// a collector gated off on Azure SQL Database is NOT a kind gap on sqlserver, because it runs on + /// every other SQL Server there is. The edition axis is what says so, with the edition named. + /// + /// This is what varying the two Azure + /// flags buys, and it is the property that would break silently if the SQL Server arm of that sweep ever + /// fixed them the way does: every + /// Azure-gated collector would become a permanent gap on every SQL Server, and the message would stop + /// naming the edition that is actually responsible. + /// + [Fact] + public void AnAzureOnlyGate_IsNotAKindGapOnSqlServer_ButIsStillAnEditionGap() + { + var collector = new SyntheticCollector + { + Gate = target => !target.IsAzureSqlDb, + TargetEngine = CollectorTargetEngine.SqlServer, + }; + + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.SqlServer)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, AzureSqlDb)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineEdition(collector, Enterprise)); + + /* A gate shut on every SQL Server shape IS a kind gap, so the assertion above is not vacuous. */ + collector.Gate = _ => false; + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(collector, MonitoredEngineKind.SqlServer)); + } + + /// + /// An absent or unrecognised token makes NO claim, whatever the definition's engine. This is the + /// guarantee #2530 was told to keep: the distinction being added is "known to be PostgreSQL", never + /// "not known to be SQL Server". A store one rung behind, and a server that has not connected since the + /// rung landed, both land here — and both must keep the miss vocabulary they had. + /// + [Theory] + [InlineData(null)] + [InlineData("")] + [InlineData("something-a-newer-build-writes")] + public void AnUnknownEngineKind_MakesNoClaim_ForEitherDialect(string? engineKind) + { + foreach (var engine in new[] { CollectorTargetEngine.SqlServer, CollectorTargetEngine.PostgreSql }) + { + var collector = new SyntheticCollector { Gate = _ => true, TargetEngine = engine }; + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(collector, engineKind)); + } + } + + /// + /// The by-name and by-definition kind overloads are the same function for every collector the catalog + /// knows, and disagree exactly where they must: on a definition the catalog has never heard of, the + /// name form has no gate to ask and makes no claim. + /// + [Fact] + public void TheTwoKindOverloads_Agree_ExceptWhereTheNameCannotBeResolved() + { + Assert.NotEmpty(CollectorCatalog.All); + + foreach (var definition in CollectorCatalog.All) + { + foreach (var kind in MonitoredEngineKind.All) + { + Assert.Equal( + CollectorEngineCapability.IsCollectedOnEngineKind(definition, kind), + CollectorEngineCapability.IsCollectedOnEngineKind(definition.Name, kind)); + } + } + + var unknown = new SyntheticCollector { Gate = _ => true, TargetEngine = CollectorTargetEngine.SqlServer }; + Assert.Null(CollectorCatalog.Find(unknown.Name)); + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(unknown, MonitoredEngineKind.Postgres)); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(unknown.Name, MonitoredEngineKind.Postgres)); + } + + /// + /// The shipped catalog, as the axis sees it: every SQL Server collector is a gap on a PostgreSQL + /// target and none on a SQL Server one, and the PostgreSQL collectors are the mirror image. + /// Both halves are counted from the catalog rather than listed, so the numbers track it. + /// + /// This is the vacuity check for everything above: if the derivation quietly answered TRUE for + /// everything, every "no claim" assertion in this file would still pass and this would not. + /// + [Fact] + public void TheShippedCatalogSplitsCleanlyOnTheKindAxis() + { + var sqlServer = CollectorCatalog.All.Where(c => c.TargetEngine == CollectorTargetEngine.SqlServer).ToArray(); + var postgres = CollectorCatalog.All.Where(c => c.TargetEngine == CollectorTargetEngine.PostgreSql).ToArray(); + + Assert.True(sqlServer.Length >= 30, $"only {sqlServer.Length} SQL Server definitions — the catalog walk is broken"); + Assert.True(postgres.Length >= 8, $"only {postgres.Length} PostgreSQL definitions — the catalog walk is broken"); + + foreach (var pgToken in new[] { MonitoredEngineKind.Postgres, MonitoredEngineKind.AuroraPostgres }) + { + Assert.All(sqlServer, c => Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(c, pgToken))); + } + + Assert.All(sqlServer, c => Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(c, MonitoredEngineKind.SqlServer))); + Assert.All(postgres, c => Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(c, MonitoredEngineKind.SqlServer))); + + /* Aurora is a strict superset of the surfaces the PostgreSQL collectors read, so every one of them + applies there. This is the half that would break if a new collector were written against something + Aurora removes rather than adds. */ + Assert.All(postgres, c => Assert.True( + CollectorEngineCapability.IsCollectedOnEngineKind(c, MonitoredEngineKind.AuroraPostgres), + $"{c.Name} is reported as a permanent gap on Aurora PostgreSQL")); + + /* Stock PostgreSQL is where the Aurora-only surfaces become a real gap (#2532), so the two tokens + genuinely differ — counted from the catalog rather than listed, and asserted as a PROPER subset so + neither "everything is a gap" nor "nothing is" passes. */ + var stockGaps = postgres + .Where(c => !CollectorEngineCapability.IsCollectedOnEngineKind(c, MonitoredEngineKind.Postgres)) + .Select(c => c.Name) + .ToArray(); + + Assert.NotEmpty(stockGaps); + Assert.True(stockGaps.Length < postgres.Length, + "every PostgreSQL collector is reported as a permanent gap on stock PostgreSQL, which would mean " + + "the sweep stopped producing shapes rather than that the gates changed: " + string.Join(", ", stockGaps)); + } + + /// + /// The measurement #2532 was filed on, named: aurora_stat_system_waits() is an Aurora built-in + /// that core PostgreSQL has no equivalent of in any version, so pg_wait_stats can never run on a + /// stock PostgreSQL target — and runs perfectly well on the Aurora one, which is the whole fleet today. + /// + /// The three collectors beside it are the control. pg_io_stats carries a PG16 floor and + /// pg_autovacuum_stats a writer-only gate; both are FIXABLE, both keep the unavailable + /// vocabulary that sends an operator to look, and a sweep that stopped varying either fact would report + /// them here as "never will". pg_blocking has no gate at all. + /// + [Fact] + public void AuroraOnlySurfaces_ArePermanentGapsOnStockPostgres_AndTheFixableGatesAreNot() + { + /* pg_statement_stats was here until #2625 gave it a vanilla pg_stat_statements path. It is now a + control, below, for the opposite reason: a collector whose SOURCE varies by flavor is not a + capability gap, and if it ever reads as one again the gate has come back. */ + foreach (var auroraOnly in new[] { "pg_wait_stats" }) + { + Assert.False(CollectorEngineCapability.IsCollectedOnEngineKind(auroraOnly, MonitoredEngineKind.Postgres), auroraOnly); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(auroraOnly, MonitoredEngineKind.AuroraPostgres), auroraOnly); + } + + foreach (var fixable in new[] { "pg_io_stats", "pg_autovacuum_stats", "pg_blocking", "pg_statement_stats" }) + { + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(fixable, MonitoredEngineKind.Postgres), fixable); + Assert.True(CollectorEngineCapability.IsCollectedOnEngineKind(fixable, MonitoredEngineKind.AuroraPostgres), fixable); + } + } + + /// + /// The two shapes of kind gap say different things, because they ARE different things: a foreign dialect + /// is stopped by the dispatch gate's engine half before the collector's own gate is consulted, while + /// pg_wait_stats on stock PostgreSQL is that collector's own AppliesTo. One sentence for + /// both would tell a PostgreSQL operator that a PostgreSQL collector is not written for PostgreSQL, + /// which is the sort of wrong that costs a reader their trust in the rest of the message. + /// + [Fact] + public void TheKindMessage_NamesTheRealReason_DialectOrTheCollectorsOwnGate() + { + var dialect = CollectorEngineCapability.NotCollectedMessage( + "aurora-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.AuroraPostgres, "wait_stats"); + + Assert.NotNull(dialect); + Assert.Contains("aurora-01 runs Aurora PostgreSQL.", dialect, StringComparison.Ordinal); + Assert.Contains("is written against SQL Server", dialect, StringComparison.Ordinal); + Assert.Contains("the dispatch gate's engine half never sends it at another engine", dialect, StringComparison.Ordinal); + + var ownGate = CollectorEngineCapability.NotCollectedMessage( + "pg-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.Postgres, "pg_wait_stats"); + + Assert.NotNull(ownGate); + Assert.Contains("pg-01 runs PostgreSQL.", ownGate, StringComparison.Ordinal); + Assert.Contains("its own AppliesTo gate excludes it", ownGate, StringComparison.Ordinal); + Assert.Contains("the aurora_stat_system_waits() cumulative wait counters", ownGate, StringComparison.Ordinal); + Assert.Contains("and never will.", ownGate, StringComparison.Ordinal); + + /* The dialect sentence must NOT appear on the own-gate message: it is the specific falsehood this + split exists to prevent. */ + Assert.DoesNotContain("is written against PostgreSQL", ownGate, StringComparison.Ordinal); + Assert.DoesNotContain("dispatch gate's engine half", ownGate, StringComparison.Ordinal); + Assert.DoesNotContain("check that collection is running", ownGate, StringComparison.OrdinalIgnoreCase); + + /* No EngineEdition claim about a PostgreSQL server, on either shape. */ + Assert.DoesNotContain("EngineEdition", dialect, StringComparison.Ordinal); + Assert.DoesNotContain("EngineEdition", ownGate, StringComparison.Ordinal); + + /* And the same read on the engine that DOES have the surface says nothing at all, so the read keeps + its own miss vocabulary. */ + Assert.Null(CollectorEngineCapability.NotCollectedMessage( + "aurora-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.AuroraPostgres, "pg_wait_stats")); + } + + /// + /// The two axes compose in the right ORDER inside the message, which is the one thing neither axis can + /// assert on its own. A PostgreSQL target's engine edition is 0 — "no claim" — so asking edition + /// first would return null for every PostgreSQL target and the read would fall back to + /// unavailable: the wrong-cause message #2530 was filed about, still there after the fix. + /// + [Fact] + public void TheKindAxisIsAskedFirst_SoAPostgresTargetsZeroEditionCannotSilenceIt() + { + /* Edition 0 AND a known PostgreSQL kind — exactly what the store holds for an Aurora target. */ + var message = CollectorEngineCapability.NotCollectedMessage( + "aurora-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.AuroraPostgres, "system_health_events"); + + Assert.NotNull(message); + Assert.Contains("Aurora PostgreSQL", message, StringComparison.Ordinal); + Assert.Contains("system_health_events", message, StringComparison.Ordinal); + Assert.Contains("is written against SQL Server", message, StringComparison.Ordinal); + Assert.Contains("and never will.", message, StringComparison.Ordinal); + + /* The edition axis must not have spoken: its sentence names an EngineEdition number, and + "EngineEdition 0" would be a claim about a property this server does not have. */ + Assert.DoesNotContain("EngineEdition", message, StringComparison.Ordinal); + Assert.DoesNotContain("check that collection is running", message, StringComparison.OrdinalIgnoreCase); + + /* Same server, same edition, kind absent — the pre-rung row — keeps its old silence. */ + Assert.Null(CollectorEngineCapability.NotCollectedMessage( + "aurora-01", CollectorEngineCapability.UnknownEngineEdition, engineKind: null, "system_health_events")); + + /* And a PostgreSQL read asked about a KNOWN SQL Server target is the mirror image, which is what + shows the axis is about engines rather than about PostgreSQL being second-class. */ + var mirrored = CollectorEngineCapability.NotCollectedMessage( + "box-01", 3, MonitoredEngineKind.SqlServer, "pg_wait_stats"); + Assert.NotNull(mirrored); + Assert.Contains("box-01 runs SQL Server.", mirrored, StringComparison.Ordinal); + Assert.Contains("is written against PostgreSQL", mirrored, StringComparison.Ordinal); + + /* A SQL Server read about a SQL Server target whose collector runs everywhere: untouched. */ + Assert.Null(CollectorEngineCapability.NotCollectedMessage( + "box-01", 3, MonitoredEngineKind.SqlServer, "wait_stats")); + } +} + +/// +/// #2518, the converse of the sweep-dimension pin: nothing asserted that a target fact the gates READ is a +/// fact the sweep VARIES. +/// +/// The hole. TheSweep_VariesEveryDimensionTheGatesRead walks a hand-written list of +/// dimensions and checks the sweep spans each. Add a property to , write a +/// gate on it, and that test still passes — the new fact sits at its CLR default across all 36 swept shapes, +/// every shape fails the new gate, and the derivation reports a permanent engine gap on every edition for a +/// collector that runs fine. Nothing is red. The message is confident, specific and wrong, which is the exact +/// defect #2511 was filed to remove. +/// +/// Why this is not a list under another name. Both halves are read out of the shipped artifacts. +/// The facts the gates read come from decoding the IL of every SQL Server definition's AppliesTo and +/// recording which getters it calls — so a gate cannot be written without +/// this seeing it, and a gate that stops reading a fact stops being evidence for it. The facts the sweep +/// varies come from running and observing +/// the values it produces — so extending the sweep is what satisfies the check, and there is no list here to +/// add a name to instead. Neither half can be brought into line by editing this file. +/// +/// What it does not catch, stated rather than assumed. A field ADDED to +/// that no SQL Server gate reads is not flagged, because it is not yet a +/// defect: the derivation only consults the gates, so an unread fact cannot make it over-claim. +/// PostgresVersionNum is the shipped proof that this state exists and is fine — it is populated by the +/// connector, read by no definition at all, and harmless. The guard fires at the moment the harm becomes +/// possible, which is the moment a gate reads the fact. +/// +public sealed class CollectorEngineCapabilitySweepDimensionTests +{ + /// The same contiguous edition domain the moving-gate tests use. + private static readonly int[] AllEngineEditions = Enumerable.Range(1, 12).ToArray(); + + /// Every public instance property of , from the type itself — + /// so a new one is in scope the moment it compiles. + private static readonly IReadOnlyList TargetFacts = + typeof(CollectorTargetInfo) + .GetProperties(BindingFlags.Public | BindingFlags.Instance) + .OrderBy(property => property.Name, StringComparer.Ordinal) + .ToArray(); + + /// Getter name → the property it reads, for mapping a decoded callvirt back to a fact. + private static readonly IReadOnlyDictionary FactByGetterName = + TargetFacts + .Where(property => property.GetGetMethod() is not null) + .ToDictionary(property => property.GetGetMethod()!.Name, StringComparer.Ordinal); + + /// Fact name → the property, for going back from a decoded read to its sweep role. + private static readonly IReadOnlyDictionary FactByName = + TargetFacts.ToDictionary(property => property.Name, StringComparer.Ordinal); + + /// + /// The two PostgreSQL tokens the store can record, which are the KIND axis's domain the way + /// is the edition axis's (#2532). Taken from the vocabulary rather than + /// typed out, so a third PostgreSQL flavour is in scope the moment it is added — and filtered by dialect + /// rather than by name, so it stays right if the tokens are ever renamed. + /// + private static readonly string[] PostgresKinds = MonitoredEngineKind.All + .Where(kind => MonitoredEngineKind.EngineOf(kind) == CollectorTargetEngine.PostgreSql) + .ToArray(); + + /// + /// How a sweep treats a fact. Derived by RUNNING the sweep and looking at what it produces, never + /// declared — a declared role would be the hand-maintained list this test exists to replace. + /// + private enum SweepRole + { + /// The sweep produces more than one value for it within a single axis value. + VariedBySweep, + + /// Constant within one axis value, but different between them — the axis decides it. + FixedByAxis, + + /// The engine discriminator itself. Fixed by the question rather than swept. + EngineDiscriminator, + + /// The same value in every shape of every axis value — a gate reading it can only ever see + /// one answer, so a gate reading it is where over-claiming begins. + ConstantEverywhere, + } + + /// + /// The role a fact plays across one axis, given the shapes that axis produces for each of its values. + /// Parameterised over the shape families rather than hard-wired to an axis, because the two axes ask the + /// identical question of different sweeps and a second copy of this arithmetic is exactly the kind of + /// near-duplicate that drifts (#2532). + /// + private static SweepRole RoleOf(PropertyInfo fact, IReadOnlyList shapesPerAxisValue) + { + /* The engine is the precondition of the whole question, not a dimension of it: "does any target of + this engine edition / engine kind run this collector" fixes the engine by construction. Recognised + by TYPE rather than by name, so it stays recognised if the property is ever renamed. */ + if (fact.PropertyType == typeof(CollectorTargetEngine)) + { + return SweepRole.EngineDiscriminator; + } + + var perValue = shapesPerAxisValue + .Select(shapes => shapes.Select(fact.GetValue).Distinct().ToArray()) + .ToArray(); + + if (perValue.Any(values => values.Length > 1)) + { + return SweepRole.VariedBySweep; + } + + return perValue.Select(values => values[0]).Distinct().Count() > 1 + ? SweepRole.FixedByAxis + : SweepRole.ConstantEverywhere; + } + + /// The shapes the EDITION sweep produces, one array per engine edition. + private static CollectorTargetInfo[][] EditionShapes() => AllEngineEditions + .Select(edition => CollectorEngineCapability.TargetsWithEngineEdition(edition).ToArray()) + .ToArray(); + + /// The shapes the KIND sweep produces, one array per PostgreSQL token. + private static CollectorTargetInfo[][] PostgresKindShapes() => PostgresKinds + .Select(kind => CollectorEngineCapability.TargetsWithEngineKind(kind).ToArray()) + .ToArray(); + + private static SweepRole EditionRoleOf(PropertyInfo fact) => RoleOf(fact, EditionShapes()); + + private static SweepRole PostgresKindRoleOf(PropertyInfo fact) => RoleOf(fact, PostgresKindShapes()); + + /// fact name → the SQL Server collectors whose AppliesTo reads it, decoded from IL. + private static SortedDictionary> FactsReadBySqlServerGates() => + FactsReadByGatesOf(CollectorTargetEngine.SqlServer); + + /// The same, for the PostgreSQL definitions — the whole difference between the two guards, as + /// #2532 said it would be. + private static SortedDictionary> FactsReadByPostgresGates() => + FactsReadByGatesOf(CollectorTargetEngine.PostgreSql); + + private static SortedDictionary> FactsReadByGatesOf(CollectorTargetEngine engine) + { + var byFact = new SortedDictionary>(StringComparer.Ordinal); + + foreach (var definition in CollectorCatalog.All.Where(c => c.TargetEngine == engine)) + { + foreach (var fact in FactsReadByGateOf(definition)) + { + if (!byFact.TryGetValue(fact, out var collectors)) + { + byFact[fact] = collectors = new SortedSet(StringComparer.Ordinal); + } + + collectors.Add(definition.Name); + } + } + + return byFact; + } + + /// The facts one definition's gate reads, following calls into + /// other collector-assembly code so a gate refactored behind a helper cannot hide what it reads. + private static SortedSet FactsReadByGateOf(ICollectorSchemaInfo definition) + { + var gate = definition.GetType().GetMethod( + nameof(ICollectorSchemaInfo.AppliesTo), + BindingFlags.Public | BindingFlags.Instance, + binder: null, + types: new[] { typeof(CollectorTargetInfo) }, + modifiers: null); + + Assert.True(gate is not null, $"{definition.Name}: no AppliesTo(CollectorTargetInfo) to decode"); + + var facts = new SortedSet(StringComparer.Ordinal); + CollectFactReads(gate!, new HashSet(), facts); + return facts; + } + + private static void CollectFactReads(MethodBase method, HashSet visited, SortedSet facts) + { + if (!visited.Add(method)) + { + return; + } + + foreach (var callee in CalleesOf(method)) + { + if (callee.DeclaringType == typeof(CollectorTargetInfo)) + { + if (FactByGetterName.TryGetValue(callee.Name, out var fact)) + { + facts.Add(fact.Name); + } + + continue; + } + + /* Follow only into code this repo ships. Walking the framework would never terminate usefully, + and a gate cannot read a target fact through a BCL call anyway — the fact only exists here. */ + if (callee.DeclaringType?.Assembly == typeof(CollectorTargetInfo).Assembly) + { + CollectFactReads(callee, visited, facts); + } + } + } + + /// + /// Every opcode the runtime defines, keyed by its encoded value — taken from + /// by reflection rather than typed out, so the operand-size table below cannot + /// drift from the instruction set it is decoding. + /// + private static readonly IReadOnlyDictionary Opcodes = + typeof(OpCodes) + .GetFields(BindingFlags.Public | BindingFlags.Static) + .Where(field => field.FieldType == typeof(OpCode)) + .Select(field => (OpCode)field.GetValue(null)!) + .GroupBy(opcode => opcode.Value) + .ToDictionary(group => group.Key, group => group.First()); + + /// + /// The methods a method calls, by walking its IL stream instruction by instruction. + /// + /// Decoded properly, not scanned for byte patterns. Searching the stream for a + /// call/callvirt byte would match those values inside other instructions' operands, and a + /// four-byte slice of an operand can resolve to a perfectly valid, entirely unrelated method token — a + /// guard that reports facts nobody reads, in a file whose whole subject is guards that stop guarding. + /// Advancing by each opcode's real operand size is the only way to know a token is a token. + /// + private static IReadOnlyList CalleesOf(MethodBase method) + { + var callees = new List(); + var il = method.GetMethodBody()?.GetILAsByteArray(); + if (il is null) + { + return callees; + } + + var typeArguments = method.DeclaringType?.GetGenericArguments(); + var methodArguments = method.IsGenericMethod ? method.GetGenericArguments() : null; + + var position = 0; + while (position < il.Length) + { + var value = il[position] == 0xFE && position + 1 < il.Length + ? unchecked((short)(0xFE00 | il[position + 1])) + : (short)il[position]; + + Assert.True( + Opcodes.TryGetValue(value, out var opcode), + $"{method.DeclaringType?.Name}.{method.Name}: undecodable opcode 0x{value:X4} at IL offset " + + $"{position}. The stream is being mis-walked from here on, and a mis-walked stream reads " + + "operand bytes as call tokens — stop rather than report facts that were never read."); + + position += opcode.Size; + var operandSize = OperandSize(opcode, il, position); + + if (opcode.OperandType is OperandType.InlineMethod or OperandType.InlineTok) + { + var token = BitConverter.ToInt32(il, position); + try + { + if (method.Module.ResolveMethod(token, typeArguments, methodArguments) is { } callee) + { + callees.Add(callee); + } + } + catch (ArgumentException) + { + /* InlineTok also carries field and type handles; those are not calls. */ + } + } + + position += operandSize; + } + + return callees; + } + + private static int OperandSize(OpCode opcode, byte[] il, int operandStart) => opcode.OperandType switch + { + OperandType.InlineNone => 0, + OperandType.ShortInlineBrTarget or OperandType.ShortInlineI or OperandType.ShortInlineVar => 1, + OperandType.InlineVar => 2, + OperandType.InlineBrTarget or OperandType.InlineField or OperandType.InlineI + or OperandType.InlineMethod or OperandType.InlineSig or OperandType.InlineString + or OperandType.InlineTok or OperandType.InlineType or OperandType.ShortInlineR => 4, + OperandType.InlineI8 or OperandType.InlineR => 8, + OperandType.InlineSwitch => 4 + (4 * BitConverter.ToInt32(il, operandStart)), + _ => throw new NotSupportedException($"unhandled operand type {opcode.OperandType} for {opcode.Name}"), + }; + + /// + /// THE assertion. Every fact a SQL Server gate reads is a fact the + /// sweep either varies or fixes by edition — so no gate can be answered by a single defaulted value. + /// + /// A gate written on a fact the sweep leaves at its default fails EVERY swept shape, and the + /// derivation then reports a permanent, unfixable engine gap for a collector that runs. This is the only + /// check in the suite that fires on that, and it fires at the moment the gate is written rather than the + /// day someone notices a read has gone quiet. + /// + [Fact] + public void EveryFactASqlServerGateReads_IsVariedBySweepOrFixedByEdition() + { + var read = FactsReadBySqlServerGates(); + + var overClaiming = read + .Where(entry => EditionRoleOf(FactByName[entry.Key]) == SweepRole.ConstantEverywhere) + .Select(entry => $" {entry.Key} — read by {string.Join(", ", entry.Value)}") + .ToArray(); + + Assert.True( + overClaiming.Length == 0, + "A SQL Server collector's AppliesTo gate reads a CollectorTargetInfo fact that " + + "CollectorEngineCapability.TargetsWithEngineEdition never varies. Every swept target therefore " + + "carries that fact's default, every one of them fails the gate, and the capability derivation " + + "reports a PERMANENT engine gap on every edition for a collector that runs perfectly well — the " + + "confident-and-wrong message #2511 exists to delete.\n\n" + + "Fix the SWEEP, not this test: add the fact to TargetsWithEngineEdition (varying it, or deriving " + + "it from the engine edition the way the two Azure flags are). There is no list here to add a name " + + "to.\n\n" + + string.Join("\n", overClaiming)); + } + + /// + /// The scan is only worth anything if it reads the real gates, so pin what it finds against what the + /// source plainly says — including a gate that reads NOTHING, which is what shows the decoder + /// discriminates rather than reporting every fact for every collector. + /// + /// A source-parsing guard that matches nothing passes for free and is worse than no guard at all: + /// it converts an open question into false confidence. That has happened in this repo more than once, so + /// the floor is asserted here rather than assumed. + /// + [Fact] + public void TheGateScan_ReadsTheGatesTheSourceActuallyCarries() + { + var sqlServerDefinitions = CollectorCatalog.All + .Where(c => c.TargetEngine == CollectorTargetEngine.SqlServer) + .ToArray(); + + Assert.True( + sqlServerDefinitions.Length >= 30, + $"only {sqlServerDefinitions.Length} SQL Server definitions found — the catalog walk is broken, not the catalog"); + + var read = FactsReadBySqlServerGates(); + + /* The facts today's gates are written on. Not an expected-set pin — a floor, so a decoder that + silently returned nothing cannot pass. */ + Assert.Contains(nameof(CollectorTargetInfo.IsAzureSqlDb), read.Keys); + Assert.Contains(nameof(CollectorTargetInfo.IsAzureManagedInstance), read.Keys); + Assert.Contains(nameof(CollectorTargetInfo.IsAwsRds), read.Keys); + Assert.Contains(nameof(CollectorTargetInfo.SqlMajorVersion), read.Keys); + + var byName = CollectorCatalog.All.ToDictionary(c => c.Name, StringComparer.Ordinal); + + /* The collector the whole feature was filed on: one fact, exactly the one its source names. */ + Assert.Equal( + new[] { nameof(CollectorTargetInfo.IsAzureSqlDb) }, + FactsReadByGateOf(byName["system_health_events"]).ToArray()); + + /* Two facts in one gate, so the decoder is not stopping at the first call it finds. This read three + until #2559 removed HasMsdbAccess from it — msdb access is a grant rather than an engine + capability, so the collector attempts and fails into PERMISSIONS instead of gating off a probe + cached for the connection's life. */ + Assert.Equal( + new[] + { + nameof(CollectorTargetInfo.IsAwsRds), + nameof(CollectorTargetInfo.IsAzureSqlDb), + }, + FactsReadByGateOf(byName["running_jobs"]).ToArray()); + + /* And a gate that reads nothing at all. Without this, "reports every fact for every collector" would + satisfy every assertion above. */ + Assert.Empty(FactsReadByGateOf(byName["wait_stats"])); + } + + /// + /// The sweep is the FULL cross product of the dimensions it varies, not a sample of them. + /// + /// The claim the derivation makes is "there is no target of this engine edition, under any + /// COMBINATION of the other facts, for which this collector runs". A sweep that varied each dimension but + /// only along one axis at a time would satisfy + /// TheSweep_VariesEveryDimensionTheGatesRead while missing combinations entirely — and a + /// conjunctive gate (running_jobs and agent_status are both conjunctions of three facts) + /// would be reported as a permanent gap because the one shape that passes it was never generated. + /// + /// Both sides are measured from the sweep's own output: the dimensions are the facts it varies, the + /// expected shape count is the product of the distinct values it produced for each. Nothing here says how + /// many dimensions there should be, so adding one needs no edit — only keeping it exhaustive does. + /// + [Fact] + public void TheSweep_IsTheFullCrossProductOfTheDimensionsItVaries() + { + foreach (var edition in AllEngineEditions) + { + var shapes = CollectorEngineCapability.TargetsWithEngineEdition(edition).ToArray(); + var varied = TargetFacts.Where(fact => EditionRoleOf(fact) == SweepRole.VariedBySweep).ToArray(); + + Assert.NotEmpty(varied); + + var expected = varied.Aggregate( + 1, + (total, fact) => total * shapes.Select(fact.GetValue).Distinct().Count()); + + Assert.Equal(expected, shapes.Length); + + /* Same size AND no repeats means every combination appears exactly once. Size alone would be + satisfied by a sweep that emitted one shape twice and skipped another. */ + var fingerprints = shapes + .Select(shape => string.Join("|", varied.Select(fact => fact.GetValue(shape)))) + .ToArray(); + + Assert.Equal(shapes.Length, fingerprints.Distinct(StringComparer.Ordinal).Count()); + } + } + + /// + /// Every fact on lands in exactly one derived role, and the roles are + /// non-degenerate: something is varied, something is fixed by edition, and there is exactly one engine + /// discriminator. + /// + /// This is the whole-type view the per-gate check above cannot give. If + /// ever collapsed — one shape per + /// edition, or the Azure flags stopped following the edition — every fact would slide into + /// and the derivation would answer every question from a + /// single target. The gap set would still look plausible; it would just have stopped being derived from + /// anything. + /// + [Fact] + public void EveryTargetFact_LandsInExactlyOneDerivedRole() + { + var roles = TargetFacts.ToDictionary(fact => fact.Name, EditionRoleOf, StringComparer.Ordinal); + + Assert.NotEmpty(roles); + Assert.Single(roles, role => role.Value == SweepRole.EngineDiscriminator); + Assert.Contains(roles, role => role.Value == SweepRole.VariedBySweep); + Assert.Contains(roles, role => role.Value == SweepRole.FixedByAxis); + + /* The two Azure flags are the edition-fixed pair, and they follow the edition in OPPOSITE directions + — asserted here because "fixed by edition" is otherwise satisfied by a flag that is fixed at the + wrong value. */ + var azureShapes = CollectorEngineCapability.TargetsWithEngineEdition( + CollectorEngineCapability.AzureSqlDatabaseEngineEdition).ToArray(); + var miShapes = CollectorEngineCapability.TargetsWithEngineEdition( + CollectorEngineCapability.AzureManagedInstanceEngineEdition).ToArray(); + + Assert.All(azureShapes, shape => Assert.True(shape.IsAzureSqlDb && !shape.IsAzureManagedInstance)); + Assert.All(miShapes, shape => Assert.True(shape.IsAzureManagedInstance && !shape.IsAzureSqlDb)); + } + + /* ───────── the same guard, one engine over (#2532) ───────── */ + + /// + /// THE assertion, PostgreSQL side. Every fact a PostgreSQL gate reads + /// is a fact either varies or fixes by the + /// kind — so no PostgreSQL gate can be answered by a single defaulted value. + /// + /// Why this had to land before the axis did. #2530 deliberately shipped only the ENGINE half + /// of the dispatch gate on this axis, because a sweep without this guard over-claims silently: a fact the + /// sweep leaves at its CLR default fails every shape, and the derivation announces a permanent, + /// unfixable engine gap for a collector that runs perfectly well — on the very engine whose support this + /// work exists to make credible. #2532 made that the explicit prerequisite, and this is it. + /// + /// Nothing here is a list. The facts come out of the gates' IL (the same decoder the edition + /// half uses, filtered to the PostgreSQL definitions) and the roles come out of running the sweep. Fix + /// the SWEEP, never this test — there is no name to add here instead. + /// + [Fact] + public void EveryFactAPostgresGateReads_IsVariedBySweepOrFixedByKind() + { + var read = FactsReadByPostgresGates(); + + var overClaiming = read + .Where(entry => PostgresKindRoleOf(FactByName[entry.Key]) == SweepRole.ConstantEverywhere) + .Select(entry => $" {entry.Key} — read by {string.Join(", ", entry.Value)}") + .ToArray(); + + Assert.True( + overClaiming.Length == 0, + "A PostgreSQL collector's AppliesTo gate reads a CollectorTargetInfo fact that " + + "CollectorEngineCapability.TargetsWithEngineKind never varies. Every swept target therefore " + + "carries that fact's default, every one of them fails the gate, and the capability derivation " + + "reports a PERMANENT engine gap on both PostgreSQL tokens for a collector that runs perfectly " + + "well — the confident-and-wrong message #2511 exists to delete, one engine over.\n\n" + + "Fix the SWEEP, not this test: add the fact to TargetsWithEngineKind (varying it, or deriving it " + + "from the engine kind the way IsAurora is). There is no list here to add a name to.\n\n" + + string.Join("\n", overClaiming)); + } + + /// + /// The PostgreSQL half of the scan reads the gates the source actually carries — the same non-vacuity + /// floor the SQL Server half has, and for the same reason: a source-decoding guard that matches nothing + /// passes for free and converts an open question into false confidence. + /// + [Fact] + public void ThePostgresGateScan_ReadsTheGatesTheSourceActuallyCarries() + { + var definitions = CollectorCatalog.All + .Where(c => c.TargetEngine == CollectorTargetEngine.PostgreSql) + .ToArray(); + + Assert.True( + definitions.Length >= 8, + $"only {definitions.Length} PostgreSQL definitions found — the catalog walk is broken, not the catalog"); + + var read = FactsReadByPostgresGates(); + + /* The three facts today's PostgreSQL gates are written on. A floor, not an expected-set pin. */ + Assert.Contains(nameof(CollectorTargetInfo.IsAurora), read.Keys); + Assert.Contains(nameof(CollectorTargetInfo.IsInRecovery), read.Keys); + Assert.Contains(nameof(CollectorTargetInfo.PostgresMajorVersion), read.Keys); + + var byName = CollectorCatalog.All.ToDictionary(c => c.Name, StringComparer.Ordinal); + + /* The collector #2532 was filed on: one fact, exactly the one its source names. */ + Assert.Equal( + new[] { nameof(CollectorTargetInfo.IsAurora) }, + FactsReadByGateOf(byName["pg_wait_stats"]).ToArray()); + + /* A version floor and a recovery gate, so the decoder is discriminating rather than reporting + IsAurora for everything. */ + Assert.Equal( + new[] { nameof(CollectorTargetInfo.PostgresMajorVersion) }, + FactsReadByGateOf(byName["pg_io_stats"]).ToArray()); + + Assert.Equal( + new[] { nameof(CollectorTargetInfo.IsInRecovery) }, + FactsReadByGateOf(byName["pg_autovacuum_stats"]).ToArray()); + + /* And a gate that reads nothing at all — without this, "reports every fact for every collector" + would satisfy every assertion above. */ + Assert.Empty(FactsReadByGateOf(byName["pg_blocking"])); + + /* No SQL Server fact leaks into the PostgreSQL scan. The two halves differ only by the filter, so a + filter that stopped filtering would look exactly like a guard that was working. */ + Assert.DoesNotContain(nameof(CollectorTargetInfo.IsAzureSqlDb), read.Keys); + Assert.DoesNotContain(nameof(CollectorTargetInfo.SqlMajorVersion), read.Keys); + + /* HasMsdbAccess was a third assertion here and has been REMOVED rather than left passing. Since + #2559 no SQL Server gate reads it either, so the line could no longer fail for the reason it was + written — a filter that stopped filtering would still not have surfaced it. A pin that cannot go + red is worse than no pin, because it reads as coverage. The two facts left are both genuinely + SQL-Server-only and still discriminate. */ + } + + /// + /// The kind sweep is the FULL cross product of the dimensions it varies, for the reason the edition + /// sweep is: the claim being made is "no target of this kind, under any COMBINATION of the other facts", + /// and a sweep that moved one axis at a time would report a conjunctive gate as a permanent gap because + /// the one shape that satisfies it was never generated. + /// + [Fact] + public void TheKindSweep_IsTheFullCrossProductOfTheDimensionsItVaries() + { + Assert.NotEmpty(PostgresKinds); + + foreach (var kind in PostgresKinds) + { + var shapes = CollectorEngineCapability.TargetsWithEngineKind(kind).ToArray(); + var varied = TargetFacts.Where(fact => PostgresKindRoleOf(fact) == SweepRole.VariedBySweep).ToArray(); + + Assert.NotEmpty(varied); + Assert.All(shapes, shape => Assert.Equal(CollectorTargetEngine.PostgreSql, shape.Engine)); + + var expected = varied.Aggregate( + 1, + (total, fact) => total * shapes.Select(fact.GetValue).Distinct().Count()); + + Assert.Equal(expected, shapes.Length); + + var fingerprints = shapes + .Select(shape => string.Join("|", varied.Select(fact => fact.GetValue(shape)))) + .ToArray(); + + Assert.Equal(shapes.Length, fingerprints.Distinct(StringComparer.Ordinal).Count()); + } + } + + /// + /// Every fact lands in exactly one derived role on the KIND axis too, and the roles are non-degenerate: + /// is the one the kind fixes — in OPPOSITE directions for the + /// two tokens, so "fixed by the kind" is not satisfied by a flag stuck at the wrong value — and the + /// version and recovery facts are the ones it varies. + /// + /// A kind sweep that collapsed to one shape per token would slide every fact into + /// and answer every question from a single target, with a gap + /// set that still looked plausible. This is the whole-type view that says so. + /// + [Fact] + public void EveryTargetFact_LandsInExactlyOneDerivedKindRole() + { + var roles = TargetFacts.ToDictionary(fact => fact.Name, PostgresKindRoleOf, StringComparer.Ordinal); + + Assert.NotEmpty(roles); + Assert.Single(roles, role => role.Value == SweepRole.EngineDiscriminator); + Assert.Contains(roles, role => role.Value == SweepRole.VariedBySweep); + + Assert.Equal(SweepRole.FixedByAxis, roles[nameof(CollectorTargetInfo.IsAurora)]); + Assert.Equal(SweepRole.VariedBySweep, roles[nameof(CollectorTargetInfo.IsInRecovery)]); + Assert.Equal(SweepRole.VariedBySweep, roles[nameof(CollectorTargetInfo.PostgresMajorVersion)]); + Assert.Equal(SweepRole.VariedBySweep, roles[nameof(CollectorTargetInfo.PostgresVersionNum)]); + + var stock = CollectorEngineCapability.TargetsWithEngineKind(MonitoredEngineKind.Postgres).ToArray(); + var aurora = CollectorEngineCapability.TargetsWithEngineKind(MonitoredEngineKind.AuroraPostgres).ToArray(); + + Assert.NotEmpty(stock); + Assert.NotEmpty(aurora); + Assert.All(stock, shape => Assert.False(shape.IsAurora)); + Assert.All(aurora, shape => Assert.True(shape.IsAurora)); + } + + /// + /// A kind this build does not recognise produces NO shapes, and the capability answer treats that as + /// silence rather than as "no shape runs it". + /// + /// This is the one place an empty sweep would be catastrophic instead of merely wrong: read as a + /// gap set, an empty sweep makes every collector a permanent gap on every unknown token — which is + /// exactly the row a store written by a NEWER service leaves behind, and exactly the guarantee #2530 + /// was told to keep. + /// + [Theory] + [InlineData(null)] + [InlineData("")] + [InlineData(" ")] + [InlineData("something-a-newer-build-writes")] + public void AnUnrecognisedKind_SweepsNothing_AndStillClaimsNothing(string? engineKind) + { + Assert.Empty(CollectorEngineCapability.TargetsWithEngineKind(engineKind)); + + Assert.All( + CollectorCatalog.All, + definition => Assert.True( + CollectorEngineCapability.IsCollectedOnEngineKind(definition, engineKind), + $"{definition.Name} was claimed as a permanent gap on an unrecognised engine kind")); + } +} diff --git a/Darling/Darling.Tests/ConfigSeedStatementArityTests.cs b/Darling/Darling.Tests/ConfigSeedStatementArityTests.cs new file mode 100644 index 000000000..f610109f7 --- /dev/null +++ b/Darling/Darling.Tests/ConfigSeedStatementArityTests.cs @@ -0,0 +1,148 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Text.RegularExpressions; +using Xunit; + +namespace Darling.Tests; + +/// +/// The config-store seed statements agree with themselves: as many values as columns, and every +/// parameter the command supplies is actually referenced by the SQL. +/// +/// Written after a live first-run failed with 42601: INSERT has more target columns than +/// expressions. config_service named sixteen columns and supplied fifteen values, because +/// $12 was bound on the command and never written into the VALUES list. Every subsequent value +/// shifted one column left, so 'seed' landed on updated_at and updated_by had +/// nothing at all. +/// +/// What that cost, which is why this is worth a pin. The seed is how a FRESH store learns +/// the file's settings. It failed, the service logged it and carried on with the file config — and then +/// the control plane, which is authoritative after first contact, answered from an unseeded row where +/// every toggle is false. So web.enabled and mcp.enabled were true in darling.json, false +/// in the store, the store won, and neither endpoint ever listened. On a container deployment that is +/// the entire first-run experience, and nothing in the suite noticed because no test drives the seed +/// against an empty store. +/// +/// Checked by parsing the shipped source rather than by executing it: the seed methods are private +/// and reaching them needs a live store, which is exactly why this went unpinned. Arity is a property of +/// the text, so the text is what gets read. +/// +public sealed class ConfigSeedStatementArityTests +{ + private static readonly Regex Insert = new( + @"INSERT\s+INTO\s+(?[a-z_.]+)\s*\((?[^)]*)\)\s*VALUES\s*\((?[^)]*)\)", + RegexOptions.IgnoreCase | RegexOptions.Singleline | RegexOptions.Compiled); + + [Fact] + public void EverySeedInsert_SuppliesExactlyAsManyValuesAsColumns() + { + var source = File.ReadAllText(ProviderPath()); + var problems = new List(); + var seen = 0; + + foreach (Match m in Insert.Matches(source)) + { + var table = m.Groups["table"].Value; + var cols = Split(m.Groups["cols"].Value); + var vals = Split(m.Groups["vals"].Value); + seen++; + + if (cols.Length != vals.Length) + { + problems.Add( + $"{table}: {cols.Length} columns but {vals.Length} values" + + $" — first unmatched column is '{cols.ElementAtOrDefault(Math.Min(vals.Length, cols.Length - 1))}'"); + } + } + + /* + A regex that stops matching passes for free, which is the failure this whole file is about. + The floor is DERIVED from the file rather than a literal: `seen >= 4` was true of today's four + statements, and would have gone on being true the day a fifth was added and one stopped + parsing. Counting the INSERTs independently of the pattern that parses them means the two + have to agree, so a statement the pattern cannot read is a failure rather than an absence. + + The specific way it could stop reading one: the column and value lists are captured with + [^)]*, so a future seed carrying a function call — COALESCE($1, 0) — closes the capture early + or fails the match outright. Loud is fine; silent is not, and this is what makes it loud. + */ + var declared = Regex.Matches(source, @"INSERT\s+INTO\s+[a-z_.]+", RegexOptions.IgnoreCase).Count; + Assert.True( + seen == declared, + $"parsed {seen} seed INSERTs but the file declares {declared} — the pattern cannot read one of " + + "them (a parenthesis inside a column or value list will do it), so its arity is unchecked"); + + Assert.True(problems.Count == 0, + "a seed INSERT disagrees with itself, which fails at RUN time as 42601 and leaves a fresh " + + "store unseeded — after which the control plane answers from defaults and the file's " + + "settings are silently overridden:" + Environment.NewLine + string.Join(Environment.NewLine, problems)); + } + + [Fact] + public void EverySeedInsert_ReferencesEveryParameterItBinds() + { + var source = File.ReadAllText(ProviderPath()); + var problems = new List(); + + foreach (Match m in Insert.Matches(source)) + { + var table = m.Groups["table"].Value; + var used = Regex.Matches(m.Groups["vals"].Value, @"\$(\d+)") + .Select(x => int.Parse(x.Groups[1].Value)) + .ToHashSet(); + + if (used.Count == 0) + { + continue; + } + + /* + Positional binding means the parameters are $1..$max with no gaps. A hole is not a + cosmetic gap: it means a value the caller computed is silently dropped, and every + parameter after the hole lands on the wrong column. + */ + var missing = Enumerable.Range(1, used.Max()).Where(n => !used.Contains(n)).ToArray(); + if (missing.Length > 0) + { + problems.Add($"{table}: binds up to ${used.Max()} but never references " + + string.Join(", ", missing.Select(n => "$" + n))); + } + } + + Assert.True(problems.Count == 0, + "a seed INSERT skips a positional parameter, so a supplied value is dropped and the ones " + + "after it shift onto the wrong columns:" + Environment.NewLine + string.Join(Environment.NewLine, problems)); + } + + private static string[] Split(string list) => + list.Split(',') + .Select(x => x.Trim()) + .Where(x => x.Length > 0) + .ToArray(); + + private static string ProviderPath() => + Path.Combine(RepoRoot(), "Darling", "PerformanceMonitor.Darling.Service", "StoreConfigProvider.cs"); + + private static string RepoRoot([CallerFilePath] string thisFile = "") + { + var dir = Path.GetDirectoryName(thisFile)!; + while (dir is not null && !File.Exists(Path.Combine(dir, "PerformanceMonitor.sln")) && !Directory.Exists(Path.Combine(dir, ".git"))) + { + dir = Path.GetDirectoryName(dir); + } + + Assert.NotNull(dir); + return dir!; + } +} diff --git a/Darling/Darling.Tests/Darling.Tests.csproj b/Darling/Darling.Tests/Darling.Tests.csproj index f134af585..84b950052 100644 --- a/Darling/Darling.Tests/Darling.Tests.csproj +++ b/Darling/Darling.Tests/Darling.Tests.csproj @@ -33,6 +33,12 @@ + + ", " ", RegexOptions.Singleline); + + var footerStart = xaml.IndexOf("x:Name=\"SidebarFooter\"", StringComparison.Ordinal); + Assert.True(footerStart >= 0, "the sidebar footer is no longer named SidebarFooter, so this pin cannot find it"); + + /* The footer ends at its closing Border. Bounded rather than open-ended so a ManageTags_Click + further down the file cannot satisfy the assertion by accident. */ + var footerEnd = xaml.IndexOf("", footerStart, StringComparison.Ordinal); + Assert.True(footerEnd > footerStart, "could not find the end of the sidebar footer"); + + var footer = xaml[footerStart..footerEnd]; + + Assert.Contains("ManageTags_Click", footer, StringComparison.Ordinal); + + /* And it is a Button, not a MenuItem smuggled into the footer — a MenuItem outside a menu is not a + thing a user can click. */ + var handlerAt = footer.IndexOf("ManageTags_Click", StringComparison.Ordinal); + var elementStart = footer.LastIndexOf('<', handlerAt); + Assert.StartsWith(" + /// The context-menu door stays too. It is the convenient one once tags exist, and removing it while + /// "fixing" discoverability would trade one missing entry point for another. + /// + [Fact] + public void TheGroupHeaderContextMenuEntryPoint_IsStillThere() + { + var xaml = Regex.Replace(File.ReadAllText(MainWindowXaml()), @"", " ", RegexOptions.Singleline); + + Assert.Contains("TagHeader_ContextMenuOpening", xaml, StringComparison.Ordinal); + Assert.Contains("TagHeaderContextMenu_NewTag_Click", xaml, StringComparison.Ordinal); + } + + private static string MainWindowXaml([CallerFilePath] string thisFile = "") => + Path.GetFullPath(Path.Combine( + Path.GetDirectoryName(thisFile)!, "..", "PerformanceMonitor.Darling.Viewer", "MainWindow.xaml")); +} diff --git a/Darling/Darling.Tests/packages.lock.json b/Darling/Darling.Tests/packages.lock.json index ca736f2fb..f66af24d2 100644 --- a/Darling/Darling.Tests/packages.lock.json +++ b/Darling/Darling.Tests/packages.lock.json @@ -11,6 +11,11 @@ "xunit.v3.mtp-v1": "[3.2.2]" } }, + "AWSSDK.Core": { + "type": "Transitive", + "resolved": "4.0.102.1", + "contentHash": "vsf+/o+S/euH/wG7bO1cha+OWaHvjgvpDeGFO6l3Vcdio9GlSfXkZDD2BogMuAJ+/jDyqqSgqmYVolPKV9ie/A==" + }, "HarfBuzzSharp": { "type": "Transitive", "resolved": "8.3.1.1", @@ -685,6 +690,8 @@ "performancemonitor.darling.service": { "type": "Project", "dependencies": { + "AWSSDK.RDS": "[4.0.104.4, )", + "BlackwellSystems.Gcf": "[0.2.0, )", "Microsoft.Data.SqlClient": "[7.0.2, )", "Microsoft.Extensions.Hosting": "[10.0.11, )", "Microsoft.Extensions.Hosting.WindowsServices": "[10.0.11, )", @@ -713,6 +720,7 @@ "Hardcodet.NotifyIcon.Wpf": "[2.0.1, )", "PerformanceMonitor.Alerting": "[1.0.0, )", "PerformanceMonitor.Analysis": "[1.0.0, )", + "PerformanceMonitor.Collectors": "[1.0.0, )", "PerformanceMonitor.Common": "[1.0.0, )", "PerformanceMonitor.Darling.Analysis": "[1.0.0, )", "PerformanceMonitor.Darling.Storage": "[1.0.0, )", @@ -746,6 +754,21 @@ "ScottPlot.WPF": "[5.1.59, )" } }, + "AWSSDK.RDS": { + "type": "CentralTransitive", + "requested": "[4.0.104.4, )", + "resolved": "4.0.104.4", + "contentHash": "tAhUJwG6SnAFT0ZHL2n+z6c7EDmnpo5fPrZO4HZsBmVj4vV7k1/+j0zoe+wMLWKl8qRTk+gipUuz8LITk9rSyA==", + "dependencies": { + "AWSSDK.Core": "[4.0.102.1, 5.0.0)" + } + }, + "BlackwellSystems.Gcf": { + "type": "CentralTransitive", + "requested": "[0.2.0, )", + "resolved": "0.2.0", + "contentHash": "HnL0i5xnmNDxv0j4kqGTgUxZ+uTEuRtM4S31NYS/XdL1kaQ77TfknqpW0yUALyt4KJRmzVKXgI2CmqdhvMQ9cA==" + }, "CredentialManagement": { "type": "CentralTransitive", "requested": "[1.0.2, )", diff --git a/Darling/PerformanceMonitor.Darling.Analysis/AnalysisShutdown.cs b/Darling/PerformanceMonitor.Darling.Analysis/AnalysisShutdown.cs index da9d4cf3c..f8720ac81 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/AnalysisShutdown.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/AnalysisShutdown.cs @@ -13,28 +13,89 @@ namespace PerformanceMonitor.Darling.Analysis; /// -/// Classifies whether an analysis-pass failure is the residue of the host shutting down, so the -/// catch sites can tell unfinished because we asked it to stop from unfinished because -/// something broke (#2299). Before this, a clean Stop-Service logged seven ERRORs from +/// Why an analysis pass stopped early, when it did (#2430). +/// +public enum AnalysisAbandonKind +{ + /// Not an abandonment at all — a genuine fault, which keeps its ERROR. + None, + + /// The host is stopping. Expected, costs one Information line, and the next start recomputes. + Shutdown, + + /// + /// The pass outran its per-run budget on a service that is otherwise fine. Expected in the sense + /// that we raised it, NOT in the sense that it is healthy — a pass that cannot finish inside its + /// budget is losing that server a cycle of findings, so it is a warning and not an Information. + /// + Timeout +} + +/// +/// Classifies whether an analysis-pass failure is the residue of an abandonment we asked for, so the +/// catch sites can tell unfinished because we called it off from unfinished because +/// something broke (#2299) — and, since #2430, which of the two ways we called it off. +/// +/// Before #2299, a clean Stop-Service logged seven ERRORs from /// work still in flight after "collection loop stopped" — the loop's data source is disposed at /// method scope exit and the managed postmaster is then pg_ctl stop -m fast-ed, so the /// abandoned pass's next store read throws (or the server /// kills its open connection with 57P01), and those seven lines were 7 of the day's 9 ERRORs, -/// burying the two that meant something. +/// burying the two that meant something. /// public static class AnalysisShutdown { /// - /// True when this failure should be ABANDONED quietly because the host is stopping: the - /// stopping token has fired AND the exception is a shape shutdown produces. Both halves are + /// True when this failure should be ABANDONED quietly rather than logged as a fault: the pass's + /// token has fired AND the exception is a shape an abandonment produces. Both halves are /// load-bearing — the same exceptions with the token NOT signalled mean a data source was - /// disposed (or a connection administratively killed) while the service was meant to be - /// running, which is a real bug whose only evidence is exactly this text, so it must stay - /// an ERROR. Catch sites use this in a when filter so shutdown residue propagates - /// (unwinding the pass to one Information line) instead of being swallowed per-metric. + /// disposed (or a connection administratively killed) while the pass was meant to be running, + /// which is a real bug whose only evidence is exactly this text, so it must stay an ERROR. + /// Component catch sites use this in a when filter so the residue PROPAGATES — unwinding + /// the pass to the one line chooses — instead of being swallowed + /// per-metric. + /// + /// This was IsShutdownAbandon until #2430, and the rename is the point. The token it + /// is handed is now the pass's armed budget, so "the token fired" no longer means "we are + /// stopping": at this level the REASON does not matter, only that nobody should see nine ERRORs + /// for work we ourselves called off. Deciding which reason is 's job, once, + /// where both tokens are in scope. /// - public static bool IsShutdownAbandon(Exception ex, CancellationToken stoppingToken) => - stoppingToken.IsCancellationRequested && IsShutdownResidue(ex); + public static bool IsExpectedAbandon(Exception ex, CancellationToken passToken) => + passToken.IsCancellationRequested && IsShutdownResidue(ex); + + /// + /// Which kind of abandonment this failure is, or if it is a + /// genuine fault (#2430). Asked ONCE, at the top of the pass, because it is the only place both + /// tokens are in scope — and because the answer decides a log LEVEL and a sentence, both of which + /// are read by someone deciding whether to investigate. + /// + /// Shutdown is tested first and wins a tie: a stop arriving during an already-overrunning + /// pass is a stop, and reporting it as a timeout would invent an incident out of a clean + /// Stop-Service. + /// + /// The timeout arm is deliberately narrower than the shutdown arm. A budget expiring produces + /// exactly one shape — the token observed properly — whereas a disposed data source or a 57P0x + /// means the store went away, which a timeout on a running service does not cause. Widening this + /// arm to the full residue set would quietly relabel the very bug #2299 kept an ERROR, for the + /// whole window after any pass overruns. Those fall through to + /// and stay faults. + /// + public static AnalysisAbandonKind Classify( + Exception ex, CancellationToken shutdownToken, CancellationToken passToken) + { + if (shutdownToken.IsCancellationRequested && IsShutdownResidue(ex)) + { + return AnalysisAbandonKind.Shutdown; + } + + if (passToken.IsCancellationRequested && ex is OperationCanceledException) + { + return AnalysisAbandonKind.Timeout; + } + + return AnalysisAbandonKind.None; + } /// /// The exception shapes a stop produces, detected structurally (the diff --git a/Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs b/Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs index afab53daf..636fb09d2 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/DarlingAnalysisService.cs @@ -97,6 +97,19 @@ public sealed class DarlingAnalysisService /// public string? InsufficientDataMessage { get; private set; } + /// + /// How the last pass ended EARLY, or null when it ran through (#2430). Set inside the pass's own + /// catch, so here means a genuine fault: the pass reached the + /// catch and the classifier said it was not an abandonment. + /// + /// Carried out to the caller because the caller cannot re-derive it. "No findings and the + /// budget token has fired" is true of a fault as well as of a timeout, and inferring a timeout from + /// it buries the fault's ERROR under a Warning that says the pass merely ran out of time. The pass + /// has already classified this once and logged the one line for it; this is how the scheduler reads + /// that answer instead of guessing at a second one. + /// + public AnalysisAbandonKind? EndedEarlyAs { get; private set; } + public DarlingAnalysisService(NpgsqlDataSource postgres, IPlanFetcher? planFetcher = null, ILogger? logger = null) { _postgres = postgres ?? throw new ArgumentNullException(nameof(postgres)); @@ -115,11 +128,25 @@ public DarlingAnalysisService(NpgsqlDataSource postgres, IPlanFetcher? planFetch /// Runs the full analysis pipeline for a server. /// Default time range is the last 4 hours. Host-UTC window (Lite's clock semantics — /// Darling's collectors stamp rows with the service host's UTC clock). + /// + /// #2506: moves the END of that window off "now" while + /// stays its LENGTH, so an incident can be analyzed where it happened. + /// Null — every caller but the anchored MCP tool — is the pre-#2506 behaviour exactly. Anchoring + /// reaches the whole pipeline through the context, including the anomaly detector's hour-of-day × + /// day-of-week baseline, which is keyed off the window rather than off the clock; that is what makes + /// the answer for a past window the same KIND of answer, and not merely a differently-filtered one. + /// An anchored pass does not persist — see . /// + [System.Diagnostics.CodeAnalysis.SuppressMessage("Design", "CA1068:CancellationToken parameters must come last", + Justification = "The two tokens are at positions 4 and 5 and the Darling worker passes them POSITIONALLY. " + + "Moving asOfUtc ahead of them to satisfy the rule would silently rebind that call site's " + + "arguments — a compiling change of meaning on the one caller that matters. Appending is the " + + "only edit that cannot do that, and Lite's twin keeps the same order so the two stay transplantable.")] public async Task> AnalyzeAsync( - int serverId, string serverName, int hoursBack = 4, CancellationToken cancellationToken = default) + int serverId, string serverName, int hoursBack = 4, CancellationToken cancellationToken = default, + CancellationToken shutdownToken = default, DateTime? asOfUtc = null) { - var timeRangeEnd = DateTime.UtcNow; + var timeRangeEnd = asOfUtc ?? DateTime.UtcNow; var timeRangeStart = timeRangeEnd.AddHours(-hoursBack); var context = new AnalysisContext @@ -128,7 +155,16 @@ public async Task> AnalyzeAsync( ServerName = serverName, TimeRangeStart = timeRangeStart, TimeRangeEnd = timeRangeEnd, - CancellationToken = cancellationToken + AsOfUtc = asOfUtc, + CancellationToken = cancellationToken, + + /* #2430. The fifth argument is what keeps the abandon classification truthful once + cancellationToken is a BUDGET rather than the stopping token. Defaulting it to None is + deliberate rather than lazy: an on-demand caller (the MCP analyze_server tool, the + Viewer) has no service stop to distinguish, so its cancellations are timeouts and + should read as timeouts. The scheduled worker is the one caller that has both, and it + is the only one that passes both. */ + ShutdownToken = shutdownToken }; return await AnalyzeAsync(context); @@ -144,6 +180,7 @@ public async Task> AnalyzeAsync(AnalysisContext context) IsAnalyzing = true; InsufficientDataMessage = null; + EndedEarlyAs = null; try { @@ -248,43 +285,80 @@ cannot unwind from inside it — these boundary checks are what turn the token i ?? FactRemediation.BuildMissingIndexAction(finding); // WS4: missing-index CREATE — copy-paste only } - // 7. Insert the survivors in one batched pass, persisting remediation_action_json. - await _findingStore.InsertFindingsAsync(findings, context); + // 7. Insert the survivors in one batched pass, persisting remediation_action_json — + // UNLESS the window was anchored at a past instant (#2506), in which case the pass is + // exploratory and writes nothing. The findings are still built, enriched and returned in + // full; only the row is withheld, because the row would claim to be a current + // observation. AnalysisContext.PersistFindings carries the whole argument. + if (context.PersistFindings) + { + await _findingStore.InsertFindingsAsync(findings, context); + } LastAnalysisTime = DateTime.UtcNow; // 8. Notify listeners — the returned/enriched findings (now action-bearing) also // flow back to the caller (the worker), which routes them to the shared - // AnalysisNotificationService. - AnalysisCompleted?.Invoke(this, new AnalysisCompletedEventArgs + // AnalysisNotificationService. Gated with the insert for the same reason and not a + // weaker one: this event is how findings reach notification, and an alert about last + // Tuesday delivered today is the persistence problem with a shorter fuse. + if (context.PersistFindings) { - ServerId = context.ServerId, - ServerName = context.ServerName, - Findings = findings, - AnalysisTime = LastAnalysisTime.Value - }); + AnalysisCompleted?.Invoke(this, new AnalysisCompletedEventArgs + { + ServerId = context.ServerId, + ServerName = context.ServerName, + Findings = findings, + AnalysisTime = LastAnalysisTime.Value + }); + } _logger?.LogInformation( - "[DarlingAnalysisService] Analysis complete for {Server}: {Count} finding(s), highest severity {Severity:F2}", - context.ServerName, findings.Count, findings.Count > 0 ? findings.Max(f => f.Severity) : 0); + "[DarlingAnalysisService] Analysis complete for {Server}: {Count} finding(s), highest severity {Severity:F2}{Exploratory}", + context.ServerName, findings.Count, findings.Count > 0 ? findings.Max(f => f.Severity) : 0, + context.PersistFindings ? string.Empty : " (anchored window — exploratory, not persisted)"); return findings; } - catch (Exception ex) when (AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) - { - /* #2299: the ONE line a stop is allowed to cost. The component catches let shutdown - residue propagate instead of logging it per-metric, so seven ERRORs collapse to - this Information — and it states the loss honestly: whatever this pass would have - written is gone, and the next scheduled pass recomputes it from the store. */ - _logger?.LogInformation( - "[DarlingAnalysisService] Analysis abandoned at shutdown for {Server} — this pass's findings are lost by design; the next pass recomputes them ({Detail})", - context.ServerName, ex.Message); - return []; - } catch (Exception ex) { - _logger?.LogError("[DarlingAnalysisService] Analysis failed for {Server}: {Message}", - context.ServerName, ex.Message); + /* #2299: the ONE line an abandonment is allowed to cost. The component catches let the + residue propagate instead of logging it per-metric, so seven ERRORs collapse to a single + line here — and it states the loss honestly: whatever this pass would have written is + gone, and the next scheduled pass recomputes it from the store. + + #2430 split that line in two, because the pass token now fires for two very different + reasons and only one of them is fine. Getting this wrong is the reason the Lite fix could + not simply be ported: arm the token with a budget while the classifier still asks "are we + stopping?", and every ordinary overrun on a healthy service reports itself at Information + as a clean stop — a wrong answer wearing a calm one's clothes, on exactly the signal + someone would use to decide the budget needs raising. + + Classified ONCE, in the catch body rather than across two exception filters, because the + three outcomes are one decision and splitting it would mean evaluating it twice and + letting the halves drift. */ + EndedEarlyAs = AnalysisShutdown.Classify(ex, context.ShutdownToken, context.CancellationToken); + + switch (EndedEarlyAs) + { + case AnalysisAbandonKind.Shutdown: + _logger?.LogInformation( + "[DarlingAnalysisService] Analysis abandoned at shutdown for {Server} — this pass's findings are lost by design; the next pass recomputes them ({Detail})", + context.ServerName, ex.Message); + break; + + case AnalysisAbandonKind.Timeout: + _logger?.LogWarning( + "[DarlingAnalysisService] Analysis for {Server} was cancelled at its per-pass budget and unwound as asked — this cycle produces no findings and the next one recomputes them. A pass that keeps hitting this is not finishing inside its budget, which is a server whose analysis is quietly getting less complete, not a stop ({Detail})", + context.ServerName, ex.Message); + break; + + default: + _logger?.LogError("[DarlingAnalysisService] Analysis failed for {Server}: {Message}", + context.ServerName, ex.Message); + break; + } + return []; } finally @@ -296,10 +370,15 @@ cannot unwind from inside it — these boundary checks are what turn the token i /// /// Runs the collect + score pipeline without graph traversal. /// Returns raw scored facts with amplifier details for direct inspection. + /// + /// #2506: anchors the END of the window; null is "now", which is + /// every caller but the anchored MCP tool. Nothing here persists, so the anchor carries no + /// write-side question — this is a read that happens to score what it read. /// - public async Task> CollectAndScoreFactsAsync(int serverId, string serverName, int hoursBack = 4) + public async Task> CollectAndScoreFactsAsync( + int serverId, string serverName, int hoursBack = 4, DateTime? asOfUtc = null) { - var timeRangeEnd = DateTime.UtcNow; + var timeRangeEnd = asOfUtc ?? DateTime.UtcNow; var timeRangeStart = timeRangeEnd.AddHours(-hoursBack); var context = new AnalysisContext @@ -307,7 +386,8 @@ public async Task> CollectAndScoreFactsAsync(int serverId, string ser ServerId = serverId, ServerName = serverName, TimeRangeStart = timeRangeStart, - TimeRangeEnd = timeRangeEnd + TimeRangeEnd = timeRangeEnd, + AsOfUtc = asOfUtc }; try @@ -379,10 +459,16 @@ public async Task> GetLatestFindingsAsync(int serverId) /// Gets recent findings for a server within the given time range. The MCP findings read /// passes so its occurrence stats cover /// the whole window; the store's default 100 stays for everyone else. + /// + /// #2506: anchors the window's END. This one is a pure read of + /// rows the SCHEDULED passes already wrote, so anchoring it asks "what did analysis say about this + /// server at the time" — the only way to see findings the retention sweep has not yet reached but + /// the default 24-hour window has scrolled past. /// - public async Task> GetRecentFindingsAsync(int serverId, int hoursBack = 24, int limit = 100) + public async Task> GetRecentFindingsAsync( + int serverId, int hoursBack = 24, int limit = 100, DateTime? asOfUtc = null) { - return await _findingStore.GetRecentFindingsAsync(serverId, hoursBack, limit); + return await _findingStore.GetRecentFindingsAsync(serverId, hoursBack, limit, asOfUtc); } /// @@ -434,7 +520,7 @@ private async Task GetTotalDataSpanHoursAsync(int serverId, Cancellation return Convert.ToDouble(result); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken)) { /* Probe failure reads as "no data yet" — EXCEPT shutdown residue, which must not be allowed to masquerade as a 0-hour history (#2299): it propagates to the pass's diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgAnomalyDetector.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgAnomalyDetector.cs index e9e82e7da..6769d082b 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgAnomalyDetector.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgAnomalyDetector.cs @@ -40,9 +40,10 @@ namespace PerformanceMonitor.Darling.Analysis; /// Postgres discipline (see PgFindingStore): the SQL is Lite-verbatim against the /// V4 passthrough views (already dialect-shared — no QUALIFY in the detector /// queries), with every window bound a naive-UTC Kind-Unspecified parameter. -/// Lite's SQL contains no bare NOW()/CURRENT_TIMESTAMP; the one C#-side "now" -/// (the 30-day baseline-data gate in 's $2) is -/// bound as a naive-UTC parameter exactly like Lite binds DateTime.UtcNow. +/// Lite's SQL contains no bare NOW()/CURRENT_TIMESTAMP; the 30-day baseline-data +/// gate in 's $2 is bound as a naive-UTC parameter +/// exactly like Lite binds it, and since #2506 comes off the window's end rather +/// than off the clock, so every bound in this class is window-derived. /// SQL is exposed const so Darling.Tests can pin the dialect ungated /// ($N positional parameters, no QUALIFY, no bare now(), no N'' literals) — /// the PgFindingStore/DarlingAlertReadAdapter pattern. @@ -106,7 +107,13 @@ public async Task> DetectAnomaliesAsync(AnalysisContext context) var anomalies = new List(); // Check if baseline period has any data at all — if not, skip all anomaly detection. - if (!await HasBaselineDataAsync(context.ServerId, context.CancellationToken)) + /* #2506: the gate's 30 days are measured back from the WINDOW's end, not from the clock. Every + other bound in this class already comes off context.TimeRangeStart/End, and the baseline this + gate is guarding is computed at context.TimeRangeStart too — so asking "was anything collected + in the 30 days before now" while the baseline reads the 30 days before an anchored window was + the one place the two could disagree. Identical for an unanchored pass, whose TimeRangeEnd IS + now. */ + if (!await HasBaselineDataAsync(context.ServerId, context.TimeRangeEnd, context.CancellationToken)) return anomalies; // Existing detection methods (upgraded to time-bucketed baselines) @@ -129,8 +136,10 @@ public async Task> DetectAnomaliesAsync(AnalysisContext context) /// /// Baseline-data gate: wait_stats as canary — if waits are collected, other data is too. - /// $2 is the ONLY "now"-derived bound in this class (DateTime.UtcNow.AddDays(-30), bound - /// naive-UTC Kind-Unspecified — never a bare now(), which would be timestamptz). + /// $2 is context.TimeRangeEnd.AddDays(-30), bound naive-UTC Kind-Unspecified — never a bare + /// now(), which would be timestamptz. It was DateTime.UtcNow.AddDays(-30) until #2506; the + /// two are the same value on every unanchored pass, and only the window-derived form stays correct + /// when the pass is anchored at a past instant. /// public const string HasBaselineDataSql = @" SELECT (SELECT COUNT(*) FROM v_wait_stats @@ -365,7 +374,7 @@ private async Task DetectObjectStatsAnomalies(AnalysisContext context, List - private async Task HasBaselineDataAsync(int serverId, CancellationToken cancellationToken) + private async Task HasBaselineDataAsync(int serverId, DateTime windowEnd, CancellationToken cancellationToken) { try { @@ -383,14 +392,14 @@ private async Task HasBaselineDataAsync(int serverId, CancellationToken ca using var cmd = new NpgsqlCommand(HasBaselineDataSql, connection); cmd.Parameters.AddWithValue(serverId); - /* Lite binds DateTime.UtcNow.AddDays(-30) here too — the parameterized "now", - made Kind-Unspecified for the naive-UTC timestamp columns. */ - cmd.Parameters.AddWithValue(AsNaive(DateTime.UtcNow.AddDays(-30))); + /* Lite binds the same window end minus 30 days here — the parameterized "now" (or the + #2506 anchor), made Kind-Unspecified for the naive-UTC timestamp columns. */ + cmd.Parameters.AddWithValue(AsNaive(windowEnd.AddDays(-30))); var count = Convert.ToInt64(await cmd.ExecuteScalarAsync(cancellationToken) ?? 0); return count > 0; } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken)) { /* Silent on a genuine fault BY DESIGN (Lite's gate posture: an unreadable canary reads as "no baseline data" and detection just sits out the pass) — but shutdown residue is @@ -462,7 +471,7 @@ private async Task DetectCpuAnomalies(AnalysisContext context, List anomal Metadata = metadata }); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] CPU anomaly detection failed: {Message}", ex.Message); } @@ -573,7 +582,7 @@ floor is what was measured WITH the 5.0 cutoff. The ratio still rides the metada Metadata = metadata }); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] Wait anomaly detection failed: {Message}", ex.Message); } @@ -671,7 +680,7 @@ so normalize them to per-hour before the ratio — otherwise the ratio scales wi }); } } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] Blocking anomaly detection failed: {Message}", ex.Message); } @@ -763,7 +772,7 @@ private async Task DetectIoAnomalies(AnalysisContext context, List anomali }); } } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] I/O anomaly detection failed: {Message}", ex.Message); } @@ -827,7 +836,7 @@ private async Task DetectBatchRequestAnomalies(AnalysisContext context, List an Metadata = metadata }); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] Session anomaly detection failed: {Message}", ex.Message); } @@ -956,7 +965,7 @@ private async Task DetectQueryDurationAnomalies(AnalysisContext context, List ano Metadata = metadata }); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgAnomalyDetector] Memory anomaly detection failed: {Message}", ex.Message); } diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgBaselineProvider.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgBaselineProvider.cs index 96534f1f7..181466795 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgBaselineProvider.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgBaselineProvider.cs @@ -214,7 +214,7 @@ internal static bool IsCommandTimeout(Exception ex) => return buckets; } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, cancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken)) { /* A command TIMEOUT and a genuine connection fault are the same message here, and that cost real diagnosis time on the dogfood box: Npgsql surfaces its own client-side timeout as diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgBlockingPairRowQuery.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgBlockingPairRowQuery.cs index 6eb37e236..78b543c9d 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgBlockingPairRowQuery.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgBlockingPairRowQuery.cs @@ -9,6 +9,7 @@ using System; using System.Collections.Generic; using System.Data.Common; +using System.Threading; using System.Threading.Tasks; using Npgsql; using PerformanceMonitor.Analysis; @@ -125,8 +126,15 @@ ORDER BY event_time DESC /// fragments as the blocked-process-report queries, against v_dmv_blocking_snapshots, so /// maps it unchanged. Takes a command factory so the caller's connection runs it — the Lite shape, kept so /// the future drill-down/viewer slices call it identically. + /// + /// #2443: the token is required, not defaulted. Three callers share this fetch and two of + /// them are on the analysis pass; a default would have let either keep passing nothing while the + /// signature claimed the read was abandonable. The viewer's call is the one that legitimately has + /// no pass to abandon, and it says so at its own call site rather than here. /// - internal static async Task AppendDmvSnapshotRowsAsync(Func createCommand, List rows, int serverId, DateTime start, DateTime end) + internal static async Task AppendDmvSnapshotRowsAsync( + Func createCommand, List rows, int serverId, DateTime start, DateTime end, + CancellationToken cancellationToken) { var dmv = new List(); using (var cmd = createCommand()) @@ -136,8 +144,8 @@ internal static async Task AppendDmvSnapshotRowsAsync(Func create cmd.Parameters.AddWithValue(DateTime.SpecifyKind(start, DateTimeKind.Unspecified)); cmd.Parameters.AddWithValue(DateTime.SpecifyKind(end, DateTimeKind.Unspecified)); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) dmv.Add(Read(reader)); } diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Blocking.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Blocking.cs index f32437d8d..333220811 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Blocking.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Blocking.cs @@ -29,7 +29,7 @@ ORDER BY collection_time DESC private async Task CollectTopDeadlocks(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TopDeadlocksSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -37,8 +37,8 @@ private async Task CollectTopDeadlocks(AnalysisFinding finding, AnalysisContext cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { /* #1140: parse the involved objects from the graph for the dedup fingerprint + a readable Objects field. The raw graph XML is NOT surfaced (it would bloat the alert detail). */ @@ -88,7 +88,7 @@ ORDER BY wait_time_ms DESC private async Task CollectTopBlockingChains(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TopBlockingChainsSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -96,8 +96,8 @@ private async Task CollectTopBlockingChains(AnalysisFinding finding, AnalysisCon cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -141,7 +141,7 @@ ORDER BY event_time DESC /// private async Task CollectReconstructedBlockingChains(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(ReconstructedChainsSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -149,16 +149,17 @@ private async Task CollectReconstructedBlockingChains(AnalysisFinding finding, A cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var rows = new List(); - using (var reader = await cmd.ExecuteReaderAsync()) + using (var reader = await cmd.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) rows.Add(PgBlockingPairRowQuery.Read(reader)); } // Always-on DMV blocking snapshot fallback. Merge BEFORE the empty check so DMV-only blocking // (blocked-process-report unavailable, e.g. AWS RDS) still reconstructs. await PgBlockingPairRowQuery.AppendDmvSnapshotRowsAsync( - connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd); + connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd, + context.CancellationToken); if (rows.Count == 0) return; @@ -207,7 +208,7 @@ ORDER BY total_wait_ms DESC private async Task CollectLockModeBreakdown(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(LockModeBreakdownSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -215,8 +216,8 @@ private async Task CollectLockModeBreakdown(AnalysisFinding finding, AnalysisCon cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Config.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Config.cs index cab40b020..59043943d 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Config.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Config.cs @@ -28,14 +28,14 @@ FROM v_database_config private async Task CollectConfigIssues(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(ConfigIssuesSql, connection); cmd.Parameters.AddWithValue(context.ServerId); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var issues = new List(); if (!reader.IsDBNull(2) && reader.GetBoolean(2)) issues.Add("auto_shrink ON"); diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Plans.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Plans.cs index 1451c3a4f..33378628a 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Plans.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Plans.cs @@ -49,22 +49,27 @@ private async Task CollectPlanAnalysis(AnalysisFinding finding, AnalysisContext string? planHandle = null; try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(PlanHandleLookupSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(queryHash); - using var reader = await cmd.ExecuteReaderAsync(); - if (await reader.ReadAsync() && !reader.IsDBNull(0)) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (await reader.ReadAsync(context.CancellationToken) && !reader.IsDBNull(0)) planHandle = reader.GetString(0); } - catch { return; } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* No plan_handle for this hash — the fetch below has nothing to ask for. An abandonment + is NOT swallowed here (#2443). */ + return; + } if (string.IsNullOrEmpty(planHandle)) return; // Fetch plan XML live from SQL Server - var planXml = await _planFetcher.FetchPlanXmlAsync(context.ServerId, planHandle); + var planXml = await _planFetcher.FetchPlanXmlAsync(context.ServerId, planHandle, context.CancellationToken); if (string.IsNullOrEmpty(planXml)) return; try @@ -142,15 +147,15 @@ private async Task CollectPlanAdvisoryDetail(AnalysisFinding finding, AnalysisCo /* PG port: Lite scopes the connection in a block to release its DuckDB read lock before the CPU-only parse; the scoping is kept so the connection closes before the parse, even though PG holds no lock. */ - await using (var connection = await _postgres.OpenConnectionAsync()) + await using (var connection = await _postgres.OpenConnectionAsync(context.CancellationToken)) { using var cmd = new NpgsqlCommand(PlanAdvisoryXmlSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { if (!reader.IsDBNull(0)) planXmls.Add(reader.GetString(0)); @@ -190,9 +195,10 @@ private async Task CollectPlanAdvisoryDetail(AnalysisFinding finding, AnalysisCo .ToList(); } } - catch + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { // Plan read/parse can fail on malformed XML — skip, the detail is best-effort. + // An abandonment is NOT swallowed here (#2443). } } } diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Queries.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Queries.cs index dc2b980dc..e6301c4e4 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Queries.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Queries.cs @@ -48,7 +48,7 @@ ORDER BY cpu_time_ms DESC private async Task CollectQueriesAtSpike(AnalysisFinding finding, AnalysisContext context) { // Find the peak CPU time, then get queries active within 2 minutes of it - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); // Step 1: Find when the spike occurred using var peakCmd = new NpgsqlCommand(SpikePeakSql, connection); @@ -59,9 +59,9 @@ private async Task CollectQueriesAtSpike(AnalysisFinding finding, AnalysisContex DateTime? peakTime = null; int peakCpu = 0; - using (var peakReader = await peakCmd.ExecuteReaderAsync()) + using (var peakReader = await peakCmd.ExecuteReaderAsync(context.CancellationToken)) { - if (await peakReader.ReadAsync()) + if (await peakReader.ReadAsync(context.CancellationToken)) { peakTime = peakReader.GetDateTime(0); peakCpu = peakReader.GetInt32(1); @@ -78,9 +78,9 @@ private async Task CollectQueriesAtSpike(AnalysisFinding finding, AnalysisContex queryCmd.Parameters.AddWithValue(AsNaive(peakTime.Value.AddMinutes(2))); var items = new List(); - using (var reader = await queryCmd.ExecuteReaderAsync()) + using (var reader = await queryCmd.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -126,7 +126,7 @@ ORDER BY total_cpu_us DESC private async Task CollectTopCpuQueries(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TopCpuQueriesSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -135,8 +135,8 @@ private async Task CollectTopCpuQueries(AnalysisFinding finding, AnalysisContext cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -168,7 +168,7 @@ ORDER BY total_spills DESC private async Task CollectTopSpillingQueries(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TopSpillingQueriesSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -177,8 +177,8 @@ private async Task CollectTopSpillingQueries(AnalysisFinding finding, AnalysisCo cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -248,7 +248,7 @@ ORDER BY worker_ratio DESC /// private async Task CollectParameterSensitiveQueries(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(ParameterSensitiveSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -257,8 +257,8 @@ private async Task CollectParameterSensitiveQueries(AnalysisFinding finding, Ana cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -492,7 +492,7 @@ ORDER BY regression_factor DESC /// private async Task CollectRegressedQueries(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(RegressedQueriesSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -505,8 +505,8 @@ private async Task CollectRegressedQueries(AnalysisFinding finding, AnalysisCont cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -567,7 +567,7 @@ private async Task CollectBadActorDetail(AnalysisFinding finding, AnalysisContex var queryHash = finding.RootFactKey.Replace("BAD_ACTOR_", ""); if (string.IsNullOrEmpty(queryHash)) return; - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(BadActorDetailSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -576,8 +576,8 @@ private async Task CollectBadActorDetail(AnalysisFinding finding, AnalysisContex cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); cmd.Parameters.AddWithValue(queryHash); - using var reader = await cmd.ExecuteReaderAsync(); - if (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (await reader.ReadAsync(context.CancellationToken)) { finding.DrillDown!["bad_actor_query"] = new { @@ -610,7 +610,7 @@ ORDER BY waiter_count DESC private async Task CollectPendingGrants(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(PendingGrantsSql, connection); cmd.CommandTimeout = DrillDownCommandTimeoutSeconds; @@ -619,8 +619,8 @@ private async Task CollectPendingGrants(AnalysisFinding finding, AnalysisContext cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Storage.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Storage.cs index 5315d70df..ca7995f9d 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Storage.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.Storage.cs @@ -31,7 +31,7 @@ ORDER BY avg_read_ms DESC NULLS LAST private async Task CollectFileLatencyBreakdown(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(FileLatencyBreakdownSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -39,8 +39,8 @@ private async Task CollectFileLatencyBreakdown(AnalysisFinding finding, Analysis cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -82,14 +82,14 @@ ORDER BY total_size_mb DESC /// private async Task CollectAutogrowthPercentFiles(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(AutogrowthPercentFilesSql, connection); cmd.Parameters.AddWithValue(context.ServerId); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var database = reader.IsDBNull(0) ? "" : reader.GetString(0); var fileType = reader.IsDBNull(1) ? "" : reader.GetString(1); @@ -125,7 +125,7 @@ ORDER BY (user_object_reserved_mb + internal_object_reserved_mb + version_store_ private async Task CollectTempDbBreakdown(AnalysisFinding finding, AnalysisContext context) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TempDbBreakdownSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -133,8 +133,8 @@ private async Task CollectTempDbBreakdown(AnalysisFinding finding, AnalysisConte cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.cs index 38f4f930a..888cdebe8 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgDrillDownCollector.cs @@ -192,7 +192,7 @@ 0.5 display gate. */ if (finding.DrillDown.Count == 0) finding.DrillDown = null; } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgDrillDownCollector] Drill-down failed for {StoryPath}: {ExceptionType}: {Message}", finding.StoryPath, ex.GetType().Name, ex.Message); diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Activity.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Activity.cs index 1b9d7a180..c412da628 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Activity.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Activity.cs @@ -56,15 +56,15 @@ private async Task CollectBadActorFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(BadActorSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var dbName = reader.IsDBNull(0) ? "" : reader.GetString(0); var queryHash = reader.IsDBNull(1) ? "" : reader.GetString(1); @@ -102,7 +102,10 @@ private async Task CollectBadActorFactsAsync(AnalysisContext context, List }); } } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string ActiveQuerySql = @" @@ -126,15 +129,15 @@ private async Task CollectActiveQueryFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(SessionStatsSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var totalConns = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (totalConns == 0) return; @@ -281,7 +290,10 @@ private async Task CollectSessionFactsAsync(AnalysisContext context, List } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Config.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Config.cs index 23147d86b..bc607b65e 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Config.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgFactCollector.Config.cs @@ -49,7 +49,7 @@ FROM latest /// private async Task CollectServerConfigFactsAsync(AnalysisContext context, List facts) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(ServerConfigSql, connection); cmd.Parameters.AddWithValue(context.ServerId); @@ -59,9 +59,9 @@ private async Task CollectServerConfigFactsAsync(AnalysisContext context, List 0) facts.Add(new Fact { Source = "config", Key = "SERVER_MAJOR_VERSION", Value = majorVersion, ServerId = context.ServerId }); } - catch { /* Columns may not exist yet (pre-migration) */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Columns may not exist yet (pre-migration). An abandonment is NOT swallowed here (#2443). */ + } } public const string DatabaseConfigSql = @" @@ -171,13 +174,13 @@ private async Task CollectDatabaseConfigFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(DatabaseConfigSql, connection); cmd.Parameters.AddWithValue(context.ServerId); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var dbCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (dbCount == 0) return; @@ -213,7 +216,10 @@ private async Task CollectDatabaseConfigFactsAsync(AnalysisContext context, List } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string TraceFlagsSql = @" @@ -235,16 +241,16 @@ private async Task CollectTraceFlagFactsAsync(AnalysisContext context, List(); var flagCount = 0; - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { var flag = Convert.ToInt32(reader.GetValue(0)); metadata[$"TF_{flag}"] = 1; @@ -264,7 +270,10 @@ private async Task CollectTraceFlagFactsAsync(AnalysisContext context, ListThe PLAN_REGRESSION comparison window, in days — how far back a query's "best known" plan + /// may have been observed. + internal const int PlanRegressionWindowDays = 14; + + /// + /// How far BELOW the comparison window the chunk-exclusion bound sits (#2387). `last_execution_time` + /// is the monitored server's clock and `collection_time` is the store's; a monitored server running + /// ahead can report an execution time later than the collection that carried it. A day absorbs any + /// plausible drift while still excluding all but ~15 days of a store that may hold months. + /// + internal const int PlanRegressionSkewMarginDays = 1; + /* PG port: any_value() below is standard SQL:2023, in Postgres since 16 — the product's minimum supported PG is 17, so it stays verbatim (DuckDB and PG agree on its semantics: an arbitrary non-null value from the group). */ @@ -235,6 +253,17 @@ FROM v_query_store_stats WHERE server_id = $1 AND execution_type_desc = 'Regular' AND last_execution_time >= $2 + -- #2387: the SEMANTIC window is last_execution_time above; this is a REDUNDANT bound on the + -- partitioning column so TimescaleDB can exclude chunks. Without it this reads the server's whole + -- history every analysis cycle, decompressing whatever is compressed, per server -- cost scaling with + -- STORE SIZE rather than with the configured window, which is why no VM size fixes it. Redundant to + -- the ANSWER because a row cannot be collected before the execution it reports, so + -- last_execution_time >= X already implies collection_time >= X. $3 carries a skew margin below X + -- because last_execution_time comes from the MONITORED server's clock and collection_time from the + -- store's: a monitored server running fast would otherwise have its newest rows excluded here, which + -- would be a silent under-count and a worse bug than the one this fixes. Do not delete as dead + -- weight -- it is doing all the pruning. + AND collection_time >= $3 ), plan_agg AS ( @@ -350,18 +379,25 @@ ORDER BY regression_factor DESC /// cost >= 2x the best plan that query is known to perform well with. Emits one /// aggregate PLAN_REGRESSION fact. Sourced from Query Store (v_query_store_stats); /// no fact when Query Store is not enabled on the monitored databases. - /// Unlike other collectors this windows on last_execution_time (14-day comparison - /// window), NOT collection_time — see plan note. + /// Windows on last_execution_time (the 14-day comparison window) because a plan regression is about + /// when the query last RAN, not when we happened to collect it. It ALSO carries a redundant + /// collection_time bound so TimescaleDB can exclude chunks — see the note in the SQL (#2387). /// private async Task CollectPlanRegressionFactsAsync(AnalysisContext context, List facts) { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(PlanRegressionSql, connection); cmd.Parameters.AddWithValue(context.ServerId); - cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart.AddDays(-14))); + cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart.AddDays(-PlanRegressionWindowDays))); + + /* #2387: the chunk-exclusion bound, one CLOCK-SKEW MARGIN below the comparison window. Bound as + its own parameter rather than written as "$2 - INTERVAL '1 day'" so the planner compares + against a bare parameter, which is the form runtime chunk exclusion handles most reliably. */ + cmd.Parameters.AddWithValue(AsNaive( + context.TimeRangeStart.AddDays(-(PlanRegressionWindowDays + PlanRegressionSkewMarginDays)))); var offenderCount = 0; var worstFactor = 0.0; @@ -372,8 +408,8 @@ private async Task CollectPlanRegressionFactsAsync(AnalysisContext context, List var worstLatestForced = 0; var worstForceFailures = 0L; - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { // Rows arrive ordered by regression_factor DESC — the first row is the worst offender. if (offenderCount == 0) @@ -417,7 +453,10 @@ private async Task CollectPlanRegressionFactsAsync(AnalysisContext context, List } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string ProcedureStatsSql = @" @@ -440,15 +479,15 @@ private async Task CollectProcedureStatsFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(ProcedureStatsSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var distinctProcs = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); var totalExecs = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -475,7 +514,10 @@ private async Task CollectProcedureStatsFactsAsync(AnalysisContext context, List } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string PlanAdvisorySql = @" @@ -504,15 +546,15 @@ private async Task CollectPlanAdvisoryFactsAsync(AnalysisContext context, List f { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(MemoryStatsSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var totalPhysical = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var bufferPool = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -52,7 +52,10 @@ private async Task CollectMemoryFactsAsync(AnalysisContext context, List f if (targetMemory > 0) facts.Add(new Fact { Source = "memory", Key = "MEMORY_TARGET_MB", Value = targetMemory, ServerId = context.ServerId }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string RunnableTaskStatsSql = @" @@ -82,15 +85,15 @@ private async Task CollectRunnableTaskFactsAsync(AnalysisContext context, List(); var totalMb = 0.0; var clerkCount = 0; - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { var clerkType = reader.GetString(0); var memoryMb = Convert.ToDouble(reader.GetValue(1)); @@ -223,7 +232,10 @@ private async Task CollectMemoryClerkFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(PerfmonSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var counterName = reader.GetString(0); var cntrValue = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -363,7 +378,10 @@ private async Task CollectPerfmonFactsAsync(AnalysisContext context, List }); } } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string PlanCacheStatsSql = @" @@ -395,14 +413,14 @@ private async Task CollectPlanCacheFactsAsync(AnalysisContext context, List 0) facts.Add(new Fact { Source = "config", Key = "DATABASE_TOTAL_SIZE_MB", Value = totalSize, ServerId = context.ServerId }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string IoLatencySql = @" @@ -73,15 +76,15 @@ private async Task CollectIoLatencyFactsAsync(AnalysisContext context, List= $2 @@ -148,15 +159,15 @@ private async Task CollectTempDbFactsAsync(AnalysisContext context, List f { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(TempDbSql, connection); cmd.Parameters.AddWithValue(context.ServerId); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); cmd.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var maxReserved = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var maxUserObj = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -164,11 +175,18 @@ private async Task CollectTempDbFactsAsync(AnalysisContext context, List f var maxVersionStore = reader.IsDBNull(3) ? 0.0 : Convert.ToDouble(reader.GetValue(3)); var minUnallocated = reader.IsDBNull(4) ? 0.0 : Convert.ToDouble(reader.GetValue(4)); var avgReserved = reader.IsDBNull(5) ? 0.0 : Convert.ToDouble(reader.GetValue(5)); + var maxSizeMb = reader.IsDBNull(6) ? 0.0 : Convert.ToDouble(reader.GetValue(6)); if (maxReserved <= 0) return; - // TempDB usage as fraction of total space (reserved + unallocated) - var totalSpace = maxReserved + minUnallocated; + /* #2515: against the CEILING where tempdb has one, against the current allocation where it does + not (-1 unlimited, or 0 for a window collected before the ceiling was captured). reserved + + unallocated is the files AS ALLOCATED, so on its own this fraction measures distance to the + next autogrow — which reads as 96% full on an Azure SQL Database target holding one temp + table. TempDbSpaceInfo.CapacityMb is the alert's twin of this, and the two must agree or + analyze_server and the pager describe the same server differently. */ + var allocatedMb = maxReserved + minUnallocated; + var totalSpace = maxSizeMb > 0 ? Math.Max(maxSizeMb, allocatedMb) : allocatedMb; var usageFraction = totalSpace > 0 ? maxReserved / totalSpace : 0; facts.Add(new Fact @@ -185,11 +203,17 @@ private async Task CollectTempDbFactsAsync(AnalysisContext context, List f ["max_internal_object_mb"] = maxInternalObj, ["max_version_store_mb"] = maxVersionStore, ["min_unallocated_mb"] = minUnallocated, + /* Added, never redefined — the existing keys are a consumer surface (FactAdvice reads + them by name). -1 unlimited, 0 not measured, positive = the ROWS ceiling in MB. */ + ["max_size_mb"] = maxSizeMb, ["usage_fraction"] = usageFraction } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string FileAutogrowthSql = @" @@ -220,13 +244,13 @@ private async Task CollectFileAutogrowthFactsAsync(AnalysisContext context, List { try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var cmd = new NpgsqlCommand(FileAutogrowthSql, connection); cmd.Parameters.AddWithValue(context.ServerId); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var fileCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (fileCount == 0) return; @@ -246,7 +270,10 @@ private async Task CollectFileAutogrowthFactsAsync(AnalysisContext context, List } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } public const string DiskSpaceSql = @" @@ -273,14 +300,14 @@ private async Task CollectDiskSpaceFactsAsync(AnalysisContext context, List private async Task CollectWaitStatsFactsAsync(AnalysisContext context, List facts) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var command = new NpgsqlCommand(WaitStatsSql, connection); command.Parameters.AddWithValue(context.ServerId); command.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); command.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await command.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var waitType = reader.GetString(0); var waitingTasks = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -95,15 +95,15 @@ FROM blocked_process_reports /// private async Task CollectBlockingFactsAsync(AnalysisContext context, List facts) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var command = new NpgsqlCommand(BlockingSql, connection); command.Parameters.AddWithValue(context.ServerId); command.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); command.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await command.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var eventCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (eventCount <= 0) return; @@ -149,15 +149,15 @@ FROM deadlocks /// private async Task CollectDeadlockFactsAsync(AnalysisContext context, List facts) { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var command = new NpgsqlCommand(DeadlocksSql, connection); command.Parameters.AddWithValue(context.ServerId); command.Parameters.AddWithValue(AsNaive(context.TimeRangeStart)); command.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); - using var reader = await command.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var deadlockCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (deadlockCount <= 0) return; @@ -211,7 +211,7 @@ private async Task CollectBlockingChainFactsAsync(AnalysisContext context, List< try { - await using var connection = await _postgres.OpenConnectionAsync(); + await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); using var command = new NpgsqlCommand(BlockingChainSql, connection); command.Parameters.AddWithValue(context.ServerId); @@ -219,16 +219,17 @@ private async Task CollectBlockingChainFactsAsync(AnalysisContext context, List< command.Parameters.AddWithValue(AsNaive(context.TimeRangeEnd)); var rows = new List(); - using (var reader = await command.ExecuteReaderAsync()) + using (var reader = await command.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) rows.Add(PgBlockingPairRowQuery.Read(reader)); } // Always-on DMV blocking snapshot fallback (works when the blocked-process-report XE is empty, // e.g. AWS RDS). Merge BEFORE the empty check so DMV-only blocking still produces facts. await PgBlockingPairRowQuery.AppendDmvSnapshotRowsAsync( - connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd); + connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd, + context.CancellationToken); if (rows.Count == 0) return; @@ -259,7 +260,10 @@ await PgBlockingPairRowQuery.AppendDmvSnapshotRowsAsync( } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgFindingStore.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgFindingStore.cs index acb2995e7..c17201ee4 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgFindingStore.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgFindingStore.cs @@ -8,6 +8,7 @@ using System; using System.Collections.Generic; +using System.Threading; using System.Threading.Tasks; using Microsoft.Extensions.Logging; using Npgsql; @@ -63,9 +64,16 @@ public sealed record MutedStory( /// /// /// -/// Error discipline mirrors the Dashboard twin: no public method throws — writes log and -/// degrade, reads log and return empty. The mute-hash read fails OPEN (an unreadable mute -/// registry lets findings through rather than suppressing them), exactly like both twins. +/// Error discipline mirrors the Dashboard twin: reads log and return empty, and the mute-hash read +/// fails OPEN (an unreadable mute registry lets findings through rather than suppressing them), +/// exactly like both twins. is the ONE exception, and #2448 is +/// why: "log and degrade" needs something to degrade TO, and once the batch became all-or-nothing +/// there is nothing between "all persisted" and "none persisted". Swallowing the second would let +/// DarlingAnalysisService set LastAnalysisTime, fire AnalysisCompleted and log "Analysis +/// complete — N finding(s)" over a store holding none of them, which is #2448's own misreading +/// moved one layer out and made louder. So it logs the detail only it can know and rethrows, and +/// the pass reports itself failed — which is what the Lite twin has always done, by never catching +/// at all. /// /// public sealed class PgFindingStore @@ -91,6 +99,16 @@ INSERT INTO analysis_findings incident_id, remediation_action_json, drill_down_json) VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13, $14, $15, $16, $17, $18, $19, $20, $21)"; + /* + #2506 added the UPPER bound ($3). Without it the read had a start and no end, so an as_of + anchor could only ever move the window's start EARLIER and every anchored read would still + return everything up to now — the anchor validated, the caller told the window had moved, and + the answer unchanged. That is the exact defect this convention exists to prevent, so the bound + is part of the read rather than something the caller filters afterwards. + + It binds ONLY when the caller anchored; unanchored, $3 is NoUpperBound and the read is the + half-open window it has always been. See that field for why "now" is the wrong default. + */ public const string GetRecentFindingsSql = @" SELECT finding_id, analysis_time, server_id, server_name, database_name, time_range_start, time_range_end, severity, confidence, category, @@ -100,8 +118,9 @@ INSERT INTO analysis_findings FROM analysis_findings WHERE server_id = $1 AND analysis_time >= $2 +AND analysis_time <= $3 ORDER BY analysis_time DESC, severity DESC -LIMIT $3"; +LIMIT $4"; public const string GetLatestFindingsSql = @" SELECT finding_id, analysis_time, server_id, server_name, database_name, @@ -167,7 +186,7 @@ public async Task> FilterMutedFindingsAsync( try { await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); - var mutedHashes = await GetMutedHashesAsync(connection, context.ServerId); + var mutedHashes = await GetMutedHashesAsync(connection, context.ServerId, context.CancellationToken); foreach (var story in stories) { @@ -208,7 +227,7 @@ public async Task> FilterMutedFindingsAsync( }); } } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { _logger?.LogError("[PgFindingStore] FilterMutedFindingsAsync failed: {Message}", ex.Message); } @@ -218,11 +237,52 @@ public async Task> FilterMutedFindingsAsync( /// /// Inserts the (already mute-filtered, enriched, and action-attached) findings in one - /// batched pass on a single connection. Each row persists its BUILT + /// batched pass on a single connection, inside ONE transaction. Each row persists its BUILT /// as remediation_action_json via the /// shared , so a Darling finding's persisted action /// round-trips byte-identically to a Dashboard one. Returns the same list for caller /// convenience; the in-memory findings are unchanged. + /// + /// #2448: the transaction is what makes this method's promise true, and it is the whole + /// reason for the shape. A finding set is one indivisible statement about a server: every row + /// shares one analysis_time, and keys + /// on MAX(analysis_time). So a batch that lands four of forty rows before the store + /// faults does NOT read as truncated — it reads as a complete analysis that found four + /// problems, and the server looks HEALTHIER for the store having failed. Rolling the batch back + /// instead leaves the PREVIOUS pass as the newest complete set: stale, stamped with its own + /// older analysis_time, and incapable of misleading anyone. + /// + /// This replaces per-row failure isolation, which was deliberate and is deliberately gone. + /// It could not have survived the transaction on both SKUs anyway — PostgreSQL refuses every + /// later statement in a transaction once one fails (25P02) and only SAVEPOINT escapes + /// that, which DuckDB 1.5.5 does not parse, so keeping it here alone would be exactly the + /// Lite/Darling drift this store has spent its life removing. It also should not survive on its + /// own merits: a batch that silently drops row 5 and commits the other 39 is this same defect at + /// row granularity, a set claiming 39 problems when the analysis found 40. One failure now ends + /// the batch and costs ONE log line naming the count — #2299's shape — instead of a line per + /// remaining row and a set nothing marks as partial. + /// + /// Measured rather than assumed, because "the batch is small" is the load-bearing claim: a + /// busy production server persists ~10 rows per pass, as small INSERTs against the LOCAL managed + /// store on loopback. Worth knowing when reading the loop below: CommitAsync on an + /// already-aborted transaction RETURNS NORMALLY on both Npgsql and DuckDB.NET while committing + /// nothing, so reaching the commit is not evidence that anything was written. That is the other + /// reason the row write must not swallow. + /// + /// It is also the one write here that THROWS, against the class's own no-throw discipline, + /// and that follows from the transaction rather than sitting beside it. A swallowed rollback returns + /// the same list a full success returns, so the caller cannot tell them apart and announces a + /// complete analysis for a set the store does not have. Before the transaction that line was only a + /// little wrong — most rows had landed — and now it would be entirely wrong, which is the same + /// defect this method exists to remove, one layer further out. The single ERROR line below carries + /// the part only this method knows (which row, or that it was the commit); the pass adds its own + /// outcome line and reports itself failed. + /// + /// #2443: the connection open is still the LAST cancellation point on this pass, unchanged + /// by the above. Past it the batch runs to completion, and cancelling before the first row costs + /// this cycle's findings and says so in the line the pass already logs. This is the same call + /// DarlingAnalysisService made one layer up in #2299 — "the post-enrichment tail carries + /// no check" — restated at the write it protects. /// public async Task> InsertFindingsAsync( List findings, AnalysisContext context) @@ -237,18 +297,56 @@ public async Task> InsertFindingsAsync( return findings; } + var row = 0; + var everyRowAccepted = false; + try { await using var connection = await _postgres.OpenConnectionAsync(context.CancellationToken); + /* #2448: one transaction for the whole set — the batch commits complete or not at all. + No token check between rows: see the note above. */ + await using var transaction = await connection.BeginTransactionAsync(); + foreach (var finding in findings) { - await InsertFindingAsync(connection, finding); + row++; + await InsertFindingAsync(connection, transaction, finding); } + + everyRowAccepted = true; + await transaction.CommitAsync(); } - catch (Exception ex) when (!AnalysisShutdown.IsShutdownAbandon(ex, context.CancellationToken)) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, context.CancellationToken)) { - _logger?.LogError("[PgFindingStore] InsertFindingsAsync failed: {Message}", ex.Message); + /* The ONE line the batch is allowed to cost, and it states the loss rather than the + mechanism: nothing was persisted, and that is the deliberate outcome. Naming the row + is what turns "how often does this actually happen" into something the log can + answer — #2448 asked for that measurement and this is where it comes from. + + Which is exactly why the commit gets its OWN line rather than sharing this one. After + the loop `row` sits at findings.Count, so a fault in CommitAsync — a blip, a + deadlock at commit, the store out of disk — would report "failed at row N of N" and + name the last finding as the bad one when every row had in fact been accepted and + only the commit failed. A diagnostic that exists to be counted must not miscount the + one case it cannot see from the row number. */ + if (everyRowAccepted) + { + _logger?.LogError( + "[PgFindingStore] InsertFindingsAsync had all {Count} row(s) accepted and then failed to COMMIT them, so the batch was rolled back — this analysis persisted NO findings, deliberately: a partial set would have read as a complete analysis that found fewer problems. {Message}", + findings.Count, ex.Message); + } + else + { + _logger?.LogError( + "[PgFindingStore] InsertFindingsAsync failed at row {Row} of {Count} and the batch was rolled back — this analysis persisted NO findings, deliberately: a partial set would have read as a complete analysis that found fewer problems. {Message}", + row, findings.Count, ex.Message); + } + + /* And it must not be swallowed: see the note above. The caller announces a completed + analysis on the strength of this returning, so eating a total rollback would move the + #2448 misreading up a layer instead of removing it. */ + throw; } return findings; @@ -274,18 +372,30 @@ public async Task> SaveFindingsAsync( /// /// Returns the most recent findings for a server within the given time range, newest and /// most severe first, including each finding's persisted remediation action. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task> GetRecentFindingsAsync( - int serverId, int hoursBack = 24, int limit = 100) + int serverId, int hoursBack = 24, int limit = 100, DateTime? asOfUtc = null) { var findings = new List(); try { + /* #2506: the window's END, from which the START is measured. Null is "now" — the pre-#2506 + read exactly. Made naive-UTC like every other bound here, because analysis_time is a + naive-UTC column and a Kind=Utc parameter would be inferred as timestamptz and silently + zone-shifted. */ + var windowEnd = DateTime.SpecifyKind(asOfUtc ?? DateTime.UtcNow, DateTimeKind.Unspecified); + await using var connection = await _postgres.OpenConnectionAsync(); using var command = new NpgsqlCommand(GetRecentFindingsSql, connection); command.Parameters.AddWithValue(serverId); - command.Parameters.AddWithValue(NaiveUtcNow().AddHours(-hoursBack)); + command.Parameters.AddWithValue(windowEnd.AddHours(-hoursBack)); + command.Parameters.AddWithValue(asOfUtc is null ? NoUpperBound : windowEnd); command.Parameters.AddWithValue(limit); using var reader = await command.ExecuteReaderAsync(); @@ -306,6 +416,11 @@ public async Task> GetRecentFindingsAsync( /// Returns the latest analysis run's findings for a server (most recent analysis_time), /// most severe first. Unlike the Dashboard twin this read also returns /// remediation_action_json — both reads share one column list and one reader. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task> GetLatestFindingsAsync(int serverId) { @@ -333,6 +448,11 @@ public async Task> GetLatestFindingsAsync(int serverId) /// /// Mutes a story pattern so it won't appear in future analysis runs. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task MuteStoryAsync(int serverId, string storyPathHash, string storyPath, string? reason = null) { @@ -359,6 +479,11 @@ public async Task MuteStoryAsync(int serverId, string storyPathHash, string stor /// /// Unmutes a story pattern (Dashboard-twin surface). + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task UnmuteStoryAsync(long muteId) { @@ -385,6 +510,11 @@ public async Task UnmuteStoryAsync(long muteId) /// too ( filters both), so an all-servers mute is visible to every /// server. A NULL/0 (global) row here is flagged muted but left un-unmutable from the per-server /// viewer, since deleting it would unmute the pattern everywhere. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task> GetMutedStoriesAsync(int serverId) { @@ -418,6 +548,11 @@ public async Task> GetMutedStoriesAsync(int serverId) /// /// Cleans up old findings beyond the retention period. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. /// public async Task CleanupOldFindingsAsync(int retentionDays = 30) { @@ -440,7 +575,8 @@ public async Task CleanupOldFindingsAsync(int retentionDays = 30) /// twins: an unreadable mute registry returns the hashes read so far (usually empty) and /// findings flow through unfiltered rather than being suppressed. /// - private async Task> GetMutedHashesAsync(NpgsqlConnection connection, int serverId) + private async Task> GetMutedHashesAsync( + NpgsqlConnection connection, int serverId, CancellationToken cancellationToken) { var hashes = new HashSet(); @@ -449,14 +585,18 @@ private async Task> GetMutedHashesAsync(NpgsqlConnection connect using var command = new NpgsqlCommand(GetMutedHashesSql, connection); command.Parameters.AddWithValue(serverId); - using var reader = await command.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) { hashes.Add(reader.GetString(0)); } } - catch (Exception ex) + catch (Exception ex) when (!AnalysisShutdown.IsExpectedAbandon(ex, cancellationToken)) { + /* #2443: fail-open is right for an unreadable registry, but NOT for an abandonment — + "the mutes could not be read" and "we stopped reading" are different answers, and + swallowing the second would let the pass go on to enrich and persist an unfiltered + finding set under a token that has already fired. */ _logger?.LogError("[PgFindingStore] GetMutedHashesAsync failed: {Message}", ex.Message); } @@ -464,59 +604,70 @@ private async Task> GetMutedHashesAsync(NpgsqlConnection connect } /// - /// Inserts one finding on an already-open connection (the caller owns it, so a batch - /// shares a single connection). Per-row failure isolation like the Dashboard twin: one - /// bad row logs and the batch continues. + /// Inserts one finding on an already-open connection, enlisted in the batch's transaction (the + /// caller owns both, so a batch shares a single connection and a single transaction). + /// + /// #2448: this throws rather than logging and continuing, which is the reverse of the + /// Dashboard twin's per-row isolation and is the point. Once one row has failed, PostgreSQL + /// refuses every later statement in the transaction (25P02), so "continue" would mean N-k more + /// ERROR lines for one event and a CommitAsync that returns normally having written + /// nothing. Letting it out gives the single line and the + /// rollback. The isolation is barely reachable here in any case — the table carries no primary + /// key, no CHECK and no foreign key, and every NOT NULL column maps to a non-nullable property + /// with a default — which leaves data-shape failures such as a NUL byte in a text column + /// (22021) as the realistic per-row fault, and one of those in a batch of ten is exactly the + /// case where publishing the other nine as a complete analysis is the wrong answer. + /// + /// #2443 exempt: this write deliberately takes no token. Cancelling inside a single-row + /// INSERT buys nothing — the row is milliseconds of work on loopback — and costs the one thing + /// worth having: a definite answer about whether it committed. Npgsql's cancel is a request to + /// the server, so a cancelled ExecuteNonQueryAsync can leave a row that did land. + /// carries the full reasoning and the abandonment point that + /// replaces this one. /// - private async Task InsertFindingAsync(NpgsqlConnection connection, AnalysisFinding finding) + private async Task InsertFindingAsync( + NpgsqlConnection connection, NpgsqlTransaction transaction, AnalysisFinding finding) { - try + using var command = new NpgsqlCommand(InsertFindingSql, connection, transaction); + command.Parameters.AddWithValue(finding.FindingId); + command.Parameters.AddWithValue(AsNaive(finding.AnalysisTime)); + command.Parameters.AddWithValue(finding.ServerId); + command.Parameters.AddWithValue(finding.ServerName); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = (object?)finding.DatabaseName ?? DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Timestamp, Value = finding.TimeRangeStart is { } rangeStart ? AsNaive(rangeStart) : (object)DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Timestamp, Value = finding.TimeRangeEnd is { } rangeEnd ? AsNaive(rangeEnd) : (object)DBNull.Value }); + command.Parameters.AddWithValue(finding.Severity); + command.Parameters.AddWithValue(finding.Confidence); + command.Parameters.AddWithValue(finding.Category); + command.Parameters.AddWithValue(finding.StoryPath); + command.Parameters.AddWithValue(finding.StoryPathHash); + command.Parameters.AddWithValue(finding.StoryText); + command.Parameters.AddWithValue(finding.RootFactKey); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Double, Value = (object?)finding.RootFactValue ?? DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = (object?)finding.LeafFactKey ?? DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Double, Value = (object?)finding.LeafFactValue ?? DBNull.Value }); + command.Parameters.AddWithValue(finding.FactCount); + command.Parameters.Add(new NpgsqlParameter { - using var command = new NpgsqlCommand(InsertFindingSql, connection); - command.Parameters.AddWithValue(finding.FindingId); - command.Parameters.AddWithValue(AsNaive(finding.AnalysisTime)); - command.Parameters.AddWithValue(finding.ServerId); - command.Parameters.AddWithValue(finding.ServerName); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = (object?)finding.DatabaseName ?? DBNull.Value }); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Timestamp, Value = finding.TimeRangeStart is { } rangeStart ? AsNaive(rangeStart) : (object)DBNull.Value }); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Timestamp, Value = finding.TimeRangeEnd is { } rangeEnd ? AsNaive(rangeEnd) : (object)DBNull.Value }); - command.Parameters.AddWithValue(finding.Severity); - command.Parameters.AddWithValue(finding.Confidence); - command.Parameters.AddWithValue(finding.Category); - command.Parameters.AddWithValue(finding.StoryPath); - command.Parameters.AddWithValue(finding.StoryPathHash); - command.Parameters.AddWithValue(finding.StoryText); - command.Parameters.AddWithValue(finding.RootFactKey); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Double, Value = (object?)finding.RootFactValue ?? DBNull.Value }); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = (object?)finding.LeafFactKey ?? DBNull.Value }); - command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Double, Value = (object?)finding.LeafFactValue ?? DBNull.Value }); - command.Parameters.AddWithValue(finding.FactCount); - command.Parameters.Add(new NpgsqlParameter - { - NpgsqlDbType = NpgsqlDbType.Text, - Value = string.IsNullOrEmpty(finding.IncidentId) ? (object)DBNull.Value : finding.IncidentId - }); - /* D2: persist the BUILT action (mirrors the alert path's ContextJson) so the - Recommendations reader can drive Apply + consent from a stored finding. */ - command.Parameters.Add(new NpgsqlParameter - { - NpgsqlDbType = NpgsqlDbType.Text, - Value = (object?)AlertContextSerializer.SerializeAction(finding.Remediation) ?? DBNull.Value - }); - /* #2060: persist the CAPPED drill-down beside the built action — same D2 rationale - (the evidence rows exist only on the write path), same degrade-to-NULL discipline. */ - command.Parameters.Add(new NpgsqlParameter - { - NpgsqlDbType = NpgsqlDbType.Text, - Value = (object?)DrillDownSerializer.Serialize(finding.DrillDown) ?? DBNull.Value - }); - - await command.ExecuteNonQueryAsync(); - } - catch (Exception ex) + NpgsqlDbType = NpgsqlDbType.Text, + Value = string.IsNullOrEmpty(finding.IncidentId) ? (object)DBNull.Value : finding.IncidentId + }); + /* D2: persist the BUILT action (mirrors the alert path's ContextJson) so the + Recommendations reader can drive Apply + consent from a stored finding. */ + command.Parameters.Add(new NpgsqlParameter { - _logger?.LogError("[PgFindingStore] InsertFindingAsync failed: {Message}", ex.Message); - } + NpgsqlDbType = NpgsqlDbType.Text, + Value = (object?)AlertContextSerializer.SerializeAction(finding.Remediation) ?? DBNull.Value + }); + /* #2060: persist the CAPPED drill-down beside the built action — same D2 rationale + (the evidence rows exist only on the write path), same degrade-to-NULL discipline. */ + command.Parameters.Add(new NpgsqlParameter + { + NpgsqlDbType = NpgsqlDbType.Text, + Value = (object?)DrillDownSerializer.Serialize(finding.DrillDown) ?? DBNull.Value + }); + + await command.ExecuteNonQueryAsync(); } /// @@ -560,6 +711,21 @@ private static AnalysisFinding ReadFinding(NpgsqlDataReader reader) private static DateTime NaiveUtcNow() => DateTime.SpecifyKind(DateTime.UtcNow, DateTimeKind.Unspecified); + /// + /// What 's upper bound is when the caller did NOT anchor: an + /// instant no row can reach, i.e. no bound at all. + /// + /// Why not "now". Two reasons, and the second is the one that would have hurt. First, + /// #2495's promise is that a caller sending only hours_back gets byte-for-byte the window it + /// always got, and this read has been half-open for its whole life. Second, analysis_time is + /// stamped by the WRITER's clock and would be filtered by the READER's; those are the same host + /// today, and the day they are not, a bounded default read would intermittently drop the newest run + /// — a findings read that "sometimes misses the analysis that just finished", with nothing in it to + /// point at a clock. An anchored read has a caller-supplied end and neither problem. + /// + private static readonly DateTime NoUpperBound = + new(9999, 12, 31, 23, 59, 59, DateTimeKind.Unspecified); + /// Kind-Unspecified for writes — Npgsql 6+ rejects Kind-Utc against timestamp. private static DateTime AsNaive(DateTime value) => DateTime.SpecifyKind(value, DateTimeKind.Unspecified); diff --git a/Darling/PerformanceMonitor.Darling.Analysis/PgPlanFetcher.cs b/Darling/PerformanceMonitor.Darling.Analysis/PgPlanFetcher.cs index 8818d1d5d..4c2607eec 100644 --- a/Darling/PerformanceMonitor.Darling.Analysis/PgPlanFetcher.cs +++ b/Darling/PerformanceMonitor.Darling.Analysis/PgPlanFetcher.cs @@ -59,7 +59,7 @@ public PgPlanFetcher(Func connectionStringResolver, ILogger? logge _logger = logger; } - public async Task FetchPlanXmlAsync(int serverId, string planHandle) + public async Task FetchPlanXmlAsync(int serverId, string planHandle, CancellationToken cancellationToken) { if (string.IsNullOrEmpty(planHandle)) { @@ -82,13 +82,13 @@ public PgPlanFetcher(Func connectionStringResolver, ILogger? logge }; await using var connection = new SqlConnection(builder.ConnectionString); - await connection.OpenAsync(); + await connection.OpenAsync(cancellationToken); await using var command = new SqlCommand(PlanQuery, connection); command.CommandTimeout = 15; command.Parameters.AddWithValue("@plan_handle", planHandle); - var result = await command.ExecuteScalarAsync(); + var result = await command.ExecuteScalarAsync(cancellationToken); if (result == null || result is DBNull) { return null; @@ -96,8 +96,12 @@ public PgPlanFetcher(Func connectionStringResolver, ILogger? logge return result.ToString(); } - catch (Exception ex) + catch (Exception ex) when (ex is not OperationCanceledException) { + /* #2443: an abandoned fetch is not a plan that failed to come back. Degrading it to null + here would log an ERROR for work we called off AND let the pass carry on enriching + against a token that has already fired — the sibling by-sql_handle fetch has excluded + cancellation from this arm since it was written, and this one now agrees. */ _logger?.LogError("[PgPlanFetcher] Failed to fetch plan for handle {PlanHandle}: {Message}", planHandle, ex.Message); return null; } diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs index 16fed51e7..c068f4ec7 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingAlertReadAdapter.cs @@ -654,7 +654,8 @@ public async Task> GetPvsPressureAsync( internal_object_reserved_mb, version_store_reserved_mb, top_session_tempdb_mb, - top_session_id + top_session_id, + max_size_mb FROM tempdb_stats WHERE server_id = $1 ORDER BY collection_time DESC @@ -680,7 +681,11 @@ ORDER BY collection_time DESC InternalObjectReservedMb = reader.IsDBNull(3) ? 0 : ToDouble(reader.GetValue(3)), VersionStoreReservedMb = reader.IsDBNull(4) ? 0 : ToDouble(reader.GetValue(4)), TopConsumerMb = reader.IsDBNull(5) ? 0 : ToDouble(reader.GetValue(5)), - TopConsumerSessionId = reader.IsDBNull(6) ? 0 : reader.GetInt32(6) + TopConsumerSessionId = reader.IsDBNull(6) ? 0 : reader.GetInt32(6), + /* NULL on every row collected before the V81 rung, and 0 is what "no ceiling measured" + is spelled as — so history keeps reporting the percentage it always did rather than + dividing by a zero cap. */ + MaxSizeMb = reader.IsDBNull(7) ? 0 : ToDouble(reader.GetValue(7)) }; } diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingCliCommands.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingCliCommands.cs index ae595e4fe..7242fae4d 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingCliCommands.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingCliCommands.cs @@ -76,6 +76,16 @@ public static bool IsValidateConfigVerb(string arg) => public static bool IsPrintViewerConnectionVerb(string arg) => string.Equals(arg, "--print-viewer-connection", StringComparison.OrdinalIgnoreCase); + /// The verb handles — reprint the MCP bearer token (#2479, item 2). + public static bool IsPrintMcpTokenVerb(string arg) => + string.Equals(arg, "--print-mcp-token", StringComparison.OrdinalIgnoreCase); + + /// The verb handles — reprint the web dashboard access token. + /// The issue asked for the MCP half; both tokens have the identical unrecoverable-loss problem and the + /// identical DPAPI shape, so shipping one and not the other would leave half a fix behind. + public static bool IsPrintWebTokenVerb(string arg) => + string.Equals(arg, "--print-web-token", StringComparison.OrdinalIgnoreCase); + /// The verb handles — write a COMPLETE viewer darling.json + /// server.crt + README.txt an operator copies to the viewer machine as-is (#1953). public static bool IsExportViewerConfigVerb(string arg) => @@ -158,6 +168,8 @@ public static bool IsKnownVerb(string arg) => IsEncryptPasswordVerb(arg) || IsValidateConfigVerb(arg) || IsPrintViewerConnectionVerb(arg) + || IsPrintMcpTokenVerb(arg) + || IsPrintWebTokenVerb(arg) || IsExportViewerConfigVerb(arg) || IsConfigureNetworkVerb(arg) || IsConfigureFirewallVerb(arg) @@ -232,6 +244,8 @@ public static string UsageText() => " PerformanceMonitor.Darling.Service.exe --test-connection Validate darling.json and probe every configured server." + Environment.NewLine + " PerformanceMonitor.Darling.Service.exe --encrypt-password Encrypt a SQL-auth password for darling.json (reads stdin)." + Environment.NewLine + " PerformanceMonitor.Darling.Service.exe --print-viewer-connection Print a remote-viewer connection string (managed store)." + Environment.NewLine + + " PerformanceMonitor.Darling.Service.exe --print-mcp-token Reprint the MCP bearer token from darling.json (run elevated; writes a LIVE token to stdout)." + Environment.NewLine + + " PerformanceMonitor.Darling.Service.exe --print-web-token Reprint the web dashboard access token from darling.json (run elevated; writes a LIVE token to stdout)." + Environment.NewLine + " PerformanceMonitor.Darling.Service.exe --export-viewer-config [dir] [--config ] Write a ready-to-copy viewer folder (darling.json + server.crt + README.txt)." + Environment.NewLine + " PerformanceMonitor.Darling.Service.exe --configure-network Interactive LAN-exposure wizard." + Environment.NewLine + " PerformanceMonitor.Darling.Service.exe --configure-firewall Create/remove the scoped firewall rules to match darling.json (run elevated)." + Environment.NewLine + @@ -309,10 +323,10 @@ public static string FormatProbeLine(string serverName, ConnectionProbeResult pr } /// - /// Prints a paste-ready remote-viewer connection string and the server TLS certificate for the opt-in store - /// network endpoint (darling-network-endpoints D8). It DPAPI-decrypts the credential of the role - /// postgres.network.role names (default viewer, read-only) and reads the generated - /// server.crt, so it must run ON the managed store's host under an account that can decrypt them — + /// Prints a paste-ready remote-viewer connection string PER ADMITTED ROLE (#2665) and the server TLS + /// certificate for the opt-in store network endpoint (darling-network-endpoints D8). It DPAPI-decrypts the + /// credential of every role postgres.network.role names (default viewer, read-only) and reads + /// the generated server.crt, so it must run ON the managed store's host under an account that can decrypt them — /// hence Windows-only (the caller is OperatingSystem.IsWindows()-guarded, mirroring /// --encrypt-password). The operator pastes the string into the VIEWER machine's darling.json /// (postgres.managed = false, into postgres.connectionString, consumed verbatim — no viewer @@ -337,9 +351,9 @@ public static async Task PrintViewerConnectionAsync( return 1; } - var handoff = await ResolveViewerHandoffAsync( + var handoffs = await ResolveViewerHandoffsAsync( config, "--print-viewer-connection", "print", error, cancellationToken); - if (handoff is null) + if (handoffs is null) { return 1; } @@ -348,8 +362,6 @@ public static async Task PrintViewerConnectionAsync( VIEWER machine (a bare filename resolves against the folder holding the viewer's darling.json — #1970; an absolute path also works). Kept as a literal so the printed string is paste-ready. */ const string clientCertificatePath = ViewerClientCertificateFileName; - var connectionString = BuildViewerConnectionString( - handoff.Host, handoff.Port, handoff.Role, handoff.Password, clientCertificatePath); /* Read the cert BEFORE anything reaches STDOUT: every STDERR line — including the missing-cert NOTE — must be emitted ahead of the payload (#1953 item 3). The field report watched the live password scroll @@ -358,27 +370,43 @@ must be emitted ahead of the payload (#1953 item 3). The field report watched th carries the fixed chain shape — that is what verify-full's Root Certificate must anchor on. A legacy store has no root.crt, and its single self-signed server.crt remains the right (if Windows-hostile) thing to print. */ - var distributableCertPath = DarlingManagedPostgres.RootCertificatePathFor(handoff.CertificatePath); + var certificateSource = handoffs[0].CertificatePath; + var distributableCertPath = DarlingManagedPostgres.RootCertificatePathFor(certificateSource); if (!File.Exists(distributableCertPath)) { - distributableCertPath = handoff.CertificatePath; + distributableCertPath = certificateSource; } var certificate = File.Exists(distributableCertPath) ? (await File.ReadAllTextAsync(distributableCertPath, cancellationToken)).Trim() : null; + var admitsAdmin = handoffs.Any(h => string.Equals(h.Role, "admin", StringComparison.Ordinal)); + var roleList = string.Join("', '", handoffs.Select(h => h.Role)); + var many = handoffs.Count > 1; + /* Guidance + the live-secret warning go to STDERR, so redirecting STDOUT to a file or the clipboard - captures the connection string + cert WITHOUT swallowing the warning (D8). */ + captures the connection string + cert WITHOUT swallowing the warning (D8). ALL of it is emitted + before the first STDOUT byte, including for a multi-role exposure — interleaving a warning between + two printed strings would put half the advice after a password had already scrolled past (#1953 + item 3). */ error.WriteLine(); error.WriteLine( - $"WARNING: the connection string below contains a LIVE database password (the '{handoff.Role}' role), written " + + $"WARNING: the connection {(many ? "strings" : "string")} below {(many ? "contain" : "contains")} a LIVE " + + $"database password (the '{roleList}' {(many ? "roles" : "role")}), written " + "to STDOUT. Redirect it to an ACL'd file or pipe it to the clipboard; do not leave it in shell " + "scrollback, CI logs, or a screenshare."); error.WriteLine(" Example (file): PerformanceMonitor.Darling.Service.exe --print-viewer-connection > viewer-connection.txt"); error.WriteLine(" Example (clipboard): PerformanceMonitor.Darling.Service.exe --print-viewer-connection | clip"); error.WriteLine(" Example (no paste): PerformanceMonitor.Darling.Service.exe --export-viewer-config (writes the whole viewer folder for you)"); - if (string.Equals(handoff.Role, "admin", StringComparison.Ordinal)) + if (many) + { + error.WriteLine( + $" NOTE: postgres.network.role admits {handoffs.Count} roles (#2665), so one string is printed per role, " + + "each a DIFFERENT credential. Give each seat only the one it needs — they are not interchangeable."); + } + + if (admitsAdmin) { error.WriteLine( " NOTE: 'admin' is a WRITE credential holding the config-table pivot surface. Prefer the default " + @@ -394,17 +422,24 @@ must be emitted ahead of the payload (#1953 item 3). The field report watched th else { error.WriteLine( - $"NOTE: the server TLS certificate ({handoff.CertificatePath}) does not exist yet — the service generates it " + + $"NOTE: the server TLS certificate ({certificateSource}) does not exist yet — the service generates it " + "on its first managed start with postgres.network exposed. Enable postgres.network, restart the " + "service, then re-run this command to emit the cert for verify-full."); } error.WriteLine(); - output.WriteLine( - "# Paste into the viewer machine's darling.json -> postgres.connectionString (with postgres.managed = false):"); - output.WriteLine(connectionString); - output.WriteLine(); + /* One paste-ready string per admitted role, each labelled with the role it authenticates as — the + strings differ only in Username= and Password=, so an unlabelled pair is impossible to tell apart + after the fact and the reader would have to guess which seat to hand where. */ + foreach (var handoff in handoffs) + { + output.WriteLine( + $"# Paste into the viewer machine's darling.json -> postgres.connectionString (with postgres.managed = false) — the '{handoff.Role}' seat:"); + output.WriteLine(BuildViewerConnectionString( + handoff.Host, handoff.Port, handoff.Role, handoff.Password, clientCertificatePath)); + output.WriteLine(); + } /* Emit the server cert PEM so the operator can copy it to the viewer machine. */ if (certificate is not null) @@ -418,6 +453,198 @@ must be emitted ahead of the payload (#1953 item 3). The field report watched th return 0; } + /* --------------------------------------------------------------------------------------------------- + REPRINTING AN ENDPOINT TOKEN (#2479, item 2). + + --configure-network shows each generated token's plaintext exactly once. Lose it and the only path + was regeneration, which invalidates every client already configured against it - so one mislaid + token during UAT meant re-onboarding every consumer of that endpoint. That is an unrecoverable + mistake made out of a recoverable one, and the recovery costs nothing to provide. + + WHAT THIS DISCLOSES, checked rather than assumed. Nothing new. darling.json stores the token as a + DarlingSecrets blob: DPAPI, DataProtectionScope.LocalMachine, with the entropy constant + "PerformanceMonitor.Darling.v1" compiled into an OPEN-SOURCE binary. LocalMachine scope means any + process on this box can Unprotect it, and the entropy is not a secret because it is published in + this repository. So the whole of this verb is four lines of PowerShell to anyone holding the file. + + WHICH IS WHY THE ELEVATION GATE IS NOT A CONFIDENTIALITY BOUNDARY, and this comment says so rather + than letting the code imply otherwise. install-darling.ps1 and DarlingFileSecurity deliberately grant + INTERACTIVE *read* on darling.json - the Viewer and these very CLI verbs are run by the interactive + operator - so an ordinary logged-on user who is NOT an administrator can already read the blob and + decrypt it. The gate is worth having for what it actually does: it keeps the verb consistent with the + other endpoint verbs that carry "(run elevated)", it stops a script or a shared shell picking a live + credential up in passing, and it makes reprinting a deliberate act. The token's real protection is + the file's ACL, and it always was. + + Emission follows --print-viewer-connection's posture exactly (D8): every warning on STDERR, ahead of + any STDOUT payload, so redirecting STDOUT to an ACL'd file or the clipboard captures the token + WITHOUT swallowing the warning. The field report on that verb watched a live password scroll past + and only then saw the redirect advice, which is precisely backwards. + --------------------------------------------------------------------------------------------------- */ + + /// Reprints the MCP bearer token from mcp.network. See the block above for the + /// disclosure analysis. Returns 0 when a token was printed, 1 otherwise. + /// + /// No Async suffix, deliberately: this reads one file and decrypts one blob, with nothing + /// to await. It follows HardenFiles rather than the PrintViewerConnectionAsync directly + /// above, which really is async - naming a synchronous method …Async invites a caller to await + /// something that never yields (review catch on #2479). + [SupportedOSPlatform("windows")] + public static int PrintMcpToken(string? configPath, TextWriter output, TextWriter error) + { + DarlingConfig config; + try + { + config = DarlingConfig.Load(configPath); + } + catch (Exception ex) + { + error.WriteLine($"Could not load configuration: {ex.Message}"); + return 1; + } + + var network = config.Mcp.Network; + return PrintEndpointToken( + "mcp", + "MCP bearer", + "--print-mcp-token", + "Remote MCP clients send it as the header: Authorization: Bearer ", + IsElevated(), + network is not null && network.IsConfigured, + () => network!.ResolveToken(out _), + output, + error); + } + + /// Reprints the web dashboard access token from web.network. Returns 0 when a token was + /// printed, 1 otherwise. + [SupportedOSPlatform("windows")] + public static int PrintWebToken(string? configPath, TextWriter output, TextWriter error) + { + DarlingConfig config; + try + { + config = DarlingConfig.Load(configPath); + } + catch (Exception ex) + { + error.WriteLine($"Could not load configuration: {ex.Message}"); + return 1; + } + + var network = config.Web.Network; + return PrintEndpointToken( + "web", + "web dashboard access", + "--print-web-token", + "A remote browser presents it once via ?token=... and gets a session cookie back.", + IsElevated(), + network is not null && network.IsConfigured, + () => network!.ResolveToken(out _), + output, + error); + } + + /// + /// The shared body. Split out so the elevation refusal, the disclosure warning and the redirect advice + /// are written once and cannot drift between the two endpoints - the same reason + /// DescribeToggleOverride takes a section name rather than existing twice. + /// + /// is passed IN rather than measured here, the same way the pure alert + /// gates take nowUtc: it makes both branches - the refusal and the disclosure - drivable in a test + /// on any runner, elevated or not. A verb whose refusal path is only exercised when the CI agent happens + /// to be unprivileged is a refusal nobody has checked. + /// + internal static int PrintEndpointToken( + string section, + string tokenName, + string verb, + string howClientsPresentIt, + bool elevated, + bool configured, + Func resolveToken, + TextWriter output, + TextWriter error) + { + if (!elevated) + { + error.WriteLine($"{verb} needs an ELEVATED PowerShell."); + error.WriteLine(); + error.WriteLine("It reprints a live credential, so it is a deliberate act rather than something a script or a"); + error.WriteLine("shared shell picks up in passing."); + error.WriteLine(); + error.WriteLine("Be clear about what this gate is NOT, though: darling.json grants INTERACTIVE read on purpose"); + error.WriteLine("(the Viewer and these CLI verbs are run by the interactive operator), and the token is DPAPI-"); + error.WriteLine("protected at LocalMachine scope with an entropy constant published in this project's source."); + error.WriteLine("Anyone who can log on to this box interactively can already decrypt it without this verb. The"); + error.WriteLine("token's protection is the FILE'S ACL - keep darling.json restricted, and run --harden-files if"); + error.WriteLine("you are not sure it still is."); + return 1; + } + + if (!configured) + { + error.WriteLine($"darling.json has no {section}.network block, so there is no {tokenName} token to print."); + error.WriteLine($"Run --configure-network to create one (it generates the token and shows it once)."); + return 1; + } + + string? token; + try + { + token = resolveToken(); + } + catch (System.Security.Cryptography.CryptographicException ex) + { + /* The DPAPI cause, and ONLY the DPAPI cause. ResolveToken also resolves env: and file: secret + references, and those fail with an InvalidOperationException whose message already names the + missing variable or the unreadable path (review catch on #2479). Appending "a DPAPI blob only + decrypts on the machine that produced it - regenerate it" to THOSE would be a false diagnosis + pointing at a destructive fix: regenerating invalidates every configured client, when the + actual repair is to set the environment variable or correct the path. So the advice is gated + on the exception that earns it, and the arm below reports what it actually knows. */ + error.WriteLine($"Could not read the {tokenName} token: {ex.Message}"); + error.WriteLine( + $"A DPAPI blob only decrypts on the machine that produced it, so a {section}.network.encryptedToken " + + "copied from another host can never be read here - regenerate it with --configure-network."); + return 1; + } + catch (Exception ex) + { + error.WriteLine($"Could not read the {tokenName} token: {ex.Message}"); + error.WriteLine( + $"That is what {section}.network resolved to. Fix the source it names (an env: reference needs " + + "the variable set in THIS shell; a file: reference needs a path this account can read), then " + + "re-run - the token itself is fine and does not need regenerating."); + return 1; + } + + if (string.IsNullOrEmpty(token)) + { + error.WriteLine($"{section}.network is configured but carries no token, so this endpoint is not LAN-exposed."); + error.WriteLine("Run --configure-network to generate one."); + return 1; + } + + /* Every stderr line before the payload (D8), so a STDOUT redirect keeps the token and still shows + the warning. */ + error.WriteLine(); + error.WriteLine( + $"WARNING: the {tokenName} token below is written to STDOUT as PLAINTEXT. It gates ALL network access to " + + $"this endpoint. Redirect it to an ACL'd file or pipe it to the clipboard; do not leave it in shell " + + "scrollback, CI logs, or a screenshare."); + error.WriteLine($" Example (file): PerformanceMonitor.Darling.Service.exe {verb} > token.txt"); + error.WriteLine($" Example (clipboard): PerformanceMonitor.Darling.Service.exe {verb} | clip"); + error.WriteLine($" {howClientsPresentIt}"); + error.WriteLine( + " If this token has actually LEAKED, reprinting it is not the fix: --configure-network generates a new " + + "one, which invalidates every client already configured against the old one."); + error.WriteLine(); + + output.WriteLine(token); + return 0; + } + /// /// The name the exported/pasted client-side certificate takes on the VIEWER machine, and therefore the /// Root Certificate= value both viewer verbs emit. Deliberately the same file name the store @@ -435,16 +662,22 @@ private sealed record ViewerHandoff( string Host, int Port, string Role, string Password, string CertificatePath); /// - /// Resolves the remote-viewer handoff material (D8), writing every refusal + warning to + /// Resolves the remote-viewer handoff material (D8) for EVERY role the exposure admits (#2665), in + /// ' order, writing every refusal + warning to /// and returning null when the caller must exit 1. Managed-mode only: the DPAPI /// credential files and the generated TLS cert it reads exist only there — in BYO the operator's own /// PostgreSQL governs exposure + credentials (D-BYO). Windows-only (DPAPI-LocalMachine). The two string /// parameters shape message WORDING only — never the logic, so both verbs resolve identically. + /// Fail-closed on the FIRST role whose credential is missing or will not decrypt, rather than + /// returning the roles that worked. Both credentials are provisioned in the same act by + /// DarlingManagedRoles, so one missing means the store's bootstrap did not complete — which is + /// what the shared missing-credential diagnostic below actually explains. A partial handoff would bury + /// that under an export that looks like it succeeded. /// /// The CLI verb to name in refusals, e.g. --export-viewer-config. /// What the verb would have produced, e.g. "print" / "export", for the no-remote-connection refusal. [SupportedOSPlatform("windows")] - private static async Task ResolveViewerHandoffAsync( + private static async Task?> ResolveViewerHandoffsAsync( DarlingConfig config, string verb, string action, TextWriter error, CancellationToken cancellationToken) { var postgres = config.Postgres; @@ -463,16 +696,17 @@ private sealed record ViewerHandoff( return null; } - /* The pg_hba login role the network exposure names — default viewer (read-only, the secure default). + /* Every pg_hba login role the network exposure names — default viewer (read-only, the secure default). An explicitly-invalid value is a hard error: the store degrades to loopback for it, so no remote connection exists at all. */ var network = postgres.Network; - var role = DarlingNetwork.NormalizeNetworkRole(network?.Role); - if (role is null) + var roles = DarlingNetwork.NormalizeNetworkRoles(network?.Role); + if (roles is null || roles.Count == 0) { error.WriteLine( - $"postgres.network.role '{network?.Role}' is invalid — it must be \"viewer\" (default, read-only) " + - $"or \"admin\". The store degrades to loopback for an unknown role, so there is no remote connection to {action}."); + $"postgres.network.role '{network?.Role}' is invalid — it must be \"viewer\" (default, read-only), " + + $"\"admin\", or both (e.g. \"admin,viewer\"). The store degrades to loopback for an unknown role, so " + + $"there is no remote connection to {action}."); return null; } @@ -488,44 +722,70 @@ but the endpoint will not accept it until postgres.network.listen is set and the var host = ResolveViewerHost(network?.Listen); - /* Decrypt the role's DPAPI-LocalMachine credential (Windows-only; the caller is IsWindows-guarded). - The cert lives in the same directory as the credential (ParentOf(dataDirectory)). */ + /* Decrypt each role's DPAPI-LocalMachine credential (Windows-only; the caller is IsWindows-guarded). + Every credential and the cert live in the same directory (ParentOf(dataDirectory)), so the cert + path is role-independent. */ var dataDirectory = DarlingManagedPostgres.ResolveDataDirectory(postgres); - var credentialPath = string.Equals(role, "admin", StringComparison.Ordinal) - ? DarlingManagedPostgres.AdminCredentialPathFor(dataDirectory) - : DarlingManagedPostgres.ViewerCredentialPathFor(dataDirectory); - - if (!File.Exists(credentialPath)) - { - /* #2197: which of the two things this means is decided from the store's own files, not assumed. - A bootstrap that has already failed produces this same absence, and telling THAT operator to - start the service again is the dead end the field report walked into. */ - error.WriteLine(DarlingStoreBootstrapEvidence.MissingCredentialMessage( - $"The '{role}' role credential ({credentialPath})", - "provisions the least-privilege roles and their credentials", - dataDirectory)); - return null; - } + var certificatePath = Path.Combine( + Path.GetDirectoryName(DarlingManagedPostgres.ViewerCredentialPathFor(dataDirectory))!, + DarlingManagedPostgres.ServerCertFileName); - string password; - try + var handoffs = new List(roles.Count); + foreach (var role in roles) { - password = DarlingSecrets.Unprotect((await File.ReadAllTextAsync(credentialPath, cancellationToken)).Trim()); + var credentialPath = string.Equals(role, "admin", StringComparison.Ordinal) + ? DarlingManagedPostgres.AdminCredentialPathFor(dataDirectory) + : DarlingManagedPostgres.ViewerCredentialPathFor(dataDirectory); + + if (!File.Exists(credentialPath)) + { + /* #2197: which of the two things this means is decided from the store's own files, not assumed. + A bootstrap that has already failed produces this same absence, and telling THAT operator to + start the service again is the dead end the field report walked into. */ + error.WriteLine(DarlingStoreBootstrapEvidence.MissingCredentialMessage( + $"The '{role}' role credential ({credentialPath})", + "provisions the least-privilege roles and their credentials", + dataDirectory)); + return null; + } + + string password; + try + { + password = DarlingSecrets.Unprotect((await File.ReadAllTextAsync(credentialPath, cancellationToken)).Trim()); + } + catch (Exception ex) + { + error.WriteLine( + $"Could not decrypt the '{role}' credential at {credentialPath}: {ex.Message} (DPAPI-LocalMachine — " + + "run this on the same machine as the service, under an account that can read the credential)."); + return null; + } + + handoffs.Add(new ViewerHandoff(host, postgres.Port, role, password, certificatePath)); } - catch (Exception ex) + + return handoffs; + } + + /// + /// The single seat --export-viewer-config writes when the exposure admits both roles (#2665): the + /// LEAST-PRIVILEGE one. The verb produces one folder holding one live credential, and the folder is what + /// gets handed to somebody else — so where there is a choice it must be the read-only seat, the same call + /// D7 makes for the default. The admin string stays one --print-viewer-connection away, and the + /// verb says so rather than leaving the choice invisible. Pure. + /// + private static ViewerHandoff SelectExportSeat(IReadOnlyList handoffs) + { + foreach (var handoff in handoffs) { - error.WriteLine( - $"Could not decrypt the '{role}' credential at {credentialPath}: {ex.Message} (DPAPI-LocalMachine — " + - "run this on the same machine as the service, under an account that can read the credential)."); - return null; + if (string.Equals(handoff.Role, "viewer", StringComparison.Ordinal)) + { + return handoff; + } } - return new ViewerHandoff( - host, - postgres.Port, - role, - password, - Path.Combine(Path.GetDirectoryName(credentialPath)!, DarlingManagedPostgres.ServerCertFileName)); + return handoffs[0]; } /// @@ -657,13 +917,26 @@ that fails later has not thrown away the previous export. */ return 1; } - var handoff = await ResolveViewerHandoffAsync( + var handoffs = await ResolveViewerHandoffsAsync( config, "--export-viewer-config", "export", error, cancellationToken); - if (handoff is null) + if (handoffs is null) { return 1; } + /* One folder holds one live credential, so a multi-role exposure has to CHOOSE — and the folder is + what gets handed to somebody else, so the choice is the read-only seat (see SelectExportSeat). + Saying which, and where the other one is, because an operator who set "admin,viewer" and got a + folder back would otherwise have no way to know a seat was picked for them. */ + var handoff = SelectExportSeat(handoffs); + if (handoffs.Count > 1) + { + error.WriteLine( + $"NOTE: postgres.network.role admits '{string.Join("', '", handoffs.Select(h => h.Role))}'. This folder is the " + + $"'{handoff.Role}' seat (the least-privileged of them). For another role's connection string, run " + + "--print-viewer-connection, which prints one per admitted role."); + } + string certificate; if (!File.Exists(handoff.CertificatePath)) { @@ -1341,11 +1614,40 @@ webNow.Reason is DarlingHostBinding.BindReason.NetworkExposed or DarlingHostBind output.WriteLine("Current exposure:"); output.WriteLine(DarlingNetworkConfigEditor.FormatExposureState( - "Store", storeNow.Exposed, storeNow.ListenIp, storeNow.Cidr, storeNow.Role, storeNow.DegradeReason)); + "Store", storeNow.Exposed, storeNow.ListenIp, storeNow.Cidr, + storeNow.Roles is null ? null : string.Join(", ", storeNow.Roles), storeNow.DegradeReason)); output.WriteLine(DarlingNetworkConfigEditor.FormatExposureState( "MCP ", mcpNowExposed, config.Mcp.Network?.Listen, config.Mcp.Network?.AllowFrom, null, mcpNowDegrade)); output.WriteLine(DarlingNetworkConfigEditor.FormatExposureState( "Web ", webNowExposed, config.Web.Network?.Listen, config.Web.Network?.AllowFrom, null, webNowDegrade)); + + /* #2562: whether the exposed dashboard is ENCRYPTED belongs in the exposure summary — this verb is + what an operator runs to answer "what is open", and "open on the LAN" reads very differently with + and without TLS. Only when exposed: TLS is meaningless on a loopback-only dashboard, and a line + saying "OFF" there would be a warning about nothing. The wizard does not PROMPT for a certificate + (web.network.tls is file-defined and restart-only, like the rest of the block); it reports it. */ + if (webNowExposed) + { + var tls = DarlingWebTls.Describe(config.Web.Network?.Tls); + output.WriteLine(tls.Shape switch + { + DarlingWebTls.TlsShape.NotConfigured => + " TLS: off — the access token and its session cookie cross the segment in the clear. " + + "Set web.network.tls to serve HTTPS.", + DarlingWebTls.TlsShape.Invalid => + $" TLS: MISCONFIGURED — {tls.Problem} The dashboard will bind loopback-only.", + /* The warning rides along: this verb is the one an operator runs to see what is open, and a + stale PKCS#12 password beside a working PEM pair is exactly the "I thought the bundle was + being served" state they came here to resolve. The service logs it at every start; nobody + reading this summary should have to go find that line. */ + DarlingWebTls.TlsShape.Pem => + $" TLS: on (PEM pair, {config.Web.Network!.Tls!.CertPath})." + + (tls.Warning is null ? string.Empty : $" NOTE: {tls.Warning}"), + _ => $" TLS: on (PKCS#12, {config.Web.Network!.Tls!.PfxPath})." + + (tls.Warning is null ? string.Empty : $" NOTE: {tls.Warning}"), + }); + } + output.WriteLine($" Service: {await DescribeServiceStateAsync(cancellationToken)}"); output.WriteLine(); @@ -1422,12 +1724,15 @@ untouched and a multi-surface run is all-or-nothing. */ { output.WriteLine(); output.WriteLine("== MCP exposure =="); - if (!config.Mcp.Enabled) - { - output.WriteLine("NOTE: mcp.enabled is currently false. The wizard writes the network block, but the endpoint"); - output.WriteLine(" stays down until you enable MCP in the Viewer's Settings (enabled/port are control-plane"); - output.WriteLine(" after first run; the wizard never edits them)."); - } + /* #2389: unconditional, and about the PRECEDENCE rather than the file value. This used to print + only when mcp.enabled was FALSE in the file, which is exactly backwards — a file that says true + while config_service.mcp_enabled says false is the combination that misleads, and it printed + nothing there. The wizard holds no store connection, so it cannot report the effective value; + what it can do honestly is say which plane decides and where the file value stops mattering. */ + output.WriteLine($"NOTE: darling.json has mcp.enabled = {(config.Mcp.Enabled ? "true" : "false")}, but after the first run that is only"); + output.WriteLine(" the SEED — config.config_service.mcp_enabled decides whether the endpoint actually runs, and"); + output.WriteLine(" this wizard neither reads nor writes it. It writes the network block only (which IS file-only"); + output.WriteLine(" and restart-only); use --enable-mcp / --disable-mcp or the Viewer's Settings to turn MCP on."); mcp = GatherMcpInputs(input, output, error, config.Mcp); if (mcp is null) @@ -1441,12 +1746,11 @@ untouched and a multi-surface run is all-or-nothing. */ { output.WriteLine(); output.WriteLine("== Web dashboard exposure =="); - if (!config.Web.Enabled) - { - output.WriteLine("NOTE: web.enabled is currently false. The wizard writes the network block, but the dashboard"); - output.WriteLine(" stays down until you enable it with --enable-web or the Viewer's Settings (enabled/port"); - output.WriteLine(" are control-plane after first run; the wizard never edits them)."); - } + /* #2389: the MCP note's twin — unconditional, and about which plane decides. */ + output.WriteLine($"NOTE: darling.json has web.enabled = {(config.Web.Enabled ? "true" : "false")}, but after the first run that is only"); + output.WriteLine(" the SEED — config.config_service.web_enabled decides whether the dashboard actually runs, and"); + output.WriteLine(" this wizard neither reads nor writes it. It writes the network block only (which IS file-only"); + output.WriteLine(" and restart-only); use --enable-web / --disable-web or the Viewer's Settings to turn it on."); web = GatherWebInputs(input, output, error, config.Web); if (web is null) @@ -1717,9 +2021,12 @@ private static async Task DescribeServiceStateAsync(CancellationToken ca } /// - /// Gathers listen / allowFrom / role for the store, RE-PROMPTING with the store resolver's own degrade + /// Gathers listen / allowFrom / role(s) for the store, RE-PROMPTING with the store resolver's own degrade /// reason until it accepts them (or the operator cancels). The whitespace-in-path degrade is a config /// problem the loop cannot fix, so it is reported and the surface is abandoned. Returns null on cancel. + /// The returned Role is the resolver's CANONICAL joined list (#2665): "viewer + admin" comes + /// back as "admin,viewer", so darling.json holds the text the service would itself compute and a + /// re-run of the wizard is a no-op rather than a reorder. /// [SupportedOSPlatform("windows")] private static (string Listen, string AllowFrom, string Role)? GatherStoreInputs( @@ -1741,15 +2048,21 @@ private static (string Listen, string AllowFrom, string Role)? GatherStoreInputs return null; } - output.WriteLine("Remote pg_hba role: 'viewer' (read-only, the secure default) or 'admin' (remote WRITES)."); - var role = Prompt(input, output, "Role", "viewer"); + output.WriteLine("Remote pg_hba role(s): 'viewer' (read-only, the secure default), 'admin' (remote WRITES),"); + output.WriteLine("or BOTH as 'admin,viewer' — that writes one hostssl rule per role inside the managed block, so"); + output.WriteLine("an admin seat and read-only seats can reach the same store and a later CIDR narrowing tightens"); + output.WriteLine("all of them together (#2665)."); + var role = Prompt(input, output, "Role(s)", "viewer"); if (role is null) { output.WriteLine("Cancelled — no changes made."); return null; } - if (string.Equals(role, "admin", StringComparison.OrdinalIgnoreCase)) + /* Warn off the NORMALIZED value, not the typed text. A check for the string being exactly "admin" + goes silent for every list form ("admin,viewer", "viewer+admin") — which is the case where the + operator most needs to read it, because they may be thinking about the viewer half. */ + if (string.Equals(DarlingNetwork.NormalizeNetworkRole(role), "admin", StringComparison.Ordinal)) { output.WriteLine(" WARNING: 'admin' is a remote WRITE credential holding the config-table service-credential pivot."); output.WriteLine(" Prefer 'viewer' unless you specifically need remote writes."); @@ -1761,7 +2074,7 @@ private static (string Listen, string AllowFrom, string Role)? GatherStoreInputs { /* Write the resolver's canonical values (parsed IP, host-bits-zeroed CIDR, normalized role) so the file matches what the service would compute. */ - return (decision.ListenIp!, decision.Cidr!, decision.Role!); + return (decision.ListenIp!, decision.Cidr!, string.Join(",", decision.Roles!)); } var reason = decision.DegradeReason ?? "the store resolver rejected these values"; @@ -2086,11 +2399,18 @@ private static void PrintNextSteps( output.WriteLine(" MCP firewall rule (run ELEVATED; scoped to the port + CIDR):"); output.WriteLine(" " + DarlingManagedPostgres.BuildFirewallEnableCommand( DarlingMcpHostService.McpFirewallRuleName(mcpPort), mcpPort, mcpCidr!)); - if (!mcpEnabled) - { - output.WriteLine(" NOTE: mcp.enabled is false, so the MCP endpoint stays down until you enable MCP in the"); - output.WriteLine(" Viewer's Settings. The network block you just wrote applies once MCP is enabled."); - } + /* #2414: this wizard holds no store connection, so the only port it can name is darling.json's, and + the port is part of the rule's NAME. Say that, and point at the verb that resolves the effective + one — pasting the line below on a box whose port has been moved opens the wrong port. */ + output.WriteLine($" (that rule is named for darling.json's mcp.port = {mcpPort}. If the port was ever changed in the"); + output.WriteLine(" Viewer's Settings, config.config_service.mcp_port is what the endpoint binds — run"); + output.WriteLine(" --configure-firewall ELEVATED instead and it resolves the effective port and moves the rule.)"); + /* #2389: unconditional and about the PLANE, not the file value. This printed only when the FILE + said false, which left the actually-misleading combination — file true, store false — with no + note at all; and mcpEnabled here is the file's value, which on a seeded box decides nothing. */ + output.WriteLine(" NOTE: the network block you just wrote is FILE-authoritative and applies on restart, but whether"); + output.WriteLine(" MCP runs at all is config.config_service.mcp_enabled — darling.json's mcp.enabled is only the"); + output.WriteLine($" first-run seed, and it currently reads {(mcpEnabled ? "true" : "false")}. Enable with --enable-mcp or the Viewer's Settings."); } if (webConfigured) @@ -2099,6 +2419,10 @@ private static void PrintNextSteps( output.WriteLine(" reconciles this rule for you):"); output.WriteLine(" " + DarlingManagedPostgres.BuildFirewallEnableCommand( DarlingWebHostService.WebFirewallRuleName(webPort), webPort, webCidr!)); + /* #2414: the MCP caveat's twin. */ + output.WriteLine($" (that rule is named for darling.json's web.port = {webPort}. If the port was ever changed in the"); + output.WriteLine(" Viewer's Settings, config.config_service.web_port is what the endpoint binds — run"); + output.WriteLine(" --configure-firewall ELEVATED instead and it resolves the effective port and moves the rule.)"); /* The one login step a human does differently for Web: a remote browser presents the access token once via ?token= and is 302'd back with a session cookie. A 0.0.0.0 bind has no single address @@ -2107,12 +2431,10 @@ private static void PrintNextSteps( output.WriteLine(" Remote browser login (after the service restarts):"); output.WriteLine($" http://{webHost}:{webPort}/?token="); output.WriteLine(" (the token is exchanged for a session cookie and stripped from the URL; loopback needs no token)"); - if (!webEnabled) - { - output.WriteLine(" NOTE: web.enabled is false, so the dashboard stays down until you enable it with --enable-web"); - output.WriteLine(" or the Viewer's Settings. The network block you just wrote applies once the web dashboard"); - output.WriteLine(" is enabled."); - } + /* #2389: the MCP note's twin — unconditional, and about which plane decides. */ + output.WriteLine(" NOTE: the network block you just wrote is FILE-authoritative and applies on restart, but whether"); + output.WriteLine(" the dashboard runs at all is config.config_service.web_enabled — darling.json's web.enabled is"); + output.WriteLine($" only the first-run seed, and it currently reads {(webEnabled ? "true" : "false")}. Enable with --enable-web or Settings."); } } @@ -2293,21 +2615,31 @@ Program dispatch is OperatingSystem.IsWindows()-guarded, mirroring --print-viewe /// The CLI's targeted store write that ENABLES the MCP endpoint on the single config_service row /// (id=1). Sets only mcp_enabled + the audit columns; the BEFORE-UPDATE self-bump trigger fires - /// config_version (deliberately NOT set here) so the worker hot-reloads. Pure — Darling.Tests pin the shape. + /// config_version (deliberately NOT set here) so the worker hot-reloads. Pure — Darling.Tests pin the shape. + /// #2414: it also RETURNS mcp_port. The firewall half of this verb has to name its rule for the + /// port the endpoint BINDS, which is this column and not darling.json's seed, and returning it from the write + /// itself makes that read free, atomic with the toggle, and impossible to skip — a separate SELECT is a second + /// resolution path, which is exactly how the two ports drifted apart in the first place. public const string EnableMcpStoreSql = - "UPDATE config.config_service SET mcp_enabled = TRUE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1"; + "UPDATE config.config_service SET mcp_enabled = TRUE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1 RETURNING mcp_port"; /// The CLI's targeted store write that DISABLES the MCP endpoint (twin of ). public const string DisableMcpStoreSql = - "UPDATE config.config_service SET mcp_enabled = FALSE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1"; + "UPDATE config.config_service SET mcp_enabled = FALSE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1 RETURNING mcp_port"; /// The CLI's targeted store write that ENABLES the read-only web dashboard endpoint (twin of ). public const string EnableWebStoreSql = - "UPDATE config.config_service SET web_enabled = TRUE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1"; + "UPDATE config.config_service SET web_enabled = TRUE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1 RETURNING web_port"; /// The CLI's targeted store write that DISABLES the web dashboard endpoint (twin of ). public const string DisableWebStoreSql = - "UPDATE config.config_service SET web_enabled = FALSE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1"; + "UPDATE config.config_service SET web_enabled = FALSE, updated_at = (now() AT TIME ZONE 'UTC'), updated_by = 'cli' WHERE id = 1 RETURNING web_port"; + + /// The elevated firewall verb's best-effort read of the CONTROL PLANE's effective endpoint toggles + /// (#2414). Read-only, and the only store contact --configure-firewall makes — see + /// for why it is best-effort rather than required. + public const string ReadEndpointTogglesSql = + "SELECT mcp_enabled, mcp_port, web_enabled, web_port FROM config.config_service WHERE id = 1"; /// Which optional endpoint a toggle verb targets — selects the store column, firewall rule name, /// darling.json network block, and seed-key note. @@ -2454,9 +2786,7 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — if (!postgres.Managed) { - error.WriteLine( - $"{verb} applies to the managed store only. In bring-your-own mode (postgres.connectionString), the " + - $"endpoint enable flags live in YOUR PostgreSQL's config.config_service ({column}) — toggle them there."); + error.WriteLine(ByoEndpointToggleMessage(verb, column, enable)); return 1; } @@ -2470,15 +2800,22 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — } /* The TARGETED store write. A DIRECT config_service UPDATE self-bumps config_version via the BEFORE-UPDATE - trigger, so the worker hot-reloads within one sweep — we never touch config_version ourselves. */ + trigger, so the worker hot-reloads within one sweep — we never touch config_version ourselves. + + #2414: the statement RETURNS the endpoint's port, so the firewall step below scopes its rule to the port + the endpoint actually BINDS instead of to darling.json's first-run seed. It is a scalar read of the row + this verb just wrote, in the same round trip, under the same lock — and it doubles as the seeded check, + since a RETURNING that yields no row is the 0-rows-updated case. Note what this makes structurally + impossible: the store is unreachable, and the verb has already returned 1 before any firewall command is + built, so these two verbs can never open a port on a guessed number. */ var sql = EndpointToggleSql(endpoint, enable); - int rows; + object? returnedPort; try { await using var connection = new NpgsqlConnection(connectionString); await connection.OpenAsync(cancellationToken); await using var command = new NpgsqlCommand(sql, connection); - rows = await command.ExecuteNonQueryAsync(cancellationToken); + returnedPort = await command.ExecuteScalarAsync(cancellationToken); } catch (OperationCanceledException) { @@ -2490,7 +2827,7 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — return 1; } - if (rows == 0) + if (returnedPort is null or DBNull) { error.WriteLine( "The control-plane store is not seeded yet (config.config_service has no id=1 row) — start the " + @@ -2498,6 +2835,8 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — return 1; } + var storePort = Convert.ToInt32(returnedPort, CultureInfo.InvariantCulture); + output.WriteLine( $"{endpointLabel} endpoint {(enable ? "ENABLED" : "DISABLED")} in the control-plane store " + $"(config.config_service.{column} = {(enable ? "true" : "false")})."); @@ -2505,7 +2844,15 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — "The running service applies this LIVE within one collection sweep (the write self-bumps the reload " + "beacon) — no restart needed."); - await ReconcileEndpointFirewallAsync(endpoint, enable, config, output, error, cancellationToken); + /* #2414: the SAME resolver the two host supervisors bind on, fed the row this verb just wrote. It carries + the provenance as well as the value, so the firewall step can state on whose authority it picked a port + and flag a file/store disagreement instead of silently acting on one. */ + var toggle = DarlingHostBinding.ResolveEndpointToggle( + (enable, storePort), + endpoint == EndpointKind.Mcp ? config.Mcp.Enabled : config.Web.Enabled, + endpoint == EndpointKind.Mcp ? config.Mcp.Port : config.Web.Port); + + await ReconcileEndpointFirewallAsync(endpoint, enable, config, toggle, output, error, cancellationToken); output.WriteLine(); output.WriteLine( @@ -2525,6 +2872,60 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — _ => throw new ArgumentOutOfRangeException(nameof(endpoint)), }; + /// + /// What a bring-your-own deployment is told instead of a toggle: the flags live in the operator's own + /// PostgreSQL, and the UPDATE is the whole procedure. Composed here rather than inline because #2626 + /// needs the SAME sentence from a second caller — the platform guard, which is the one a non-Windows + /// operator actually reaches. + /// + private static string ByoEndpointToggleMessage(string verb, string column, bool enable) => + $"{verb} applies to the managed store only. In bring-your-own mode (postgres.connectionString), the " + + $"endpoint enable flags live in YOUR PostgreSQL's config.config_service ({column}) — toggle them " + + $"there:\n UPDATE config.config_service SET {column} = {(enable ? "true" : "false")};\n" + + "The service picks that up within one sweep; no restart is needed."; + + /// + /// The message for an endpoint-toggle verb invoked on a non-Windows host (#2626). + /// + /// + /// "requires Windows" alone is true of the verb and MISLEADING about the situation. The two Windows + /// dependencies — the DPAPI-protected owner credential and the firewall reconcile — are both MANAGED-mode + /// concerns, and a non-Windows deployment is necessarily bring-your-own, where the verb would refuse + /// anyway with the message that actually helps. Reading the platform message on macOS or Linux, an + /// operator concludes the DASHBOARD is Windows-only; it is not, and one UPDATE turns it on. + /// + /// + /// + /// So the config is loaded here, best effort, purely to decide which sentence to print. A config that + /// cannot be loaded, or a managed one (which on a non-Windows host should not exist), gets the platform + /// sentence — the honest answer when we cannot tell. + /// + /// + public static int WriteEndpointVerbPlatformRefusal( + bool isMcp, bool enable, string? configPath, TextWriter error) + { + var endpoint = isMcp ? EndpointKind.Mcp : EndpointKind.Web; + var verb = VerbName(endpoint, enable); + var column = isMcp ? "mcp_enabled" : "web_enabled"; + + var managed = true; + try + { + managed = DarlingConfig.Load(configPath).Postgres?.Managed ?? true; + } + catch + { + /* Deliberately swallowed: this method exists to pick a sentence, and a config we cannot read is + not a reason to fail differently than we already are. */ + } + + error.WriteLine(managed + ? $"{verb} requires Windows (DPAPI + firewall)." + : ByoEndpointToggleMessage(verb, column, enable)); + + return 1; + } + /// The verb spelling for a toggle (for error + handoff text). private static string VerbName(EndpointKind endpoint, bool enable) => (endpoint, enable) switch { @@ -2544,14 +2945,25 @@ managed concerns. In BYO the operator's own PostgreSQL holds config_service — /// and the SAME pure command builders. Elevated -> runs the rule; otherwise prints the exact elevated command /// — the store toggle already succeeded, so a non-elevated shell is a handoff, never a failure. A firewall /// failure is likewise non-fatal. + /// + /// #2414: the port comes from , i.e. from the control plane, NOT from + /// config.Mcp.Port/config.Web.Port. Those are the first-run seed; the supervisor binds the store's + /// value, so naming the rule from the file opened one port while the endpoint served another. The caller + /// resolves the toggle from the row it just wrote, so this method cannot be reached with a guessed port. /// [SupportedOSPlatform("windows")] private static async Task ReconcileEndpointFirewallAsync( - EndpointKind endpoint, bool enable, DarlingConfig config, TextWriter output, TextWriter error, CancellationToken cancellationToken) + EndpointKind endpoint, bool enable, DarlingConfig config, DarlingHostBinding.EndpointToggle toggle, + TextWriter output, TextWriter error, CancellationToken cancellationToken) { - var (port, listen, allowFrom, ruleName) = endpoint == EndpointKind.Mcp - ? (config.Mcp.Port, config.Mcp.Network?.Listen, config.Mcp.Network?.AllowFrom, DarlingMcpHostService.McpFirewallRuleName(config.Mcp.Port)) - : (config.Web.Port, config.Web.Network?.Listen, config.Web.Network?.AllowFrom, DarlingWebHostService.WebFirewallRuleName(config.Web.Port)); + var (section, surface, filePort, listen, allowFrom) = endpoint == EndpointKind.Mcp + ? ("mcp", "MCP", config.Mcp.Port, config.Mcp.Network?.Listen, config.Mcp.Network?.AllowFrom) + : ("web", "web dashboard", config.Web.Port, config.Web.Network?.Listen, config.Web.Network?.AllowFrom); + + var port = toggle.Port; + var ruleName = endpoint == EndpointKind.Mcp + ? DarlingMcpHostService.McpFirewallRuleName(port) + : DarlingWebHostService.WebFirewallRuleName(port); var exposed = DarlingNetwork.IsExposedListenAddress(listen); var plan = ClassifyFirewallPlan(exposed, IsElevated()); @@ -2594,53 +3006,66 @@ there is nothing to open. Point at the wizard rather than emit a malformed New-N } } - var command = enable - ? DarlingManagedPostgres.BuildFirewallEnableCommand(ruleName, port, canonicalCidr) - : DarlingManagedPostgres.BuildFirewallDisableCommand(ruleName); - - if (plan == EndpointFirewallPlan.RunElevated) + /* #2414: state which plane named the port BEFORE doing anything with it. Here the control plane always + answered — the caller could not have got this far otherwise — so this is either a confirmation or the + report of a file/store disagreement the operator has never been shown. Only on ENABLE: a disable sweeps + every port of the surface, so which one the store named decides nothing and claiming otherwise would be + a line that describes work this verb is not doing. */ + if (enable) { - await RunFirewallCommandAsync(command, ruleName, enable, output, error, cancellationToken); - return; + output.WriteLine(DarlingHostBinding.DescribeFirewallPortAuthority(toggle, section, surface, filePort, null)); } - /* Handoff (not elevated) — the store toggle already succeeded; print the exact command to run elevated. */ - output.WriteLine(enable - ? "Firewall: this shell is not elevated, so the endpoint was enabled but its firewall rule was NOT opened. " + - "Run this in an ELEVATED PowerShell to open the port (scoped to the port + CIDR):" - : "Firewall: this shell is not elevated, so the firewall rule was NOT removed. Run this in an ELEVATED PowerShell to close the port:"); - output.WriteLine(" " + command); - } + /* #2414, the security half: sweep EVERY port of this surface before ensuring the desired rule, rather than + reconciling the one exact DisplayName. The port lives in the rule's name, so moving a port does not + update a rule — it creates a second one and strands the first as an inbound allow rule on a port nothing + serves. Exact-name reconciliation can by construction never reach that rule; it is not even aware of it. + The wildcard comes from DarlingFirewallCheck.SurfaceRuleWildcard, which returns the name UNCHANGED when + it does not parse, so the widening is provably confined to other ports of THIS surface and can never + reach a rule this product did not create. --configure-firewall has swept this way since #1771; the + toggle verbs, which are the ones an operator reaches for after changing a port, did not. */ + var wildcard = DarlingFirewallCheck.SurfaceRuleWildcard(ruleName); + var sweepCommand = DarlingManagedPostgres.BuildFirewallSweepCommand(wildcard); + var openCommand = enable + ? DarlingManagedPostgres.BuildFirewallEnableCommand(ruleName, port, canonicalCidr) + : null; - /// Runs a scoped firewall command via the shared PowerShell runner and reports the outcome. NEVER - /// throws (except on cancellation) — the store toggle already succeeded, so a firewall failure degrades to a - /// printed elevated hand-off, never a non-zero exit. - [SupportedOSPlatform("windows")] - private static async Task RunFirewallCommandAsync( - string command, string ruleName, bool enable, TextWriter output, TextWriter error, CancellationToken cancellationToken) - { - try + if (plan == EndpointFirewallPlan.RunElevated) { - var (exitCode, psOutput) = await DarlingManagedPostgres.RunPowerShellAsync(command, cancellationToken); - if (exitCode == 0) + /* Its own step, never concatenated ahead of the open: the sweep command ends in `exit 0` and would + otherwise terminate the shell before the rule was created. A failure in either step is reported with + the exact command and is non-fatal — the store toggle already succeeded. */ + await TryRunFirewallStepAsync( + sweepCommand, + enable + ? $"Firewall: cleared any previous {surface} rule matching '{wildcard}'." + : $"Firewall: removed every {surface} rule matching '{wildcard}'.", + $"Firewall: could not remove the rule(s) matching '{wildcard}'", + output, error, cancellationToken); + + if (openCommand is not null) { - output.WriteLine($"Firewall rule '{ruleName}' {(enable ? "opened" : "removed")}."); - return; + await TryRunFirewallStepAsync( + openCommand, + $"Firewall rule '{ruleName}' opened (TCP {port}, inbound, from {canonicalCidr}).", + $"Firewall rule open did not confirm for '{ruleName}'", + output, error, cancellationToken); } - error.WriteLine( - $"Firewall rule {(enable ? "open" : "removal")} did not confirm (exit {exitCode}: {psOutput}). " + - "Run this in an elevated PowerShell:"); - error.WriteLine(" " + command); + return; } - catch (OperationCanceledException) + + /* Handoff (not elevated) — the store toggle already succeeded; print the exact commands to run elevated. */ + output.WriteLine(enable + ? "Firewall: this shell is not elevated, so the endpoint was enabled but its firewall rule was NOT opened. " + + "Run these in an ELEVATED PowerShell — the first clears any rule left on a previously configured port, " + + "the second opens this one (scoped to the port + CIDR):" + : "Firewall: this shell is not elevated, so the firewall rule was NOT removed. Run this in an ELEVATED " + + "PowerShell to close the port (it covers every port this surface has ever been configured on):"); + output.WriteLine(" " + sweepCommand); + if (openCommand is not null) { - throw; - } - catch (Exception ex) - { - error.WriteLine($"Firewall rule {(enable ? "open" : "removal")} failed ({ex.Message}). Run this in an elevated PowerShell:"); - error.WriteLine(" " + command); + output.WriteLine(" " + openCommand); } } @@ -2654,9 +3079,20 @@ deployment mode — that meant the rule was never created at all and remote clie calls this verb, uninstall-darling.ps1 removes the rules, and the running service only VERIFIES (DarlingFirewallCheck). - It reconciles all THREE surfaces (store, MCP, web) in one pass, from darling.json alone: no store - connection, no credentials, so it is safe to run at install time before the store has ever booted — - unlike --enable-mcp/--enable-web, which write the control-plane store and therefore need it running. + It reconciles all THREE surfaces (store, MCP, web) in one pass. #2414 gave it ONE optional store read: + the MCP and web PORTS are control-plane-authoritative (config_service.mcp_port/web_port), and naming a + rule from darling.json's seed opened one port while the endpoint served another. The read is best-effort + and never a precondition — this verb still has to work at install time, before the store exists, where + the file's port is the value the store will be seeded WITH. When the read fails, the verb says which port + it used and why, and it never falls back silently: see TryReadEndpointTogglesAsync. + + #2436 followed the same read one step further, into what the verb is entitled to DO with it. The store + row carries an enable flag as well as a port, so a surface the control plane has switched off gets its + rule removed rather than opened — the posture --disable-mcp already had, which this verb used to undo on + every upgrade (see DescribeDisabledSurface). And when the read FAILS, the verb stops removing a + LAN-exposed surface's existing rules at all — neither the port nor the enable flag is knowable then, and + both are ways to end up deleting the rule the endpoint is actually being served on. It still creates the + file's rule, which is what a fresh install needs (see FirewallRulePlan.SweepOtherPorts). ================================================================================================ */ /// What --configure-firewall will do to one surface's rule. @@ -2671,9 +3107,31 @@ public enum FirewallRuleAction } /// One surface's desired firewall state. is a human explanation printed - /// alongside, non-null only where the reason is not self-evident (a fail-closed exposure). + /// alongside, non-null only where the reason is not self-evident (a fail-closed exposure, or a surface the + /// control plane has switched off). + /// (#2414) says which PLANE supplied and, when the store + /// could not be read, what would make the chosen port wrong — non-null only on an Open plan, since a Remove + /// sweeps every port of the surface and does not depend on picking the right one. + /// (#2436) is permission to remove this surface's rules on ports + /// OTHER than . That sweep is how a rule stranded by a port change gets collected, + /// and its whole justification is that the sweeper knows what this surface is actually doing — so it is + /// withheld on exactly one combination: a surface darling.json exposes on the LAN, on a run that could not + /// read the control plane. There, BOTH store-backed values are a guess. The port may have moved, so another + /// port's rule may be the live one; and mcp_enabled may have been turned on with --enable-mcp + /// or in the Viewer, which never writes back to darling.json — so the file's enabled = false may be + /// years stale and the surface may be serving right now. Either way the wildcard would delete the rule the + /// endpoint is being served on, which is the outage this whole change exists to prevent. Nothing is exposed + /// by deferring: the store is unreadable because the service is stopped, so no Darling endpoint is + /// listening on any port until a run that CAN read the control plane is possible again. + /// The other half is certain and always sweeps. Whether a surface is network-exposed at all is + /// decided by mcp.network / web.network, which are FILE-ONLY and have no config_service + /// equivalent by design (#2389) — the control plane can switch an exposed surface off, never switch an + /// unexposed one on. So a bind that resolves loopback-only is loopback-only whatever the store would have + /// said, no rule belongs on any port, and collecting them needs no knowledge this run is missing. Same for + /// the store surface, whose postgres.port has no config_service column to disagree with it. public readonly record struct FirewallRulePlan( - string Surface, string RuleName, int Port, FirewallRuleAction Action, string? Cidr, string? Note); + string Surface, string RuleName, int Port, FirewallRuleAction Action, string? Cidr, string? Note, string? PortNote, + bool SweepOtherPorts = true); /// /// PURE desired-state for all three rules, so the whole decision pins without a live firewall. @@ -2686,9 +3144,19 @@ public readonly record struct FirewallRulePlan( /// otherwise every start would report a rule this verb had just deliberately created. /// BYO mode resolves loopback for MCP/web (their resolvers take managed and refuse exposure /// without it) and skips the store entirely, whose exposure the operator's own PostgreSQL governs. + /// #2414: the MCP and web PORTS come from / when + /// the caller could read config.config_service, because that is the port the supervisor binds; they fall + /// back to darling.json's seed when it could not, which is the normal state at install time and carries a + /// saying so. The store surface has no such split — postgres.port + /// is file-only, with no config_service column to disagree with it — so it is resolved from the file, full + /// stop, and gets no note. /// [SupportedOSPlatform("windows")] - public static IReadOnlyList PlanFirewallRules(DarlingConfig config) + public static IReadOnlyList PlanFirewallRules( + DarlingConfig config, + (bool Enabled, int Port)? mcpStore = null, + (bool Enabled, int Port)? webStore = null, + string? storeUnavailableReason = null) { var plans = new List(); var managed = config.Postgres?.Managed ?? false; @@ -2703,28 +3171,59 @@ public static IReadOnlyList PlanFirewallRules(DarlingConfig co config.Postgres.Port, store.Exposed ? FirewallRuleAction.Open : FirewallRuleAction.Remove, store.Exposed ? store.Cidr : null, - store.DegradeReason)); + store.DegradeReason, + null, + /* postgres.port is file-only — there is no config_service column able to disagree with it — so + this surface's port is never a guess and the sweep is always authoritative. */ + SweepOtherPorts: true)); } + /* #2414: the SAME resolver the MCP/web supervisors bind on, so the rule this verb writes is named for the + port the endpoint serves. A null store answer resolves to the file value with Origin = File, which the + describer turns into the loud fallback disclosure rather than a silent guess. */ + var mcpToggle = DarlingHostBinding.ResolveEndpointToggle(mcpStore, config.Mcp.Enabled, config.Mcp.Port); var mcpBind = DarlingMcpHostService.ResolveMcpBind(config.Mcp, managed); var mcpExposed = mcpBind.Mode == DarlingMcpHostService.McpBindMode.NetworkAndLoopback; + var mcpOpen = mcpExposed && mcpToggle.Enabled; plans.Add(new FirewallRulePlan( "MCP", - DarlingMcpHostService.McpFirewallRuleName(config.Mcp.Port), - config.Mcp.Port, - mcpExposed ? FirewallRuleAction.Open : FirewallRuleAction.Remove, - mcpExposed ? CanonicalCidrOrNull(config.Mcp.Network?.AllowFrom) : null, - mcpExposed ? null : DescribeLoopbackReason(mcpBind.Reason, "mcp"))); - + DarlingMcpHostService.McpFirewallRuleName(mcpToggle.Port), + mcpToggle.Port, + mcpOpen ? FirewallRuleAction.Open : FirewallRuleAction.Remove, + mcpOpen ? CanonicalCidrOrNull(config.Mcp.Network?.AllowFrom) : null, + mcpOpen + ? null + : mcpExposed + ? DescribeDisabledSurface(mcpToggle, "mcp", storeUnavailableReason) + : DescribeLoopbackReason(mcpBind.Reason, "mcp"), + mcpOpen + ? DarlingHostBinding.DescribeFirewallPortAuthority(mcpToggle, "mcp", "MCP", config.Mcp.Port, storeUnavailableReason) + : null, + /* NOT !mcpOpen: a Remove reached because the FILE said disabled is exactly as much a guess as a + stale port, because --enable-mcp and the Viewer write only the store. mcpExposed is the half + that is certain — network.* is file-only, so the control plane can never expose what the file + does not. */ + SweepOtherPorts: !mcpExposed || mcpToggle.Origin == DarlingHostBinding.EndpointToggleOrigin.ControlPlane)); + + var webToggle = DarlingHostBinding.ResolveEndpointToggle(webStore, config.Web.Enabled, config.Web.Port); var webBind = DarlingWebHostService.ResolveWebBind(config.Web, managed); var webExposed = webBind.Mode == DarlingHostBinding.BindMode.NetworkAndLoopback; + var webOpen = webExposed && webToggle.Enabled; plans.Add(new FirewallRulePlan( "web dashboard", - DarlingWebHostService.WebFirewallRuleName(config.Web.Port), - config.Web.Port, - webExposed ? FirewallRuleAction.Open : FirewallRuleAction.Remove, - webExposed ? CanonicalCidrOrNull(config.Web.Network?.AllowFrom) : null, - webExposed ? null : DescribeLoopbackReason((DarlingMcpHostService.McpBindReason)webBind.Reason, "web"))); + DarlingWebHostService.WebFirewallRuleName(webToggle.Port), + webToggle.Port, + webOpen ? FirewallRuleAction.Open : FirewallRuleAction.Remove, + webOpen ? CanonicalCidrOrNull(config.Web.Network?.AllowFrom) : null, + webOpen + ? null + : webExposed + ? DescribeDisabledSurface(webToggle, "web", storeUnavailableReason) + : DescribeLoopbackReason((DarlingMcpHostService.McpBindReason)webBind.Reason, "web"), + webOpen + ? DarlingHostBinding.DescribeFirewallPortAuthority(webToggle, "web", "web dashboard", config.Web.Port, storeUnavailableReason) + : null, + SweepOtherPorts: !webExposed || webToggle.Origin == DarlingHostBinding.EndpointToggleOrigin.ControlPlane)); return plans; } @@ -2735,6 +3234,47 @@ public static IReadOnlyList PlanFirewallRules(DarlingConfig co private static string? CanonicalCidrOrNull(string? allowFrom) => ClassifyAllowFrom(allowFrom, out var canonical) == EndpointAllowFromVerdict.Valid ? canonical : null; + /// + /// Why a surface that IS configured for LAN exposure still gets no rule: it is switched OFF (#2436), so + /// the supervisor never starts it and there is nothing behind the port. + /// + /// The decision, because the alternative is arguable. Leaving the rule open would mean an + /// operator who enables the endpoint from the Viewer's Settings — a store write the running service picks + /// up within one poll interval, with no elevated step anywhere in it — finds the LAN path already open. + /// That convenience is real, and it is what this verb used to do. It loses to three things. The product's + /// own posture everywhere else is that no listener means no rule: --disable-mcp sweeps the surface, + /// 's stop path documents an admin removing the rule with THIS verb, + /// and #2414 was fixed precisely because an inbound allow on a port nothing serves is the inverse of what + /// scoping a rule to a port is for. Worse, the two disagreed: --disable-mcp closed the port and the + /// next --configure-firewall — which every upgrade runs — silently re-opened it, so the disable verb + /// did not stay done. And the convenience is not lost, only deferred to one elevated action the service + /// already asks for by name: its start-up check reports the missing rule and prints the command. + /// + /// Two messages, on the same three-state honesty + /// applies to the port — and for a reason review had to point out. Origin == File is NOT "fresh + /// install": collapses BYO, a missing credential, a timeout and + /// every other connection failure into the same answer, so this branch is also reached on a long-lived box + /// whose store is merely unreachable this minute and may well hold mcp_enabled = true. Asserting the + /// endpoint is off there would contradict the sweep-declined line printed a few lines later in the same + /// run, which says the opposite — that the off may be stale and the rule may be the live one — and it is + /// the sweep-declined line that is right. + /// + private static string DescribeDisabledSurface( + DarlingHostBinding.EndpointToggle toggle, string section, string? storeUnavailableReason) + => toggle.Origin == DarlingHostBinding.EndpointToggleOrigin.ControlPlane + ? $"{section}.network exposes this endpoint, but the CONTROL PLANE has it off " + + $"(config.config_service.{section}_enabled = false), so the service does not start it and no rule " + + $"belongs on that port — turn it on with --enable-{section} or in the Viewer's Settings, then " + + "re-run --configure-firewall from an elevated prompt" + : $"{section}.network exposes this endpoint, but the control plane could NOT be read " + + $"({storeUnavailableReason ?? "reason unknown"}), so this run goes on darling.json's " + + $"{section}.enabled = false and opens nothing. On a box whose store has never been written — the " + + $"normal state at install time — that is right: config.config_service.{section}_enabled is SEEDED " + + "from this value, so the endpoint will not start. On a box that has run before it may be stale, " + + $"because --enable-{section} and the Viewer's Settings write only the control plane and never back " + + "to the file; if the endpoint IS enabled there, re-run --configure-firewall once the store is up " + + "and it will open the port"; + /// Why a surface is loopback-only, when the reason is a DEGRADE worth printing. A plain /// loopback-by-default config is the normal case and gets no note. private static string? DescribeLoopbackReason(DarlingMcpHostService.McpBindReason reason, string section) => @@ -2749,13 +3289,6 @@ public static IReadOnlyList PlanFirewallRules(DarlingConfig co _ => null, }; - /// - /// Creates or removes every scoped Darling firewall rule so the live firewall matches darling.json. - /// Requires elevation (that is the entire point of the verb) and is idempotent — safe to re-run on every - /// upgrade, which is exactly how install-darling.ps1 uses it. Returns 0 when the firewall ends up matching - /// the config, 1 when it could not be made to. - /// - [SupportedOSPlatform("windows")] /// /// One target of : a path, whether the interactive operator legitimately reads it, /// and what it is called in the report. Kept as data so the list is readable as a policy rather than as @@ -2949,6 +3482,119 @@ private static IEnumerable SafeEnumerate(string? directory, string patte } } + /// What the NOT-elevated --configure-firewall branch owes the operator for one surface + /// (#2445), and therefore whether an elevated shell really has work to do. + public enum FirewallHandoff + { + /// Nothing to hand over — the desired state already holds, or this run is not entitled to ask + /// for the change. Prints no command and does not make the verb fail. + Nothing, + + /// A rule has to be CREATED. Prints the enable command and fails the verb, exactly as this + /// branch always has for an Open plan. + OpenCommand, + + /// A rule the desired state forbids was FOUND by the read-only probe. Prints the same sweep + /// command the elevated path would have run, and fails the verb — there is measured elevated work. + SweepCommand, + + /// The probe could not answer, so whether such a rule exists is unknown. Prints the sweep + /// command, because it is a no-op when there is nothing to remove and the operator can only act on what + /// they are handed — but does NOT fail the verb. An exit code here is a claim about the FIREWALL, and + /// "the probe did not answer" is not one. + SweepCommandUnverified, + } + + /// + /// PURE: what the not-elevated branch owes ONE plan. is the read-only probe's + /// answer for this surface's wildcard — null when it could not answer — and is only ever consulted for a + /// Remove, because only there does the answer change anything. + /// + /// Why the Remove half is measured and the Open half is not. An Open plan already knows the + /// verb's whole job: the rule has to exist with these exact parameters, the enable command is idempotent + /// (it removes its own DisplayName first), and a probe could at best say a rule with that name exists — + /// not that its port, direction and RemoteAddress are the ones this config asks for. So the probe would buy + /// nothing there and this keeps the pre-#2445 behaviour exactly. A Remove is the opposite: the desired + /// state is the ABSENCE of a rule, absence is the overwhelmingly common case (every default loopback + /// install), and the two things this branch has to decide — whether to print a command at all, and whether + /// to report failure — are BOTH answered by "is one actually there". The alternative considered and + /// rejected was gating on Note is not null, which is a proxy for "a rule might exist" and is wrong in + /// both directions: it prints a sweep for a fail-closed surface that never had a rule, and stays silent + /// about a plain-loopback surface holding a stale rule from an exposure someone turned off by hand — the + /// exact state exists to name. + /// + /// Why a withheld sweep hands over nothing at all. #2436 withholds + /// for a surface darling.json exposes on a run that could not + /// read the control plane, because there the file's "switched off" may be years stale and the rule matching + /// the wildcard may be the one the endpoint is being served on. Printing that sweep command would hand the + /// operator, by copy and paste, the very outage the elevated path just declined to cause — and the operator + /// pasting it has no way to see that it was declined. The plan's is + /// printed above and says why; the remedy is a re-run with the store up, not a command. + /// + public static FirewallHandoff ClassifyNotElevatedHandoff( + FirewallRuleAction action, bool sweepOtherPorts, bool? rulesFound) => + action == FirewallRuleAction.Open ? FirewallHandoff.OpenCommand + : !sweepOtherPorts ? FirewallHandoff.Nothing + : rulesFound switch + { + true => FirewallHandoff.SweepCommand, + false => FirewallHandoff.Nothing, + _ => FirewallHandoff.SweepCommandUnverified, + }; + + /// + /// PURE: whether a not-elevated run has to report failure — true only where an elevated shell really has + /// work to do, which is what the exit code is supposed to mean. + /// deliberately does not count. A non-zero exit + /// from this verb is a signal other things act on — install-darling.ps1 prints a re-run banner on + /// it — and turning "the probe did not answer" into that signal would report drift the run never saw. + /// Under-reporting there is the safe direction: the running service probes the same rule on every start + /// and WARNs a stale one by name (), so an unmeasured rule is found + /// within one service start, whereas a bogus failure teaches operators to ignore the banner. + /// + public static bool NotElevatedRunHasWork(IEnumerable handoffs) => + handoffs.Any(h => h is FirewallHandoff.OpenCommand or FirewallHandoff.SweepCommand); + + /// + /// Read-only "does ANY rule for this surface exist" for the not-elevated branch (#2445). True/false when the + /// probe answered, null when it could not. Never writes anything and never throws except on cancellation. + /// + /// It runs — the probe the RUNNING service + /// already uses, shaped to exit 0 whether or not a rule matches — against + /// , which is the same wildcard the elevated sweep + /// deletes by. That pairing is the point rather than convenience: a narrower question would report "nothing + /// to do" about rules the sweep would still collect, and a wider one would offer a command covering rules + /// this product does not own. Probe and remedy share the builders, so they cannot come to disagree. + /// + /// Affordable precisely HERE, in the one branch that by definition holds no privilege: + /// Get-NetFirewallRule needs no elevation — reads succeed under a restricted token where writes + /// return PermissionDenied — and this costs at most one bounded PowerShell round trip per surface, only on + /// a path install-darling.ps1 cannot take, since it refuses to run unelevated at all. + /// + [SupportedOSPlatform("windows")] + private static async Task ProbeSurfaceRulesAsync(string wildcard, CancellationToken cancellationToken) + { + try + { + var (exitCode, psOutput) = await DarlingManagedPostgres.RunPowerShellAsync( + DarlingFirewallCheck.BuildProbeCommand(wildcard), cancellationToken); + + /* The probe exits 0 for BOTH answers, so a non-zero exit means the probe itself broke; do not read + a count out of a run that failed. Same reasoning as DarlingFirewallCheck.CheckAsync. */ + return exitCode == 0 ? DarlingFirewallCheck.TryParseProbeOutput(psOutput) : null; + } + catch (OperationCanceledException) + { + throw; + } + catch (Exception) + { + /* A missing or broken powershell.exe is "could not tell", never "no rule" — guessing absence would + manufacture a clean bill of health for a firewall this run never read. */ + return null; + } + } + /// /// Creates or removes every scoped Darling firewall rule so the live firewall matches darling.json. /// Requires elevation (that is the entire point of the verb) and is idempotent — safe to re-run on every @@ -2970,7 +3616,15 @@ public static async Task ConfigureFirewallAsync( return 1; } - var plans = PlanFirewallRules(config); + /* #2414: ask the control plane which ports the MCP/web endpoints are actually bound to, because that — + not darling.json's seed — is what the rule has to be named for. BEST EFFORT on purpose. This verb is + called by install-darling.ps1 before the service has ever run, when there is no store to ask and + darling.json's port is the only truth there is (the store row is SEEDED from it), so requiring a + reachable store would fail every fresh install. What it must not do is fall back QUIETLY: the plan + carries a PortNote that names the port it used, why it could not do better, and what makes that port + wrong, printed below on the same footing as the rules themselves. */ + var (mcpStore, webStore, storeUnavailable) = await TryReadEndpointTogglesAsync(config, cancellationToken); + var plans = PlanFirewallRules(config, mcpStore, webStore, storeUnavailable); var toOpen = plans.Count(p => p.Action == FirewallRuleAction.Open); output.WriteLine("Reconciling the scoped Windows Firewall rules to match darling.json."); @@ -2979,26 +3633,116 @@ public static async Task ConfigureFirewallAsync( "running service only verifies them and reports what it finds."); output.WriteLine(); + /* Print the port provenance BEFORE the not-elevated branch below, so the operator who is about to paste + these commands by hand can see which plane chose the ports baked into them. */ + var portNotes = plans.Select(p => p.PortNote).Where(n => n is not null).ToList(); + foreach (var note in portNotes) + { + output.WriteLine(note); + } + + if (portNotes.Count > 0) + { + output.WriteLine(); + } + if (!IsElevated()) { /* Not elevated. When nothing is exposed there is genuinely nothing to do and a hard failure would be a lie (and would fail an otherwise fine loopback install); when something IS exposed this is a real, actionable failure, so exit non-zero AND print every command to run by hand. */ - if (toOpen == 0) + /* #2436: the per-surface notes, for EVERY plan and before either branch below. The elevated path + prints them as it works; this path used to print none at all, so the operator who most needs + them — the one who cannot act on them yet — was the one who never saw them. They are also the + only place a surface says it is fully LAN-configured and switched OFF, which is a state that did + not exist here before this change: "no rule is needed" would otherwise read as "your exposure + config did not take", and in the mixed case a disabled surface's stale rule would be dropped + silently while its exposed sibling got a command. */ + foreach (var plan in plans.Where(p => p.Note is not null)) + { + output.WriteLine($"{plan.Surface}: {plan.Note}."); + } + + /* #2445: MEASURE the Remove half instead of guessing at it. This branch used to print a command + only for Open plans, so a run whose whole work was a removal said "nothing needs an open port", + exited 0, and left a stale inbound allow rule sitting there with nothing offered to close it. + #2436 made that shape ordinary rather than exotic — a surface is now a Remove when the control + plane has it switched OFF, so it covers an admin who just ran --disable-mcp and re-ran this verb + from a normal prompt. + + The read-only probe is what makes all three behaviours honest at once, and it is affordable here + for a reason worth stating: Get-NetFirewallRule needs no elevation, so this branch could always + have looked and simply never did. Without it the only gate available is a proxy for "a rule might + exist", which would print sweep commands on every default loopback install and STILL miss a stale + rule on a surface that has no note. With it, a clean box prints nothing extra and exits 0 exactly + as before, and a box with a stale rule gets that rule's own sweep command and a non-zero exit. + Only the Remove plans that are permitted to sweep are probed, so the cost is at most one bounded + PowerShell round trip per such surface, on a path install-darling.ps1 cannot reach — it fails + outright when not elevated, so the re-run banner this could have fired is not on that path. */ + var handoffs = new List<(FirewallRulePlan Plan, FirewallHandoff Handoff)>(); + foreach (var plan in plans) + { + bool? rulesFound = null; + if (plan.Action == FirewallRuleAction.Remove && plan.SweepOtherPorts) + { + rulesFound = await ProbeSurfaceRulesAsync( + DarlingFirewallCheck.SurfaceRuleWildcard(plan.RuleName), cancellationToken); + } + + handoffs.Add((plan, ClassifyNotElevatedHandoff(plan.Action, plan.SweepOtherPorts, rulesFound))); + } + + if (handoffs.All(h => h.Handoff == FirewallHandoff.Nothing)) { output.WriteLine( - "Every endpoint is loopback-only, so no firewall rule is needed and none was changed. " + - "(This shell is not elevated, but there was nothing to do.)"); + "No endpoint wants an open port, and no scoped Darling rule is open that should not be, so " + + "there is nothing for an elevated shell to do. (This shell is not elevated, but reading the " + + "firewall does not need elevation — so that is measured, not assumed.)"); return 0; } + var hasWork = NotElevatedRunHasWork(handoffs.Select(h => h.Handoff)); + error.WriteLine("This shell is not elevated, so NO firewall rule was changed. Run these in an ELEVATED PowerShell:"); - foreach (var plan in plans.Where(p => p.Action == FirewallRuleAction.Open)) + foreach (var (plan, handoff) in handoffs) { - error.WriteLine(" " + DarlingManagedPostgres.BuildFirewallEnableCommand(plan.RuleName, plan.Port, plan.Cidr!)); + switch (handoff) + { + case FirewallHandoff.OpenCommand: + error.WriteLine(" " + DarlingManagedPostgres.BuildFirewallEnableCommand(plan.RuleName, plan.Port, plan.Cidr!)); + break; + + /* The same builder and the same wildcard the elevated path would have used, so the command + an operator pastes and the command the verb runs cannot drift apart. */ + case FirewallHandoff.SweepCommand: + error.WriteLine( + $" # {plan.Surface}: a scoped rule is open on this surface and none belongs on any port."); + error.WriteLine(" " + DarlingManagedPostgres.BuildFirewallSweepCommand( + DarlingFirewallCheck.SurfaceRuleWildcard(plan.RuleName))); + break; + + case FirewallHandoff.SweepCommandUnverified: + error.WriteLine( + $" # {plan.Surface}: the read-only rule probe gave no usable answer, so this may be a no-op."); + error.WriteLine(" " + DarlingManagedPostgres.BuildFirewallSweepCommand( + DarlingFirewallCheck.SurfaceRuleWildcard(plan.RuleName))); + break; + + case FirewallHandoff.Nothing: + default: + break; + } } - return 1; + if (!hasWork) + { + error.WriteLine( + "Returning 0: nothing above is CONFIRMED work. The read-only probe could not answer for at " + + "least one surface, so those commands are offered as a precaution — and a non-zero exit is a " + + "claim about the firewall, which this run has no measurement to support."); + } + + return hasWork ? 1 : 0; } var failures = 0; @@ -3013,17 +3757,43 @@ public static async Task ConfigureFirewallAsync( name, so changing a port does not update a rule — it makes a different one and strands the old as an inbound allow rule on a port nothing serves. Reconciling by exact name could never reach that. Its own step, not concatenated ahead of the open below, because the sweep command ends in exit 0 - and would otherwise terminate the shell before the rule was created. */ + and would otherwise terminate the shell before the rule was created. + + #2436: and the sweep's authority to remove OTHER ports comes from knowing which port is live. On + an Open plan whose port fell back to darling.json's seed because the store could not answer, that + is precisely the knowledge missing — a rule on another port is as likely to be the one the + endpoint is being served on as it is to be stale, and removing it turns the elevated verb an + operator was told to run into an outage of the surface it was meant to repair. That is not + hypothetical: install-darling.ps1 runs this verb with the service STOPPED, so on an upgrade of a + box whose port was moved in the Viewer it is the working rule that matched the wildcard. The + enable command below removes its own exact DisplayName first, so declining the sweep costs + nothing in idempotence — it defers the cleanup to a run that can see the control plane, which is + what the installer's post-start reconcile is for. */ var wildcard = DarlingFirewallCheck.SurfaceRuleWildcard(plan.RuleName); - if (!await TryRunFirewallStepAsync( - DarlingManagedPostgres.BuildFirewallSweepCommand(wildcard), - plan.Action == FirewallRuleAction.Remove - ? $"{plan.Surface}: loopback-only — no rule needed (removed any '{wildcard}')." - : $"{plan.Surface}: cleared any previous rule matching '{wildcard}'.", - $"{plan.Surface}: could not remove the rule(s) matching '{wildcard}'", - output, error, cancellationToken)) + if (plan.SweepOtherPorts) { - failures++; + if (!await TryRunFirewallStepAsync( + DarlingManagedPostgres.BuildFirewallSweepCommand(wildcard), + plan.Action == FirewallRuleAction.Remove + ? $"{plan.Surface}: no rule belongs on any port (removed any '{wildcard}')." + : $"{plan.Surface}: cleared any previous rule matching '{wildcard}'.", + $"{plan.Surface}: could not remove the rule(s) matching '{wildcard}'", + output, error, cancellationToken)) + { + failures++; + } + } + else + { + output.WriteLine(plan.Action == FirewallRuleAction.Open + ? $"{plan.Surface}: leaving any rule on ANOTHER port alone. This run could not read the control " + + "plane, so it cannot tell a rule stranded by a port change from the rule the endpoint is " + + "actually being served on — re-run --configure-firewall with the service up to collect it." + : $"{plan.Surface}: leaving this surface's existing rules alone rather than removing them. " + + "darling.json exposes this endpoint and says it is switched off, but the --enable-* verbs and " + + "the Viewer write only the control plane, which this run could not read — so that off may be " + + "stale and the rule may be the one the endpoint is being served on. Nothing is listening while " + + "the service is stopped; re-run --configure-firewall with it up to reconcile."); } if (plan.Action == FirewallRuleAction.Remove) @@ -3049,12 +3819,92 @@ and would otherwise terminate the shell before the rule was created. */ return 1; } + /* Not "every endpoint is loopback-only" any more: since #2436 a surface is also a Remove when it is + fully LAN-configured and the control plane has it switched OFF. That is the same misleading summary + #2442 fixed one branch up, and the per-surface lines above already carry the real reason. */ output.WriteLine(toOpen == 0 - ? "Done. Every endpoint is loopback-only, so no port was opened." + ? "Done. No endpoint wants an open port, so no port was opened." : $"Done. {toOpen} endpoint(s) exposed on the LAN; their scoped rules are in place."); return 0; } + /// + /// BEST-EFFORT read of the control plane's effective endpoint toggles for --configure-firewall (#2414). + /// Returns nulls plus a human reason on every failure and NEVER throws except on caller cancellation. + /// + /// Why best-effort and not required. Every other consumer of these values runs inside the + /// service, where the store is a precondition. This verb is the opposite: install-darling.ps1 calls it while + /// elevated, before the service has ever started, on a box where config.config_service does not exist + /// yet — and it is correct there, because the store row is SEEDED from darling.json's ports, so at that moment + /// the file IS the control plane's future answer. Refusing would fail every fresh networked install to close a + /// window in which the file cannot be wrong. + /// + /// Why it must still say so. The same fallback on a box whose port HAS been moved in the Viewer is + /// the whole of #2414: a rule for a port nothing serves, recreated on every run of the verb that is supposed to + /// fix it. So the reason travels back with the nulls and is printed, rather than being swallowed into a silent + /// default. Bounded to ten seconds because "the store is down" and "the store is slow" must not differ in how + /// long an installer hangs. + /// + [SupportedOSPlatform("windows")] + private static async Task<((bool Enabled, int Port)? Mcp, (bool Enabled, int Port)? Web, string? Unavailable)> + TryReadEndpointTogglesAsync(DarlingConfig config, CancellationToken cancellationToken) + { + var postgres = config.Postgres; + if (postgres is null || !postgres.Managed) + { + /* BYO: config_service lives in the operator's own PostgreSQL and this verb holds no credential for it. + Harmless in practice — LAN exposure for MCP/web is managed-mode only, so both surfaces plan Remove, + which sweeps every port and needs no port at all. */ + return (null, null, + "postgres.managed is false, so config.config_service lives in YOUR PostgreSQL and this verb has no credential for it"); + } + + string? connectionString; + try + { + connectionString = DarlingManagedPostgres.TryBuildConnectionStringFromStoredCredential(postgres); + } + catch (Exception ex) + { + return (null, null, $"the stored store credential could not be read ({ex.Message})"); + } + + if (connectionString is null) + { + return (null, null, + "the service has not initialized the managed store yet, so there is no owner credential to read it with"); + } + + using var budget = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken); + budget.CancelAfter(TimeSpan.FromSeconds(10)); + + try + { + await using var connection = new NpgsqlConnection(connectionString); + await connection.OpenAsync(budget.Token); + await using var command = new NpgsqlCommand(ReadEndpointTogglesSql, connection); + await using var reader = await command.ExecuteReaderAsync(budget.Token); + if (!await reader.ReadAsync(budget.Token)) + { + return (null, null, "config.config_service has no id = 1 row yet — the service has never seeded the store"); + } + + return ((reader.GetBoolean(0), reader.GetInt32(1)), (reader.GetBoolean(2), reader.GetInt32(3)), null); + } + catch (OperationCanceledException) when (!cancellationToken.IsCancellationRequested) + { + return (null, null, "the store did not answer within 10 seconds"); + } + catch (OperationCanceledException) + { + throw; + } + catch (Exception ex) + { + return (null, null, $"the store did not answer ({ex.Message})"); + } + } + /// Runs one reconcile step and reports it. Returns false on failure, after printing the exact /// command so an operator can finish by hand. Never throws except on cancellation. [SupportedOSPlatform("windows")] diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingCollectorRunner.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingCollectorRunner.cs index d0dd78424..23b76fe7e 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingCollectorRunner.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingCollectorRunner.cs @@ -32,7 +32,13 @@ namespace PerformanceMonitor.Darling.Service; /// Darling twin of Lite's _lastCollectionNote; null (the default) leaves the row's message column /// null exactly as before. /// -public sealed record CollectorRunResult(int Rows, long SqlMs, long StorageMs, string? Note = null); +/// +/// The per-database rollup for a run that fanned out, null for one that did not (#2472). Defaulted because +/// the great majority of construction sites here are the early returns of runs that never reached a fan-out +/// — a single query, an enumeration that yielded nothing — and null is their correct answer rather than a +/// value they forgot to supply. The one site that MUST set it is the success return. +/// +public sealed record CollectorRunResult(int Rows, long SqlMs, long StorageMs, string? Note = null, FanoutCost? Fanout = null); /// /// Runs a shared collector definition against one monitored server and binary-COPYs the rows @@ -181,6 +187,40 @@ public DarlingCollectorRunner(NpgsqlDataSource postgres, CollectorDeltaCalculato _compressPlanContent = compressPlanContent ?? (() => true); } + /* One ingestor for the process, so the resume marker survives between cycles - it is per-file and + in-memory by design (#2538), and a fresh instance every cycle would silently re-read the same tail + forever while looking like it was making progress. */ + private RdsPlanIngestor? _rdsPlans; + + /// + /// Plan capture for Aurora and RDS, where the log is only reachable through the AWS API (#2538). + /// + /// Reported as a normal so the cycle accounts for it exactly like + /// a collector: same collection_log row, same rows-collected number, same health surface. The + /// TRANSPORT differs; the bookkeeping should not, or an operator would have to know which route a + /// target used before they could read its collection history. + /// + public async Task IngestRdsPlansAsync( + ServerRuntime server, CancellationToken cancellationToken) + { + _rdsPlans ??= new RdsPlanIngestor(_postgres, logger: _logger); + + var host = new NpgsqlConnectionStringBuilder(server.ConnectionString).Host ?? string.Empty; + + var started = Stopwatch.GetTimestamp(); + + var rows = await _rdsPlans.IngestAsync( + server.ServerId, server.StorageName, host, cancellationToken); + + var elapsedMs = (long)Stopwatch.GetElapsedTime(started).TotalMilliseconds; + + /* Counted as STORAGE time rather than SQL time: no query ran against the monitored server, and + filing an HTTPS round trip under sql_duration_ms would make one target's numbers mean something + different from every other target's. */ + return new CollectorRunResult(rows, 0, elapsedMs, + rows == 0 ? "no new auto_explain plans in the RDS log window" : null); + } + public async Task RunAsync( ICollectorDefinition definition, ServerRuntime server, @@ -328,6 +368,13 @@ constant. Converted MB -> bytes here so the store knob stays operator-friendly. long storageMs = 0; var rowsWritten = 0; + /* The per-database rollup (#2472). Both fan-out shapes feed it — the enumeration driver's + onItemComplete hook and the Azure per-database connection loop — so a collector that fans out on + one branch on Azure and the other on-prem reports the same shape either way. A run that never + fans out never calls Observe and the accumulator stays empty, which is how the columns end up + NULL on ~98 percent of collection_log rows. */ + var fanout = new FanoutCostAccumulator(); + /* The collection_log note for this run (#1837) — null on every ordinary path. Only the enumeration branch sets it, but it is declared here so the note reaches the single success return below when items WERE found and merely some of their probes failed. Lite's twin is _lastCollectionNote. */ @@ -378,6 +425,11 @@ turn a permissions problem into a silent one-database collection. */ var failed = 0; Exception? firstFailure = null; + /* #2623: the names, not just the count. A partial loss composes a note naming which databases + were skipped, because the count alone does not tell an operator whether the ONE database + that matters is in the collected set or the skipped one. */ + var failedDatabases = new List(); + /* #1875: this path reads the trailing probe-failure set once PER DATABASE, so the note and the log cap are decided for the cycle after the loop rather than inside it — see CycleProbeFailures for why neither generalizes from the single-read plain path. */ @@ -417,7 +469,7 @@ into a sibling database's landing. */ documented first-run window, per database. No clamp is applied HERE because this branch also serves the XE ring-buffer collectors (deadlocks / BPR), where flooring a stale watermark would WRONGLY truncate legitimate catch-up - — those sources roll past 24h on their own. query_store also reaches this + — those sources roll past the catch-up horizon on their own. query_store also branch on Azure SQL DB (#1836) and does need the bound, so it applies WatermarkPolicy.ClampCatchup inside its own cutoff computation: the clamp travels with the collector that needs it instead of with the path. */ @@ -540,16 +592,30 @@ set and the loop simply never advanced the reader to it — the rows were built await EnumeratedCollectorDriver.ReadPayloadProbeFailuresAsync(dbReader, dbToken)); } } - sqlMs += sqlSlice.ElapsedMilliseconds; + /* Read ONCE. The stopwatch is still running, so a second read a few statements later + returns a larger number, and the per-item total would then exceed the blended total + it is a ratio against — a dominance a hair above the truth, on every Azure run + (#2472). Small, and wrong in the direction that matters. */ + var dbSqlMs = sqlSlice.ElapsedMilliseconds; + sqlMs += dbSqlMs; /* Flush this database before reading the next — peak memory is one database's rows. */ + long dbStorageMs = 0; if (batch.Count > 0) { var storageSlice = Stopwatch.StartNew(); rowsWritten += await WriteBatchAsync(pgConnection, definition, batch, server, collectionTime, context, cancellationToken); - storageMs += storageSlice.ElapsedMilliseconds; + dbStorageMs = storageSlice.ElapsedMilliseconds; + storageMs += dbStorageMs; } + /* #2472: this database's slice, counted even when its batch was empty — an empty batch + still paid for its read, and that read is in the blended total the rollup is a ratio + against. Observed here rather than beside the log line below for the same reason the + completion hook fires after the flush: both slices are only known once the write is + done. */ + fanout.Observe(databaseName, dbSqlMs + dbStorageMs); + /* Same per-database bounded-cycle WARNING the enumeration path emits from onItemComplete, mirroring Lite. Reachable here since #1836 put query_store — the only collector that declares either bound — on this branch for Azure SQL DB; @@ -604,6 +670,7 @@ context signal stays this database's until the next read resets it. */ var budgetFailure = EnumeratedCollectorDriver.ItemBudgetException( definition.PerItemWallClockBudget!.Value); failed++; + failedDatabases.Add(databaseName); firstFailure ??= budgetFailure; /* Same #2111 stamp the generic arm makes, and it MATTERS more here: this is what turns @@ -626,6 +693,7 @@ database is ordinary and this is a collector that could not finish its work. */ /* OOM is filtered OUT of this per-database skip and propagates: it is fatal, not a routine one-database miss. */ failed++; + failedDatabases.Add(databaseName); firstFailure ??= ex; /* #2111: the yield-to-live stamp + adaptive-shrink count for the Azure SQL DB @@ -647,7 +715,10 @@ guard as the hole recording above. */ /* #1875: ONE note for the cycle and ONE capped log burst, composed from every database's failures together. Assigned unconditionally — a cycle where nothing failed composes null, which is exactly what this path carried before. */ - collectionNote = cycleProbeFailures.Note; + collectionNote = EnumeratedCollectorDriver.MergeNotes( + cycleProbeFailures.Note, + EnumeratedCollectorDriver.BuildPartialFailureNote( + failed, attempted, failedDatabases, firstFailure?.Message)); LogEnumerationProbeFailures(definition, server, cycleProbeFailures.Failures); /* One database failing is routine (offline, mid-restore, a permissions oddity) and stays a @@ -738,7 +809,7 @@ cycle that captured nothing (the review catch): PendingState flushes as long as var driverResult = await EnumeratedCollectorDriver.RunAsync( items, - /* Per-database watermark refresh + the 24h catch-up clamp, computed INSIDE the loop — + /* Per-database watermark refresh + the catch-up clamp, computed INSIDE the loop — this is the per-item cutoff site the plan's LOUD FLAG requires the clamp to live at. Only query_store (the sole enumeration collector with a per-database timestamp watermark) reaches this; the two snapshot collectors are watermark-less. */ @@ -909,6 +980,12 @@ await FetchAndStoreQueryTextAsync(textFetchConnection, writeBatch: (batch, ct) => WriteBatchAsync(pgConnection, definition, batch, server, collectionTime, context, ct), onItemComplete: (item, batchCount, itemSqlMs, itemStorageMs) => { + /* #2472: the per-database cost the blended collection_log row cannot carry. Counted + for every completed item, including the quiet ones the log line below skips — + their read time is in the blended total, so leaving them out would inflate the + dominance ratio of whichever database happened to have rows. */ + fanout.Observe(item, itemSqlMs + itemStorageMs); + /* #2111: a completed item resets the adaptive-shrink count — recovery returns the member to the full catch-up width on its next cycle. */ if (string.Equals(definition.Name, QueryStoreCollector.Instance.Name, StringComparison.Ordinal)) @@ -1085,7 +1162,7 @@ await SaveCollectorStateAsync( _logger?.LogDebug("Collected {RowCount} {Collector} rows for server '{Server}'", rowsWritten, definition.Name, server.Config.DisplayName); - return new CollectorRunResult(rowsWritten, sqlMs, storageMs, collectionNote); + return new CollectorRunResult(rowsWritten, sqlMs, storageMs, collectionNote, fanout.Result); } /// diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingCommandExecutor.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingCommandExecutor.cs index f6821fcaa..64c45c8c6 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingCommandExecutor.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingCommandExecutor.cs @@ -308,6 +308,21 @@ the structured request — mirrors the FetchPlan branch. The payload is an IDENT return await _host.ExecuteActualPlanAsync(command.TargetServerId!.Value, actualPlanRequest, cancellationToken); } + case CommandKind.TestHypotheticalIndex: + { + /* Re-parsed here so the host receives the structured request, mirroring the two branches + above. The payload is an IDENTIFIER plus a candidate; the host resolves the statement + text from the store and never takes it from the caller. */ + if (!HypotheticalIndexRequest.TryParse(command.ArgsJson, out var hypotheticalRequest)) + { + return new CommandOutcome(false, "invalid args_json", + ErrorJson("test_hypothetical_index args_json was missing or malformed")); + } + + return await _host.TestHypotheticalIndexAsync( + command.TargetServerId!.Value, hypotheticalRequest, cancellationToken); + } + case CommandKind.FetchActiveQueries: /* Worker-delegated like fetch_plan (needs the target's LIVE connection to read the running-request DMVs). Takes NO args — the target_server_id on the command row is the whole request; the host @@ -409,6 +424,24 @@ every other command — a read-only viewer cannot reach it. */ ? new CommandPlan(CommandKind.ExecuteActualPlan, null, null, "actual plan captured", null) : Fail("execute_actual_plan requires args_json with a query_hash"); + case "test_hypothetical_index": + /* Worker-delegated (needs the target's LIVE connection to plan against its statistics) and the + store (to resolve the statement text from the queryid). #2612: the ONE PostgreSQL command + that makes the product act rather than read, so it rides the same read-write-seat-only + enqueue as execute_actual_plan — but it is a lesser class than that one, and the message + says so rather than leaving the reader to infer it: EXPLAIN without ANALYZE does not run + the statement, a hypothetical index is never written anywhere, and both are undone before + the call returns. */ + if (command.TargetServerId is null) + { + return Fail("test_hypothetical_index requires target_server_id"); + } + + return HypotheticalIndexRequest.TryParse(command.ArgsJson, out _) + ? new CommandPlan(CommandKind.TestHypotheticalIndex, null, null, "hypothetical index tested", null) + : Fail("test_hypothetical_index requires args_json with queryid, schemaName, tableName and a " + + "non-empty columns array of plain identifiers"); + case "fetch_active_queries": /* Worker-delegated like fetch_plan (needs the target's LIVE runtime connection to read what is running now). Requires a target server; takes NO args_json (the whole request is "read this @@ -629,6 +662,11 @@ public enum CommandKind /// the host, running the shared Active Queries collector's query) and return the rows in result_json. Read-only. FetchActiveQueries, + /// test_hypothetical_index: plan one stored statement twice on a PostgreSQL target — with and + /// without a candidate index the planner can see but nothing builds — and report whether it would be used + /// (#2612). Nothing is executed, nothing is written, and the experiment is undone before the call returns. + TestHypotheticalIndex, + /// Bad arguments or an unknown command_type — report failed without touching the store. Fail, } @@ -676,6 +714,15 @@ public interface IDarlingCommandHost /// Task ExecuteActualPlanAsync(int serverId, ActualPlanRequest request, CancellationToken cancellationToken); + /// + /// Plans one stored PostgreSQL statement with and without a candidate index (#2612). + /// + /// Consent-class like in that it makes the product ACT on a + /// monitored server, and unlike it in what that costs: no statement is executed, nothing is written, + /// the candidate exists only inside one session, and the session is cleaned before this returns. + /// + Task TestHypotheticalIndexAsync(int serverId, HypotheticalIndexRequest request, CancellationToken cancellationToken); + /// /// fetch_active_queries: read the LIVE running-request DMV snapshot from the target server on demand /// (the shared Active Queries collector's query, run on the server's runtime connection) and return the rows diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs index 2c153efe5..995067992 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingConfig.cs @@ -184,7 +184,8 @@ public sealed class DarlingConfig /// from the MCP server (they gate different blast radii). Default OFF: a headless service should not open a /// port unless the operator asks. Loopback-only by default; an opt-in block /// exposes it on the LAN behind a required token (once) → an HMAC session cookie + an in-app CIDR check - /// (managed-mode only). Loopback stays tokenless even while LAN-exposed (the web surface is read-only). + /// (managed-mode only), optionally over TLS (, #2562). While LAN-exposed EVERY + /// request authenticates, loopback included — loopback is exempt from the CIDR test only (#1649). /// Optional — omit the section entirely. /// [JsonPropertyName("web")] @@ -721,10 +722,17 @@ public sealed class SmtpConfig /// public sealed class McpConfig { - /// Default OFF — the headless twin of both apps' mcp_enabled=false default. + /// Default OFF — the headless twin of both apps' mcp_enabled=false default. + /// A first-run SEED only (#2389): once the store is seeded, config.config_service.mcp_enabled + /// is authoritative and this value is never read again except before the worker's first publish. Editing it + /// on a seeded box changes nothing — use --enable-mcp/--disable-mcp or the Viewer's Settings. + /// , sitting right beside it, is the OPPOSITE: file-only with no store equivalent. + /// The supervisor reports the disagreement when the two planes differ. [JsonPropertyName("enabled")] public bool Enabled { get; set; } + /// A first-run SEED only, like — config.config_service.mcp_port wins + /// on a seeded store. [JsonPropertyName("port")] public int Port { get; set; } = 5152; @@ -913,10 +921,16 @@ public sealed class WebConfig /// Omit the whole network object for the secure default = loopback-only HTTP. When present with a /// non-loopback AND managed mode AND a token AND a valid , the web /// host binds the network interface behind an in-app CIDR check and a token gate: a browser presents the token -/// once via ?token=, which is exchanged for an HMAC-signed session cookie. Loopback requests stay -/// TOKENLESS even while exposed (the dashboard is read-only). Any missing precondition keeps the dashboard -/// loopback-only + LogCritical (fail-closed, enforced in the web host). No TLS on the dashboard (the same -/// reverse-proxy story as MCP); the token/cookie travels cleartext on-segment. +/// once via ?token=, which is exchanged for an HMAC-signed session cookie. While exposed EVERY request +/// authenticates, loopback included — loopback is exempt from the CIDR test only, never from the credential +/// (#1649). Any missing precondition keeps the dashboard loopback-only + LogCritical (fail-closed, enforced in +/// the web host). +/// +/// TLS is opt-in via (#2562). Without it the network listener is plain HTTP and +/// the token and session cookie cross the segment in the clear — the web host warns about exactly that at +/// every exposed start. MCP still has no TLS of its own on the older rationale that a self-signed certificate +/// breaks real MCP clients; that argument is about MCP clients rather than about the wire, and it does not +/// carry to a surface whose only client is a browser. /// public sealed class WebNetworkConfig { @@ -942,12 +956,21 @@ public sealed class WebNetworkConfig public string? EncryptedToken { get; set; } /// - /// Plaintext access token — dev convenience only; the caller warns when it is used. Prefer - /// . + /// The access token as a literal (dev convenience only; the caller warns) or an env:/file: + /// reference (#1804 — , which does not count as plaintext-in-config, and + /// is how the compose distribution mounts it). Prefer on Windows. /// [JsonPropertyName("token")] public string? Token { get; set; } + /// + /// Opt-in TLS for the network listener (#2562). Omit for plain HTTP, which is the zero-config default and + /// the right answer on loopback; supply a certificate before exposing the dashboard on a segment where the + /// access token crossing in the clear matters. See . + /// + [JsonPropertyName("tls")] + public WebTlsConfig? Tls { get; set; } + /// /// True when any field is set — used only for the BYO "network.* is ignored" caller warning (D-BYO); /// NOT the same as "exposed". @@ -957,7 +980,8 @@ public sealed class WebNetworkConfig !string.IsNullOrWhiteSpace(Listen) || !string.IsNullOrWhiteSpace(AllowFrom) || !string.IsNullOrWhiteSpace(EncryptedToken) - || !string.IsNullOrWhiteSpace(Token); + || !string.IsNullOrWhiteSpace(Token) + || (Tls?.IsConfigured ?? false); /// /// The access token, preferring (DPAPI-decrypted; Windows-only) over the @@ -993,6 +1017,113 @@ public sealed class WebNetworkConfig } } +/// +/// Opt-in TLS for the web dashboard's network listener (#2562). Omit the whole tls object for plain +/// HTTP — the zero-config default, and the correct one for a loopback-only dashboard, which has nothing to +/// encrypt. Supply a certificate to close the gap the exposure block otherwise leaves open: the access token +/// and the HMAC session cookie it is exchanged for both cross the segment in the clear over HTTP, and the +/// in-app CIDR check bounds who can ROUTE to the port, never what an on-path attacker can read off the wire. +/// +/// Two forms, exactly one at a time. A PKCS#12 bundle (, with the password +/// in whichever of the three slots suits the platform) or a PEM pair ( + +/// , which is what a container mounts). Configuring both is refused rather than resolved +/// by precedence — see . +/// +/// The product consumes a certificate; it does not manage a PKI. No issuance, no ACME, no +/// self-signed fallback: a certificate the product minted itself would buy encryption without authentication +/// and train the operator to click through a browser warning, which is a worse habit than the plain HTTP it +/// replaces. An internal CA is the normal answer on the LAN this feature is for. +/// +/// Fail-closed, like every other exposure precondition. A missing, unreadable, mismatched or +/// EXPIRED certificate keeps the dashboard loopback-only and logs Critical. It never falls back to serving +/// the LAN over HTTP: an operator who configured TLS and silently got cleartext would be in precisely the +/// state this block exists to prevent. +/// +public sealed class WebTlsConfig +{ + /// + /// Path to a PKCS#12 (.pfx/.p12) bundle holding the certificate AND its private key. + /// Mutually exclusive with /. + /// + [JsonPropertyName("pfxPath")] + public string? PfxPath { get; set; } + + /// + /// DPAPI-LocalMachine-protected password for , base64 — produced by + /// --encrypt-password, and the preferred slot on Windows. Read via + /// . + /// + [JsonPropertyName("encryptedPfxPassword")] + public string? EncryptedPfxPassword { get; set; } + + /// + /// The PKCS#12 password as a literal (dev convenience only; the caller warns) or an + /// env:/file: reference (#1804 — , which does not count as + /// plaintext-in-config). Omit entirely for a bundle that has no password. + /// + [JsonPropertyName("pfxPassword")] + public string? PfxPassword { get; set; } + + /// + /// Path to a PEM-encoded certificate (the leaf first; a chain may follow). Requires + /// , and is mutually exclusive with . + /// + [JsonPropertyName("certPath")] + public string? CertPath { get; set; } + + /// + /// Path to the PEM-encoded PRIVATE KEY for . A separate file rather than a combined + /// PEM because that is how both Docker secrets and every certificate tool in this space emit them, and + /// because the two want different file permissions. + /// + [JsonPropertyName("keyPath")] + public string? KeyPath { get; set; } + + /// True when any field is set. Drives the "this block is configured" half of the decision — a + /// block set but unusable must FAIL rather than read as "TLS was never asked for". + [JsonIgnore] + public bool IsConfigured => + !string.IsNullOrWhiteSpace(PfxPath) + || !string.IsNullOrWhiteSpace(EncryptedPfxPassword) + || !string.IsNullOrWhiteSpace(PfxPassword) + || !string.IsNullOrWhiteSpace(CertPath) + || !string.IsNullOrWhiteSpace(KeyPath); + + /// + /// The PKCS#12 password, preferring (DPAPI-decrypted; Windows-only) + /// over — the same shape as . + /// Returns null when neither is set, which is correct for a bundle with no password rather than an error. + /// + public string? ResolvePfxPassword(out bool usedPlaintext) + { + usedPlaintext = false; + + if (!string.IsNullOrWhiteSpace(EncryptedPfxPassword)) + { + /* DarlingSecrets.Unprotect is DPAPI (Windows-only). Unlike the token slots this one is reachable + from the container path too (exposure is honored in a container since #1804), so the guard is a + real branch here, not just analyzer bookkeeping: say which slot to use instead. */ + if (!OperatingSystem.IsWindows()) + { + throw new PlatformNotSupportedException( + "web.network.tls.encryptedPfxPassword requires Windows (DPAPI); use \"pfxPassword\" with a " + + "file:/env: reference on other platforms."); + } + + return DarlingSecrets.Unprotect(EncryptedPfxPassword); + } + + if (!string.IsNullOrWhiteSpace(PfxPassword)) + { + /* An env:/file: reference (#1804) is not plaintext-in-config — no warning for it. */ + usedPlaintext = !DarlingSecretSource.IsReference(PfxPassword); + return DarlingSecretSource.Resolve(PfxPassword, "web.network.tls.pfxPassword"); + } + + return null; + } +} + /// /// Shared network-exposure helpers for the opt-in store/MCP endpoints (darling-network-endpoints). /// These pure functions are the single source of truth for "is this listen value a network bind?" @@ -1037,19 +1168,90 @@ loopback because it cannot IPAddress.Parse it. */ /// default viewer (read-only remote, D7); viewer/admin pass through /// (case-insensitively); anything else ⇒ null (invalid — the store degrades to loopback). NEVER /// returns darling (the superuser/owner is service-only). + /// + /// Kept for the callers that ask a yes/no question about the exposure — "is this store reachable + /// as admin" (see DarlingWorker's startup warning). Where BOTH roles are admitted it answers + /// admin, because the question those callers are really asking is whether write-capable access + /// is reachable from the network, and it is. /// public static string? NormalizeNetworkRole(string? role) + { + var roles = NormalizeNetworkRoles(role); + + /* Count == 0 cannot happen today — the plural returns null or a non-empty list — but this is the one + caller that would INDEX the result, and a future "return the ones we understood" would turn that + invariant into an IndexOutOfRangeException at service startup rather than a degrade. Checked here + because every other caller already checks it. */ + if (roles is null || roles.Count == 0) + { + return null; + } + + return roles.Contains("admin", StringComparer.Ordinal) ? "admin" : roles[0]; + } + + /// + /// Every pg_hba login role postgres.network.role admits, in a stable order (#2665). One rule is + /// written per role, so a team can reach the same store with an admin Viewer and a set of + /// read-only viewer ones — which the roles have always supported and the single generated + /// hostssl line did not. + /// + /// Accepts "admin", "viewer", or both in one string separated by a comma, a plus or + /// whitespace ("admin,viewer", "admin+viewer"). A JSON array would be the tidier surface + /// and is not worth a breaking change to a field that has shipped: everything already written stays + /// valid, and a list is expressible without it. + /// + /// Absent/blank still means viewer alone. A field learning to take a list must not + /// become a way for an existing configuration to quietly gain admin-capable network access. + /// + /// Returns null when ANY element is unrecognised, rather than silently keeping the ones it + /// understood — the store then degrades to loopback with the reason named. A typo in one of two roles + /// should not open the store with the other and leave somebody believing both are reachable. + /// + public static IReadOnlyList? NormalizeNetworkRoles(string? role) { /* Literals rather than DarlingManagedPostgres.ViewerRoleName/AdminRoleName: that type is [SupportedOSPlatform("windows")] and this classifier is platform-neutral, so referencing its consts would raise CA1416 here. The names mirror those consts (pinned equal by test). */ if (string.IsNullOrWhiteSpace(role)) { - return "viewer"; + return new[] { "viewer" }; + } + + var parts = role.Split( + new[] { ',', '+', ' ', '\t', '\r', '\n' }, + StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries); + + /* Separators and nothing else (",", "+") is NOT the blank case: the operator wrote a value and none + of it names a role, so it takes the same fail-closed path as a typo rather than being read as the + viewer default. Guessing here would expose a store off the back of a malformed field. */ + if (parts.Length == 0) + { + return null; + } + + var seen = new List(); + + foreach (var part in parts) + { + var normalized = part.ToLowerInvariant(); + + if (normalized is not ("viewer" or "admin")) + { + return null; + } + + if (!seen.Contains(normalized, StringComparer.Ordinal)) + { + seen.Add(normalized); + } } - var normalized = role.Trim().ToLowerInvariant(); - return normalized is "viewer" or "admin" ? normalized : null; + /* Stable order regardless of how it was written, so the generated pg_hba block is identical for + "admin,viewer" and "viewer,admin" — otherwise the reconciler rewrites the file and reloads the + server every time somebody reorders the field. */ + seen.Sort(StringComparer.Ordinal); + return seen; } } diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingFirewallCheck.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingFirewallCheck.cs index 3ad5bf8a3..4fb0f5672 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingFirewallCheck.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingFirewallCheck.cs @@ -115,6 +115,13 @@ internal static string SurfaceRuleWildcard(string ruleName) /// stores one per profile) and a single match would otherwise have no .Count. /// /// The name goes through the same single-quoted-literal escaping as the write builders. + /// Takes either ONE rule's exact DisplayName — what asks, because the service + /// verifies the specific rule its own config wants — or a , which is what + /// --configure-firewall's not-elevated branch asks (#2445): there the question is whether the sweep + /// it is about to hand over would remove anything, so the probe has to cover exactly what the sweep would + /// delete. Single quoting is literal to the SHELL and not to the cmdlet, so a * reaches + /// -DisplayName as the wildcard it is — the same way it does for the sweep's + /// Remove-NetFirewallRule -DisplayName, which is where that behaviour is already load-bearing. /// internal static string BuildProbeCommand(string ruleName) => $"$c = @(Get-NetFirewallRule -DisplayName {DarlingManagedPostgres.SingleQuotedPowerShell(ruleName)} " + diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingInstallDirectoryReport.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingInstallDirectoryReport.cs index ba94a168f..e5b346488 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingInstallDirectoryReport.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingInstallDirectoryReport.cs @@ -35,6 +35,26 @@ namespace PerformanceMonitor.Darling.Service; /// test, so a directory it cannot account for may be anything at all — an operator's snapshot, a deploy /// tool's staging folder, someone's notes. Reporting is the only verdict that stays harmless when the /// classification is wrong, and here it will sometimes be wrong by construction. +/// +/// Rollback backups are recognised, and that is not the same as being owned (#2525). A +/// dogfood box was found carrying forty-six _rollback_manual_* directories holding 5.48 GB, the +/// oldest three weeks old — and this report named every one of them on its own line, every start. Each +/// line was individually true and the pile of them was useless: a real layout problem, a stray DLL or a +/// half-extracted upgrade, arrives as warning forty-seven in a list of forty-six identical ones. That is a +/// guard that has stopped guarding by being too loud, which is the same failure as a pin that never bites +/// wearing the opposite clothes. So a directory in the deploy procedure's own namespace +/// () now gets ONE line for the whole set — count, total size, the age +/// of the oldest, and the command that prunes them — leaving the per-directory lines for the directories +/// that genuinely have something to diagnose. Nothing is hidden and nothing is deleted: the set is still +/// reported every start, and the size still accumulates in a total an operator can act on. +/// +/// Silent in the healthy state, loud exactly once when there is a backlog. Under the deploy +/// script's retention the steady state is a handful of backups by design, and warning about the correct +/// state on every start is how this report would talk itself back into being ignored. At or under +/// the line drops to informational; past it, it is a +/// warning that names the excess. The recognition half matters even where retention is already running, +/// because the boxes carrying the backlog today got it before any script pruned anything, and no upgrade +/// removes a directory the product did not create. /// [SupportedOSPlatform("windows")] internal static class DarlingInstallDirectoryReport @@ -109,55 +129,41 @@ internal static void Report(string installDirectory, ILogger logger, TimeSpan si } var deadline = DateTime.UtcNow + sizeProbeBudget; - var found = new List<(string Path, long Bytes, bool Measured)>(); + var unaccountedFor = new List(); + var rollbackBackups = new List(); foreach (var candidate in Directory.GetDirectories(root)) { - if (IsProductDirectory(candidate) || IsSatelliteResourceDirectory(candidate)) + if (IsProductDirectory(candidate)) { continue; } - var bytes = DarlingStoreUpgrade.MeasureDirectoryBytes(candidate, deadline, out var measured); - found.Add((candidate, bytes, measured)); - } - - if (found.Count == 0) - { - return; - } - - long total = 0; - var allMeasured = true; - foreach (var (_, bytes, measured) in found) - { - total += bytes; - allMeasured &= measured; - } - - found.Sort(static (left, right) => right.Bytes.CompareTo(left.Bytes)); - - logger.LogWarning( - "{Count} director(ies) in the install directory {Root} are not part of the product's layout, and are holding {Approximately}{Size}. NONE of them is deleted automatically — the product does not know what they are, only that it did not put them there. Remove the ones you no longer need: a major store upgrade in copy mode needs roughly twice the data directory in free space.", - found.Count, root, allMeasured ? string.Empty : "at least ", DarlingStoreUpgrade.FormatBytes(total)); + /* Ahead of the satellite test on purpose, and the order is worth a line. The satellite test + is a directory WALK; the rollback test is a string comparison. On the box that produced + #2525 that ordering is the difference between classifying forty-six backups for nothing + and walking 5.48 GB of binaries to learn what their name already said. */ + if (DarlingRollbackBackups.IsRollbackBackup(candidate)) + { + rollbackBackups.Add(candidate); + continue; + } - foreach (var (path, bytes, measured) in found) - { - /* An exhausted budget degrades a directory's SIZE and never its presence in this list. The - whole point of the report is that the operator learns these exist; letting a slow walk - silence one would be the report failing at the only job it has. */ - if (!measured) + if (IsSatelliteResourceDirectory(candidate)) { - logger.LogWarning( - "Directory not part of the product's layout: {Path} (at least {Size}; the {Budget}-second size probe did not finish walking it). It is never deleted automatically.", - path, DarlingStoreUpgrade.FormatBytes(bytes), (int)sizeProbeBudget.TotalSeconds); continue; } - logger.LogWarning( - "Directory not part of the product's layout: {Path} ({Size}). It is never deleted automatically.", - path, DarlingStoreUpgrade.FormatBytes(bytes)); + unaccountedFor.Add(candidate); } + + /* Classify everything first, then measure the unaccounted-for directories BEFORE the backups. + The budget is finite and shared, and the enumeration order of a directory is the filesystem's + business — so measuring in discovery order lets forty-six backups spend the whole budget and + leave the one stray directory, the only line here anybody has to act on, reporting "at least + 0 bytes". The report that carries the signal gets first call on the budget. */ + ReportUnaccountedFor(root, Measure(unaccountedFor, deadline), logger, sizeProbeBudget); + ReportRollbackBackups(root, Measure(rollbackBackups, deadline), logger); } catch (Exception ex) { @@ -167,6 +173,161 @@ silence one would be the report failing at the only job it has. */ } } + /// + /// Sizes each directory against the shared , carrying whether the walk + /// finished so a caller can say "at least" rather than state a number it did not finish computing. + /// + private static List<(string Path, long Bytes, bool Measured)> Measure(List directories, DateTime deadline) + { + var measured = new List<(string Path, long Bytes, bool Measured)>(directories.Count); + foreach (var directory in directories) + { + var bytes = DarlingStoreUpgrade.MeasureDirectoryBytes(directory, deadline, out var complete); + measured.Add((directory, bytes, complete)); + } + + return measured; + } + + /// + /// The original report, unchanged in shape: one summary line and one line per directory, biggest first, + /// naming every directory the product's layout cannot account for. + /// + /// Per-directory lines are right HERE and only here. The product does not know what any of these + /// are, so the path is the whole message — there is nothing to summarise and nobody can act on a count. + /// The rollback backups moved out from under this exactly so these lines stay findable. + /// + private static void ReportUnaccountedFor( + string root, + List<(string Path, long Bytes, bool Measured)> found, + ILogger logger, + TimeSpan sizeProbeBudget) + { + if (found.Count == 0) + { + return; + } + + long total = 0; + var allMeasured = true; + foreach (var (_, bytes, measured) in found) + { + total += bytes; + allMeasured &= measured; + } + + found.Sort(static (left, right) => right.Bytes.CompareTo(left.Bytes)); + + logger.LogWarning( + "{Count} director(ies) in the install directory {Root} are not part of the product's layout, and are holding {Approximately}{Size}. NONE of them is deleted automatically — the product does not know what they are, only that it did not put them there. Remove the ones you no longer need: a major store upgrade in copy mode needs roughly twice the data directory in free space.", + found.Count, root, allMeasured ? string.Empty : "at least ", DarlingStoreUpgrade.FormatBytes(total)); + + foreach (var (path, bytes, measured) in found) + { + /* An exhausted budget degrades a directory's SIZE and never its presence in this list. The + whole point of the report is that the operator learns these exist; letting a slow walk + silence one would be the report failing at the only job it has. */ + if (!measured) + { + logger.LogWarning( + "Directory not part of the product's layout: {Path} (at least {Size}; the {Budget}-second size probe did not finish walking it). It is never deleted automatically.", + path, DarlingStoreUpgrade.FormatBytes(bytes), (int)sizeProbeBudget.TotalSeconds); + continue; + } + + logger.LogWarning( + "Directory not part of the product's layout: {Path} ({Size}). It is never deleted automatically.", + path, DarlingStoreUpgrade.FormatBytes(bytes)); + } + } + + /// + /// The whole set of deploy rollback backups, in ONE line: how many, what they cost together, which is + /// the oldest, and the command that prunes them (#2525). + /// + /// No per-directory lines, deliberately. Their name says what they are and our own deploy script + /// put every one of them there, so a line each adds a path and no information — while costing the + /// per-directory lines above the only property that makes them worth reading, which is that a line in + /// this report means something happened that nobody planned. + /// + /// Severity is the retention decision, not the disk cost. At or under the deploy script's + /// retention this is the intended state of a box that has been upgraded a few times, and a warning + /// about the intended state is noise with extra steps; past it the box is carrying backups that can + /// only roll back to a version nobody wants, which is worth one warning a start until someone prunes + /// them. + /// + /// The total still says "at least" when the budget ran out. A size that stopped being measured + /// must never be reported as though it were measured — the same rule the unaccounted-for report follows, + /// and it matters more here because this number is the one an operator sizes a cleanup against. + /// + private static void ReportRollbackBackups( + string root, + List<(string Path, long Bytes, bool Measured)> backups, + ILogger logger) + { + if (backups.Count == 0) + { + return; + } + + long total = 0; + var allMeasured = true; + foreach (var (_, bytes, measured) in backups) + { + total += bytes; + allMeasured &= measured; + } + + var approximately = allMeasured ? string.Empty : "at least "; + + if (backups.Count <= DarlingRollbackBackups.DefaultRetained) + { + logger.LogInformation( + "{Count} deploy rollback backup(s) in the install directory {Root}, holding {Approximately}{Size} — within the {Keep} the deploy procedure keeps. It creates and prunes them; the service never touches one.", + backups.Count, root, approximately, DarlingStoreUpgrade.FormatBytes(total), DarlingRollbackBackups.DefaultRetained); + return; + } + + logger.LogWarning( + "{Count} deploy rollback backups in the install directory {Root}, holding {Approximately}{Size} — {Excess} more than the deploy procedure keeps, and the oldest of them is {Oldest}. Reported once with a running total rather than one warning each: our own deploy script made every one of these, so there is nothing in the list to diagnose, and a line each would bury the directories above that DO need looking at. Prune them with {Prune}, which keeps the newest {Keep}. The service never creates or deletes one.", + backups.Count, + root, + approximately, + DarlingStoreUpgrade.FormatBytes(total), + backups.Count - DarlingRollbackBackups.DefaultRetained, + OldestBackupName(backups), + DarlingRollbackBackups.PruneCommand, + DarlingRollbackBackups.DefaultRetained); + } + + /// + /// The NAME of the oldest backup, chosen by last-write time rather than by sorting the stamps in the + /// names. + /// + /// The filesystem knows when a directory was written; the name only carries whatever spelling the + /// procedure used that month, and #2525's box holds more than one spelling. Ordering on the timestamp + /// and printing the name gives an operator both — a stamp they can recognise, chosen by something that + /// cannot be wrong about which came first — and it is the same ordering upgrade-darling.ps1 + /// prunes by, so the directory this line calls oldest is the one the script would remove first. + /// + private static string OldestBackupName(List<(string Path, long Bytes, bool Measured)> backups) + { + var oldestPath = backups[0].Path; + var oldestWhen = Directory.GetLastWriteTimeUtc(oldestPath); + + for (var i = 1; i < backups.Count; i++) + { + var when = Directory.GetLastWriteTimeUtc(backups[i].Path); + if (when < oldestWhen) + { + oldestWhen = when; + oldestPath = backups[i].Path; + } + } + + return Path.GetFileName(oldestPath); + } + private static bool IsProductDirectory(string candidate) { var name = Path.GetFileName(candidate); diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingManagedPostgres.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingManagedPostgres.cs index 61dc23017..2c29af4ad 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingManagedPostgres.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingManagedPostgres.cs @@ -8,6 +8,7 @@ using System; using System.Collections.Generic; +using System.Linq; using System.Diagnostics; using System.Globalization; using System.IO; @@ -1699,7 +1700,7 @@ private sealed record NetworkPlan( NetworkMode Mode, string? ListenIp, string? Cidr, - string? Role, + IReadOnlyList? Roles, string? CertPath, string? KeyPath, string? DegradeReason) @@ -1707,8 +1708,8 @@ private sealed record NetworkPlan( public static NetworkPlan Loopback(string? degradeReason = null) => new(NetworkMode.Loopback, null, null, null, null, null, degradeReason); - public static NetworkPlan Exposed(string listenIp, string cidr, string role, string certPath, string keyPath) - => new(NetworkMode.Exposed, listenIp, cidr, role, certPath, keyPath, null); + public static NetworkPlan Exposed(string listenIp, string cidr, IReadOnlyList roles, string certPath, string keyPath) + => new(NetworkMode.Exposed, listenIp, cidr, roles, certPath, keyPath, null); } /// @@ -1718,7 +1719,7 @@ public static NetworkPlan Exposed(string listenIp, string cidr, string role, str /// ListenIp/Cidr/Role. /// internal sealed record NetworkExposureDecision( - bool Exposed, string? ListenIp, string? Cidr, string? Role, string? DegradeReason); + bool Exposed, string? ListenIp, string? Cidr, IReadOnlyList? Roles, string? DegradeReason); /// /// PURE validation of postgres.network into a — NO I/O and NO cert @@ -1755,11 +1756,15 @@ internal static NetworkExposureDecision ResolveNetworkExposure(PostgresNetworkCo $"postgres.network.allowFrom '{network.AllowFrom}' address family does not match listen '{listenRaw}'"); } - var role = DarlingNetwork.NormalizeNetworkRole(network.Role); - if (role is null) + /* #2665: one rule per role, so an admin Viewer and read-only ones can reach the same store. Null + when ANY element is unrecognised rather than keeping the ones it understood — a typo in one of two + roles must not open the store with the other and leave somebody believing both are reachable. */ + var roles = DarlingNetwork.NormalizeNetworkRoles(network.Role); + if (roles is null || roles.Count == 0) { return Degrade( - $"postgres.network.role '{network.Role}' must be 'viewer' or 'admin' (never the superuser)"); + $"postgres.network.role '{network.Role}' must be 'viewer', 'admin', or both " + + "(e.g. \"admin,viewer\") — never the superuser"); } /* The cert path rides the -o string, which cannot nest-quote a space -> a spaced path fail-DEADS @@ -1772,7 +1777,7 @@ the postmaster. Gate on it (D6) and degrade rather than ever start ssl=on with a } /* Canonical base/prefix form (IPNetwork requires zeroed host bits) for the pg_hba line + firewall. */ - return new NetworkExposureDecision(true, listenIp.ToString(), $"{cidr.BaseAddress}/{cidr.PrefixLength}", role, null); + return new NetworkExposureDecision(true, listenIp.ToString(), $"{cidr.BaseAddress}/{cidr.PrefixLength}", roles, null); static NetworkExposureDecision Degrade(string reason) => new(false, null, null, null, reason); } @@ -1804,7 +1809,7 @@ private NetworkPlan BuildNetworkPlan() return NetworkPlan.Loopback($"could not generate/read the store TLS cert ({ex.Message})"); } - return NetworkPlan.Exposed(decision.ListenIp!, decision.Cidr!, decision.Role!, certPath, keyPath); + return NetworkPlan.Exposed(decision.ListenIp!, decision.Cidr!, decision.Roles!, certPath, keyPath); } private static bool ContainsWhitespace(string value) @@ -2000,6 +2005,19 @@ internal static string BuildListenAddresses(string? networkListenIp) internal static string BuildNetworkPgHbaLine(string role, string cidr) => $"hostssl {DatabaseName} {role} {cidr} scram-sha-256"; + /// + /// The managed pg_hba block for every admitted role (#2665) — one hostssl line each, in the + /// order settled on, so the text is identical however + /// the field was written and the reconciler does not rewrite-and-reload on a reordering. + /// + /// Each line still names exactly one role and one CIDR: never all, never the superuser + /// (D5/D6). A list admits more roles, it does not widen what any one line grants — and because these + /// live INSIDE the managed block, tightening allowFrom narrows all of them, which is exactly what + /// a hand-added second line outside the markers would not do. + /// + internal static string BuildNetworkPgHbaLines(IReadOnlyList roles, string cidr) + => string.Join("\n", roles.Select(role => BuildNetworkPgHbaLine(role, cidr))); + /// /// Whether the live pg_hba.conf needs reconciling (darling-network-endpoints): true when there is a rule /// to apply ( non-null = exposing) OR the file still carries a Darling @@ -2015,9 +2033,12 @@ public static bool NeedsPgHbaReconcile(string current, string? desiredLine) /// /// Pure marked-block reconcile for pg_hba.conf (D5): returns with the /// Darling-managed block (between and ) set - /// to a single when non-empty, or REMOVED when null/empty (disable). + /// to when non-empty, or REMOVED when null/empty (disable). /// Every non-marked line is preserved verbatim; a narrowing CIDR REPLACES the block (old removed, new /// appended); idempotent (re-running with the same desired yields identical text). + /// Since #2665 the desired text may be SEVERAL newline-separated rules (one per admitted role, + /// from ). That needs no code change here — this always replaced the + /// whole block rather than a line — but it is why the parameter is a block, not a rule. /// public static string ReconcilePgHba(string existing, string? desiredRuleLine) { @@ -2087,7 +2108,7 @@ private async Task ReconcileNetworkAsync( try { var exposed = plan.Mode == NetworkMode.Exposed; - var desiredLine = exposed ? BuildNetworkPgHbaLine(plan.Role!, plan.Cidr!) : null; + var desiredLine = exposed ? BuildNetworkPgHbaLines(plan.Roles!, plan.Cidr!) : null; var hbaPath = Path.Combine(_dataDirectory, "pg_hba.conf"); if (!File.Exists(hbaPath)) @@ -2165,9 +2186,9 @@ private async Task ReconcileNetworkAsync( /// /// Confirms the live pg_hba rules match intent via pg_hba_file_rules: no rule has a parse error, - /// and the Darling hostssl rule for the network role is PRESENT when exposed / ABSENT when - /// loopback. A mismatch is logged critical (a reload delivers a SIGHUP that Postgres may then reject). - /// Best-effort — a query failure degrades to a warning, not a throw. + /// and a Darling hostssl rule exists for EVERY admitted network role when exposed / for none of them + /// when loopback (#2665). A mismatch is logged critical (a reload delivers a SIGHUP that Postgres may + /// then reject). Best-effort — a query failure degrades to a warning, not a throw. /// private async Task VerifyPgHbaAsync(string ownerConnectionString, NetworkPlan plan, bool reloaded, CancellationToken cancellationToken) { @@ -2193,24 +2214,35 @@ private async Task VerifyPgHbaAsync(string ownerConnectionString, NetworkPlan pl long present; if (exposed) { + /* EVERY admitted role must be live, not just one (#2665): a rule that failed to apply for the + second role leaves those clients locked out while the exposure looks healthy. + + Counting DISTINCT ROLE NAMES rather than matching ROWS, because rows do not answer the + question. `count(*) ... WHERE user_name && $2` is satisfied by two rules naming the SAME + role, and that is the expected shape on an upgraded box: #2665's own workaround was a + hand-added second hostssl line outside the markers, which ReconcilePgHba deliberately + preserves. Such a file with a stale viewer line and no live admin rule counts 2 of 2 and + passes, which is precisely the failure this check exists to catch. Unnesting and counting + distinct names is >= Roles.Count only when every admitted role really has a rule. */ await using var command = new NpgsqlCommand( - "SELECT count(*) FROM pg_hba_file_rules WHERE type = 'hostssl' AND $1 = ANY(database) AND $2 = ANY(user_name)", + "SELECT count(DISTINCT u) FROM pg_hba_file_rules AS r, unnest(r.user_name) AS u " + + "WHERE r.type = 'hostssl' AND $1 = ANY(r.database) AND u = ANY($2::text[])", connection); command.Parameters.AddWithValue(DatabaseName); - command.Parameters.AddWithValue(plan.Role!); + command.Parameters.AddWithValue(plan.Roles!.ToArray()); present = Convert.ToInt64(await command.ExecuteScalarAsync(cancellationToken) ?? 0L); - if (present == 0) + if (present < plan.Roles!.Count) { _logger.LogCritical( - "pg_hba verification: the expected 'hostssl {Db} {Role} {Cidr} scram-sha-256' rule is NOT live after reload — the store is not accepting network clients as intended", - DatabaseName, plan.Role, plan.Cidr); + "pg_hba verification: only {Present} of the {Expected} configured role(s) '{Roles}' have a live 'hostssl {Db} {Cidr} scram-sha-256' rule after reload — the store is not accepting network clients as intended", + present, plan.Roles!.Count, string.Join(", ", plan.Roles!), DatabaseName, plan.Cidr); return; } _logger.LogInformation( - "Store network access reconciled: exposed to {Cidr} as '{Role}' over TLS (pg_hba {Verb}).", - plan.Cidr, plan.Role, reloaded ? "reloaded + verified" : "already current"); + "Store network access reconciled: exposed to {Cidr} as '{Roles}' over TLS (pg_hba {Verb}).", + plan.Cidr, string.Join(", ", plan.Roles!), reloaded ? "reloaded + verified" : "already current"); } else { diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingNetworkConfigEditor.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingNetworkConfigEditor.cs index 05d69120d..29e3a1360 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingNetworkConfigEditor.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingNetworkConfigEditor.cs @@ -423,7 +423,7 @@ internal static string BuildStoreNetworkBlock(string listen, string allowFrom, s "\"network\": {\n" + FieldIndent + $"\"listen\": {JsonString(listen)}, // bind IP; 0.0.0.0 = all interfaces (connect by a cert SAN name).\n" + FieldIndent + $"\"allowFrom\": {JsonString(allowFrom)}, // pg_hba + firewall CIDR (address family must match listen).\n" + - FieldIndent + $"\"role\": {JsonString(role)} // remote pg_hba role: viewer (read-only, default) or admin (remote writes).\n" + + FieldIndent + $"\"role\": {JsonString(role)} // remote pg_hba role(s): viewer (read-only, default), admin (remote writes), or both (\"admin,viewer\").\n" + ChildIndent + "}"; /// diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingObservability.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingObservability.cs index adf49669a..6005c4120 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingObservability.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingObservability.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -30,17 +30,24 @@ public static class DarlingObservability only enabled servers are in the loop). The ON CONFLICT re-connect deliberately does NOT touch is_enabled, so a control-plane disable (config_monitored_servers.is_enabled = FALSE, mirrored onto this observed row by SyncServerEnabledStatesAsync) is never clobbered back to TRUE on the next - connect. Before Stage 2 this forced is_enabled = TRUE on every connect and nothing ever read it. */ - private const string UpsertServerSql = @" -INSERT INTO servers (server_id, server_name, display_name, is_enabled, sql_engine_edition, sql_major_version, created_date, modified_date, monthly_cost_usd) -VALUES ($1, $2, $3, TRUE, $4, $5, $6, $6, $7) + connect. Before Stage 2 this forced is_enabled = TRUE on every connect and nothing ever read it. + + engine_kind (V82, #2530) is written on BOTH arms, unlike is_enabled: it is a probed fact about the + target rather than an operator decision, so the re-connect arm is exactly where a re-pointed + registration (same storage name, different engine) has to correct it. Internal so a pure test can pin + the shape without a live store. */ + internal const string UpsertServerSql = @" +INSERT INTO servers (server_id, server_name, display_name, is_enabled, sql_engine_edition, sql_major_version, engine_kind, created_date, modified_date, monthly_cost_usd, postgres_major_version) +VALUES ($1, $2, $3, TRUE, $4, $5, $6, $7, $7, $8, $9) ON CONFLICT (server_id) DO UPDATE SET server_name = EXCLUDED.server_name, display_name = EXCLUDED.display_name, sql_engine_edition = EXCLUDED.sql_engine_edition, sql_major_version = EXCLUDED.sql_major_version, + engine_kind = EXCLUDED.engine_kind, modified_date = EXCLUDED.modified_date, - monthly_cost_usd = EXCLUDED.monthly_cost_usd;"; + monthly_cost_usd = EXCLUDED.monthly_cost_usd, + postgres_major_version = EXCLUDED.postgres_major_version;"; /* Mirror the DESIRED config (config.config_monitored_servers) onto the OBSERVED registry (collect.servers) for the two fields the viewer/FinOps read straight off collect.servers: is_enabled @@ -77,8 +84,8 @@ AND NOT EXISTS (SELECT 1 FROM config.config_monitored_servers c WHERE c.server_id = s.server_id);"; private const string InsertCollectionLogSql = @" -INSERT INTO collection_log (log_id, server_id, server_name, collector_name, collection_time, duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) -VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11);"; +INSERT INTO collection_log (log_id, server_id, server_name, collector_name, collection_time, duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms, fanout_item_count, slowest_item, slowest_item_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13, $14);"; /* The fleet-sentinel server_id the daily retention purge writes its run-record under. collection_log's server_id is NOT NULL and Collection Health reads per real server_id, but the purge is fleet-wide (per @@ -128,10 +135,23 @@ public static async Task UpsertServerAsync(NpgsqlDataSource postgres, ServerRunt just the 5/8 Azure classifications. */ command.Parameters.AddWithValue(server.EngineEdition); command.Parameters.AddWithValue(server.Target.SqlMajorVersion); + /* The engine KIND (#2530), derived from the target the connector probed rather than from the + configured engine string: Aurora-ness is not configurable — it comes from aurora_version being + present in pg_proc — and it is half of what this column exists to carry. */ + command.Parameters.AddWithValue(MonitoredEngineKind.For(server.Target)); /* Naive-UTC storage: Npgsql 6+ rejects Kind=Utc against `timestamp` — see PgCollectorRowWriter. */ command.Parameters.AddWithValue(DateTime.SpecifyKind(DateTime.UtcNow, DateTimeKind.Unspecified)); /* Per-server FinOps budget from darling.json (0 = hide cost in the viewer, like Lite). */ command.Parameters.AddWithValue(server.Config.MonthlyCostUsd); + /* The probed PostgreSQL major (V100, #2653), so a READ can explain a column the target's version + does not have. DBNull rather than 0 off a PostgreSQL target: the reads treat NULL as "no claim", + and a 0 would be a version nobody runs asserted as fact. Guarded on Engine as well as on the + value because 0 is also what a PostgreSQL probe that failed before reading server_version_num + leaves behind. */ + command.Parameters.AddWithValue( + server.Target.Engine == CollectorTargetEngine.PostgreSql && server.Target.PostgresMajorVersion > 0 + ? server.Target.PostgresMajorVersion + : (object)DBNull.Value); await command.ExecuteNonQueryAsync(cancellationToken); } catch (Exception ex) @@ -188,6 +208,13 @@ public static async Task SyncServerEnabledStatesAsync(NpgsqlDataSource postgres, /// the storage phase; duckdb_duration_ms carries the storage (Postgres) milliseconds under /// Lite's column name so analysis SQL can twin. /// + /// + /// The per-database rollup for a run that fanned out, null for one that did not (#2472). REQUIRED rather + /// than defaulted on purpose: five collectors fan out and the rest do not, so every call site has to say + /// which it is. A default would have let the five failure sites below — which genuinely have no fan-out — + /// stand in for a success site that forgot, and the whole point of the columns is that a blended number + /// stops being the only thing recorded. + /// public static async Task LogCollectionAsync( NpgsqlDataSource postgres, ServerRuntime server, @@ -197,6 +224,7 @@ public static async Task LogCollectionAsync( long sqlMs, long storageMs, string? errorMessage, + FanoutCost? fanout, ILogger? logger, CancellationToken cancellationToken) { @@ -222,6 +250,12 @@ public static async Task LogCollectionAsync( command.Parameters.AddWithValue(rowsCollected); command.Parameters.AddWithValue((int)sqlMs); command.Parameters.AddWithValue((int)storageMs); + + /* All three NULL together or all three set: a slowest item with no count cannot be turned into + the dominance ratio the columns exist for, so half an answer is worse than none. */ + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Integer, Value = fanout.HasValue ? fanout.Value.ItemCount : (object)DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = fanout.HasValue ? fanout.Value.SlowestItem : (object)DBNull.Value }); + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Integer, Value = fanout.HasValue ? fanout.Value.SlowestItemMs : (object)DBNull.Value }); await command.ExecuteNonQueryAsync(cancellationToken); } catch (Exception ex) @@ -276,6 +310,14 @@ public static async Task LogRetentionRunAsync( command.Parameters.AddWithValue(rowsPurged); // rows_collected command.Parameters.AddWithValue(0); // sql_duration_ms (no SQL-target phase) command.Parameters.AddWithValue(elapsed); // duckdb_duration_ms (storage phase = whole sweep) + + /* The fan-out rollup columns (#2472). The retention sweep is fleet-wide over shared tables, not + a per-database fan-out, so it has nothing to attribute and says so with NULL rather than a + zero that would read as "fanned out over nothing". These three exist because this INSERT is + shared with the collector writer above; the shape is one statement on purpose. */ + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Integer, Value = DBNull.Value }); // fanout_item_count + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Text, Value = DBNull.Value }); // slowest_item + command.Parameters.Add(new NpgsqlParameter { NpgsqlDbType = NpgsqlDbType.Integer, Value = DBNull.Value }); // slowest_item_ms await command.ExecuteNonQueryAsync(cancellationToken); } catch (Exception ex) diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs index fa6fcf25f..d67ea7d44 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingRetention.cs @@ -465,9 +465,24 @@ statement that never had text rather than as one whose text expired. */ is ~40 MB against the plan dim's 127 GB — shortening it buys nothing and would quietly break "text stays analyzable for the facts' full retention", which is half the knob's own justification. The router is pure so the scoping is pinned by tests. */ + /* #2386: the plan dim is capped by ROWS, every other dim keeps the one-day slice. + Only this table carries multi-kilobyte TOAST payloads, so only this one has a day + that cannot be deleted inside the command timeout — query_text_dim is ~40 MB + total and drains in a single slice. Scoped rather than global so a table that is + fine keeps the compressed-chunk-safe shape it needs. */ + var isPlanDim = string.Equals( + dimTable, PayloadDimensions.QueryPlanDimTable, StringComparison.Ordinal); + var dimDeleted = await PurgeOneAsync( - postgres, dimTable, TimeSlicedDeleteSql(dimTable, PayloadDimensions.LastSeenColumn), - ComputeDimTableCutoff(dimTable, dimensionCutoff, planDimensionCutoff), logger, cancellationToken); + postgres, + dimTable, + isPlanDim + ? RowCappedDeleteSql(dimTable, PayloadDimensions.LastSeenColumn, PlanDimDeleteRowCap) + : TimeSlicedDeleteSql(dimTable, PayloadDimensions.LastSeenColumn), + ComputeDimTableCutoff(dimTable, dimensionCutoff, planDimensionCutoff), + logger, + cancellationToken, + batchSize: isPlanDim ? PlanDimDeleteRowCap : 1); if (dimDeleted is not null) { tablesPurged++; @@ -663,6 +678,46 @@ internal static string TimeSlicedDeleteSql(string table, string timeColumn, stri + $" AND {timeColumn} < (SELECT min({timeColumn}) FROM {table} WHERE {expired}) + INTERVAL '{TimescaleSupport.ChunkIntervalDays} days'"; } + /// + /// Rows per statement for the plan dimension's purge (#2386). Measured on the use2 store, whose + /// query_plan_dim is 133 GB over 12.4 M rows: deletes run at ~1,000 rows/sec, linear + /// from 10 k to 50 k, worst observed 834 rows/sec. 50 k therefore costs ~50 s against + /// 's 300 s — about 5x margin, which is the point: a loaded box is + /// slower than the idle one this was measured on, and 100 k (~97 s nominal) would sit at ~291 s under + /// a 3x slowdown, i.e. against the wall. + /// + internal const int PlanDimDeleteRowCap = 50_000; + + /// + /// The plan dimension's purge statement: capped by ROW COUNT rather than by a time slice (#2386). + /// + /// Why the time slice cannot work here. bounds work at + /// one day, justified as "exactly what a steady-state daily purge deletes in total". True, and that is + /// the problem for this table: a day is ~755 k rows whose gzipped plan XML averages ~9.5 KB, so one + /// statement must remove ~7 GB of TOAST. Measured, that needs ~755 s against a 300 s command + /// timeout — 2.5x over, not borderline. It times out, the statement rolls back, nothing is deleted, + /// and the next sweep retries the identical doomed slice. Retention on the largest table in the store + /// stops permanently, and the table only grows. No choice of horizon avoids it: the slice width is set + /// by the data at the old end, not by where the cutoff sits, so stepping the horizon down one day at a + /// time meets the same full day at the first step. + /// + /// Why a row cap is available here specifically. The ctid row-cap idiom was + /// abandoned in #1564 because reading the ctid system column through TimescaleDB's transparent + /// decompression is unsupported, so it errored the moment any in-range chunk was compressed. + /// query_plan_dim is a PLAIN table, never a hypertable — that constraint has never applied to + /// it, and the fact tables that do need the compressed-safe shape keep + /// . + /// + /// ORDER BY {timeColumn} keeps the delete oldest-first, which the index on that column + /// serves directly (measured: the planner takes an index-only scan, and the sibling min() probe + /// costs 0.364 ms — the scan was never the expense). Oldest-first matters because progress has to be + /// monotonic: an unordered cap would nibble arbitrary rows and leave the floor where it was. + /// + internal static string RowCappedDeleteSql(string table, string timeColumn, int cap) => + $"DELETE FROM {table} WHERE ctid IN (" + + $"SELECT ctid FROM {table} WHERE {timeColumn} < $1 " + + $"ORDER BY {timeColumn} LIMIT {cap})"; + /// /// The dimension GC's cutoff (#1795): the ASSUMED horizon (widest dim-feeding fact retention + /// drop_chunks granularity + 1 day for the hourly @@ -882,8 +937,17 @@ avoidable work on a path that already runs per table per day. */ string deleteSql, DateTime cutoff, ILogger? logger, - CancellationToken cancellationToken) + CancellationToken cancellationToken, + int batchSize = 1) { + /* Accumulated OUTSIDE the try so the catch can report progress (#2386). Each statement + autocommits, so a timeout on the fifth batch does not undo the first four — but the old + catch returned null and threw the running total away, and the sweep's summary then said + "0 row(s) deleted, 1 failed" for a purge that had removed 755k rows. That reads as total + paralysis, which is what made this bug look worse than it was and hid that progress was + being made one slice per sweep. */ + var deleted = 0; + try { await using var connection = await postgres.OpenConnectionAsync(cancellationToken); @@ -902,14 +966,52 @@ avoidable work on a path that already runs per table per day. */ using var command = new NpgsqlCommand(deleteSql, connection) { CommandTimeout = DeleteTimeoutSeconds }; command.Parameters.AddWithValue(cutoff); - /* batchSize 1: the time-sliced statement has no row cap, so "fewer than the cap" degenerates - to "deleted zero rows" — a slice that clears anything means older slices may remain. */ - return await DrainBatchesAsync(ct => command.ExecuteNonQueryAsync(ct), batchSize: 1, cancellationToken); + /* batchSize 1 for the TIME-SLICED statement: it has no row cap, so "fewer than the cap" + degenerates to "deleted zero rows" — a slice that clears anything means older slices may + remain. A ROW-capped caller passes its cap instead, which restores the drain loop's real + contract (a full-cap batch means there may be more). */ + var batches = 0; + var drained = await DrainBatchesAsync( + async ct => + { + batches++; + var rows = await command.ExecuteNonQueryAsync(ct); + deleted += rows; + return rows; + }, + batchSize, + cancellationToken); + + /* A row-capped drain reports the two facts the sweep summary cannot carry, because both are + per-table and the summary is fleet-wide. + + The CUTOFF, because it is not the retention knob and reading it as the knob is a live trap: + ComputeDimensionCutoff subtracts the configured days PLUS a one-day margin for the hourly + last_seen refresh, so counting rows older than the knob value overstates what is eligible by + a full day of ingest — on this table that is hundreds of thousands of rows, which reads as a + backlog retention is failing to clear when it is simply not due yet. + + And the BATCH COUNT, because rows-deleted alone cannot distinguish a drain from a peel. That + is the #2386 failure mode exactly: a purge that removed one bounded slice and reported + success looked identical in the log to one that cleared everything expired. One batch means + the table was already inside its horizon; many means there was a backlog and it is gone. */ + if (batchSize > 1) + { + logger?.LogInformation( + "Retention purge drained {Rows} row(s) from {Table} in {Batches} batch(es) (cap {Cap}), cutoff {Cutoff:yyyy-MM-dd HH:mm}Z", + drained, tableName, batches, batchSize, cutoff); + } + + return drained; } catch (Exception ex) when (ex is not OperationCanceledException) { - /* Failure-isolated per table — one stuck DELETE must not stop the sweep. */ - logger?.LogWarning("Retention purge failed for {Table}: {Message}", tableName, ex.Message); + /* Failure-isolated per table — one stuck DELETE must not stop the sweep. Reports what DID + land, because those batches are committed and saying otherwise sends an operator looking + for a stall that is really a throughput limit. */ + logger?.LogWarning( + "Retention purge failed for {Table} after removing {Rows} row(s): {Message}", + tableName, deleted, ex.Message); return null; } } diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingRollbackBackups.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingRollbackBackups.cs new file mode 100644 index 000000000..996b25326 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingRollbackBackups.cs @@ -0,0 +1,95 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; + +namespace PerformanceMonitor.Darling.Service; + +/// +/// The single place the deploy procedure's rollback-backup naming convention is written down. Two very +/// different pieces of the product need to agree on it: upgrade-darling.ps1, which CREATES one of +/// these directories before it overwrites the install tree and PRUNES the ones past retention, and +/// , which has to RECOGNISE them so it can stop shouting about +/// them one line at a time. +/// +/// Why a shared constant and not two literals. #2525 is what happens when the two halves do +/// not agree: the deploy procedure had been writing _rollback_manual_<stamp> into the install +/// root since before the layout report existed, the layout report classifies by ELIMINATION against the +/// product's own directories, and so every backup the procedure made came back as a directory nobody could +/// account for. Forty-six of them, forty-six warnings, every start. If the script's spelling and the +/// service's matcher are ever allowed to drift apart the same failure returns silently, so the C# side owns +/// the string and DarlingDeployRollbackRetentionTests runs the script's own predicate against this +/// one to prove they still answer alike. +/// +/// What this does NOT cover. The store-copy report (#1775) sees directories named +/// <datadir>_rollback_manual_<stamp> beside the store's DATA directory under +/// %ProgramData% — the prefix as a suffix, on a different tree, holding a whole PostgreSQL cluster. +/// Those stay reported individually and deliberately: one of them can be hundreds of gigabytes, they are +/// identified structurally by PG_VERSION rather than by name, and no script of ours creates or prunes +/// them. This convention is about the INSTALL tree, where a backup is ~120 MB of binaries and the deploy +/// script both makes and removes them. +/// +internal static class DarlingRollbackBackups +{ + /// + /// What the deploy procedure names its pre-copy backup of the install tree. Matched + /// case-insensitively: Windows paths are, and a matcher stricter than the filesystem is a matcher that + /// misses a directory the operator will tell you is right there. + /// + internal const string Prefix = "_rollback_manual_"; + + /// + /// How many backups upgrade-darling.ps1 keeps, and the number the report quotes when it says a + /// box is carrying more than that. + /// + /// Three, not one and not ten. One is not a retention policy — it is gone the moment you deploy + /// the fix for the bad deploy, which is exactly when you want it. Three covers "the release, the one + /// before it, and the one before that", which is as far back as a rollback is ever actually wanted: + /// beyond it you are not rolling back, you are restoring a version nobody asked for. The field box in + /// #2525 had forty-six, and the forty-third could only have returned it to a build from three weeks + /// earlier. + /// + /// The script's -KeepRollbacks default is this number and + /// DarlingDeployRollbackRetentionTests pins the two together — a report that advertises a + /// retention the script does not implement is worse than no advice at all. + /// + internal const int DefaultRetained = 3; + + /// The command an operator runs to act on what the report just told them. + internal const string PruneCommand = "upgrade-darling.ps1 -PruneOnly"; + + /// + /// True when is one of the deploy procedure's rollback backups. + /// + /// Prefix plus at least one more character, and nothing more clever than that. A stamp format is + /// deliberately NOT parsed: the backlog this exists to recognise was made over months by a procedure + /// that has spelled its stamp more than one way, and a matcher that only accepted today's spelling + /// would leave the old ones being reported one line at a time — which is the entire complaint. The + /// prefix is a namespace: anything inside it belongs to the deploy procedure, and a directory an + /// operator named into someone else's namespace is treated as theirs. + /// + /// The trailing-character requirement keeps a bare _rollback_manual_ — a directory with no + /// stamp at all, which no procedure produces — outside the convention, so it stays in the + /// unaccounted-for report where a thing nobody can explain belongs. + /// + internal static bool IsRollbackBackup(string directoryPath) + { + if (string.IsNullOrEmpty(directoryPath)) + { + return false; + } + + /* GetFileName on a path with a trailing separator returns empty, so the separator comes off first — + the caller may be handing over a directory path spelled either way. */ + var name = Path.GetFileName(Path.TrimEndingDirectorySeparator(directoryPath)); + + return name.Length > Prefix.Length + && name.StartsWith(Prefix, StringComparison.OrdinalIgnoreCase); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs index d4809bb53..d8e86f94d 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingSelfAlertEvaluator.cs @@ -838,28 +838,48 @@ await FireAsync( if (!string.IsNullOrEmpty(replica.ConnectedStateDesc)) { _agReplicaConnectedState.TryGetValue(key, out var previousState); - var connection = AgAlertPolicy.DecideConnection(previousState, replica.ConnectedStateDesc); - _agReplicaConnectedState[key] = replica.ConnectedStateDesc; /* #1696 (V37): "AG Replica Disconnected" was a pure edge, so a replica that stayed disconnected for a week announced it ONCE. The #1659 treatment: re-announce every N minutes while it is still down (0 = off, the shipped default, so nothing starts re-alerting on upgrade). Re-fires deliver under the SAME metric name, because webhook - automation keyed on it is exactly what a re-fire exists to re-trigger. */ - bool stillDisconnected = - connection == AgConnectionDecision.None - && AgAlertPolicy.IsDisconnected(replica.ConnectedStateDesc) - && _agDisconnectRefireMinutes() is int refire - && refire > 0 - && (!_lastAgDisconnectAlert.TryGetValue(key, out var lastDown) - || _utcNow() - lastDown >= TimeSpan.FromMinutes(refire)); + automation keyed on it is exactly what a re-fire exists to re-trigger. + + #2426 moved the combined decision into the shared policy rather than leaving it as an + expression here: Lite grew the same knob, and the parts with sharp corners — a re-fire + must not double up with the edge that just fired, and an unrecognized state string must + not count as down — are precisely the parts that drift when written twice. What stays + here is what is genuinely this app's: the stamp, and the delivery it is stamped on. */ + int refireMinutes = _agDisconnectRefireMinutes(); + var connection = AgAlertPolicy.DecideConnection( + previousState, + replica.ConnectedStateDesc, + refireMinutes > 0 ? TimeSpan.FromMinutes(refireMinutes) : null, + _lastAgDisconnectAlert.TryGetValue(key, out var lastDown) ? lastDown : (DateTime?)null, + _utcNow()); + _agReplicaConnectedState[key] = replica.ConnectedStateDesc; + + bool stillDisconnected = connection == AgConnectionDecision.StillDisconnected; if (connection == AgConnectionDecision.Disconnected || stillDisconnected) { + /* #2426: the re-fire says so in the DETAIL, not only in the short message. ShortMessage is + the interactive toast body and reaches neither the history row nor the email — + DarlingAlertDeliverer forwards DetailText — so an operator opening Alert Detail on the + sixth re-announcement read text byte-identical to the first notice, which is precisely + the "a week-long outage reads like a blip" problem this knob exists to end. The + connection re-fire above already bakes it into detail; this is its AG twin, worded to + match Lite's so the two SKUs' history rows say the same thing. */ + var opening = stillDisconnected + ? $"Availability Group '{replica.AgName}': replica {replica.ReplicaServerName} is STILL " + + $"DISCONNECTED from the primary (re-alerting every {refireMinutes.ToString(CultureInfo.InvariantCulture)} min)." + : $"Availability Group '{replica.AgName}': replica {replica.ReplicaServerName} is DISCONNECTED " + + "from the primary."; + await FireAsync( Key(serverId), serverName, AgReplicaDisconnectedMetric, replica.ConnectedStateDesc, "CONNECTED", - detail: $"Availability Group '{replica.AgName}': replica {replica.ReplicaServerName} is " + - "DISCONNECTED from the primary. A disconnected replica receives no log at all, so it falls " + + detail: opening + + " A disconnected replica receives no log at all, so it falls " + "further behind every second and cannot be failed over to without losing whatever the primary " + "has committed since. If it is a synchronous-commit replica, the primary also loses its " + "automatic-failover partner. Check the replica's SQL Server service, the availability endpoint " + diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingServerConnector.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingServerConnector.cs index 9bbe764fa..1869b82a8 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingServerConnector.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingServerConnector.cs @@ -13,6 +13,7 @@ using Microsoft.Data.SqlClient; using Microsoft.Extensions.Logging; using Npgsql; +using PerformanceMonitor.Darling.Service.Targets; using PerformanceMonitor.Collectors; using PerformanceMonitor.Common; @@ -370,6 +371,15 @@ private static async Task ConnectPostgresAsync( PostgresMajorVersion = majorVersion, PostgresVersionNum = versionNum, IsAurora = isAurora, + /* #2633: derived from the ENDPOINT, because nothing else here can see it. IsAwsRds is + probed with a T-SQL detection query on the SQL Server path, so before this it was + silently false for every PostgreSQL target — which made the second half of + pg_plan_capture's `IsAurora || IsAwsRds` dispatch unreachable and sent plain RDS + PostgreSQL down the pg_read_file route, where a managed instance has no filesystem to + read and the failure names a grant that would never have helped. Aurora was unaffected + because IsAurora carries it, which is why the fleet never showed this. */ + IsAwsRds = RdsEndpoint.TryParse( + new NpgsqlConnectionStringBuilder(connectionString).Host) is not null, IsInRecovery = isInRecovery, }, StorageName = storageName, @@ -481,20 +491,13 @@ public static string DescribeProbeFacts(ConnectionProbeResult probe) $"{role}, {flavour} — {applies}"; } - /// Human-readable SERVERPROPERTY('EngineEdition') description for the probe result. - public static string DescribeEngineEdition(int engineEdition) => engineEdition switch - { - 1 => "Personal/Desktop", - 2 => "Standard", - 3 => "Enterprise", - 4 => "Express", - 5 => "Azure SQL Database", - 6 => "Azure Synapse Analytics", - 8 => "Azure SQL Managed Instance", - 9 => "Azure SQL Edge", - 11 => "Azure Synapse serverless SQL pool", - _ => $"Unknown ({engineEdition})", - }; + /// Human-readable SERVERPROPERTY('EngineEdition') description for the probe result. + /// Delegates to (#2511) rather than + /// keeping a second switch: the capability messages both MCP surfaces return name the edition, and two + /// edition tables in one repo drift — with the copy nobody is reading being the one that drifts. + /// + public static string DescribeEngineEdition(int engineEdition) => + CollectorEngineCapability.DescribeEngineEdition(engineEdition); } /// diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingWebEndpoints.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingWebEndpoints.cs index 96c4ca53f..245f62e89 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingWebEndpoints.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingWebEndpoints.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -32,7 +32,7 @@ namespace PerformanceMonitor.Darling.Service; /// each endpoint calls the SAME public static tool method the MCP server exposes (the [McpServerTool] /// attributes are inert — the tool bodies are plain statics), so there is ZERO SQL / projection drift between the /// two surfaces. Query-string values bind to the tool's parameters (server/server_name, -/// hours/hours_back, top, limit, ...); missing optional parameters fall back to the +/// hours/hours_back, as_of, top, limit, ...); missing optional parameters fall back to the /// method's own defaults. The excluded tools are the ones that are not read-only-over-the-store: /// analyze_server (a live monitored-server touch), mute_analysis_finding (a write), and the /// analyze_*_plan compute family (phase 2). A reflection test pins the endpoint set == the tool catalog @@ -1043,6 +1043,14 @@ internal sealed record CatalogRead(string Category, string Description, IReadOnl private static CatalogParam PServer() => new("server", TypeServer, false, null); private static CatalogParam PHours(int def) => new("hours", TypeInt, false, def); + + /// + /// The window ANCHOR (#2495): ?as_of= moves the END of a window off "now" so a + /// caller can ask about a past incident. Typed text rather than a new catalog type because it IS a + /// string on the wire — an ISO-8601 instant the tool parses and REFUSES if it cannot, which is where the + /// error message belongs; a new type would only teach the picker to render a box it already renders. + /// + private static CatalogParam PAsOf() => new("as_of", TypeText, false, null); private static CatalogParam PLimit(int def) => new("limit", TypeInt, false, def); private static CatalogParam PTop(int def) => new("top", TypeInt, false, def); private static CatalogParam PText(string name) => new(name, TypeText, false, null); @@ -1067,100 +1075,132 @@ private static CatalogRead R(string category, string description, params Catalog { /* ── analysis reads (AuditConfig/CompareAnalysis/GetAnalysisFacts/GetAnalysisFindings) ── */ ["audit_config"] = R(CatAnalysis, "Configuration-audit findings for a server.", PServer()), - ["compare_analysis"] = R(CatAnalysis, "Compare a window's analysis facts against an earlier baseline.", PServer(), PHours(4), PInt("baseline_hours_back", 28)), - ["get_analysis_facts"] = R(CatAnalysis, "Raw analysis facts for a window, filtered by source and minimum severity.", PServer(), PHours(4), PText("source"), PDouble("min_severity", 0)), - ["get_analysis_findings"] = R(CatAnalysis, "Persisted analysis findings for a server.", PServer(), PHours(24)), + ["compare_analysis"] = R(CatAnalysis, "Compare a window's analysis facts against an earlier baseline.", PServer(), PHours(4), PInt("baseline_hours_back", 28), PAsOf()), + ["get_analysis_facts"] = R(CatAnalysis, "Raw analysis facts for a window, filtered by source and minimum severity.", PServer(), PHours(4), PText("source"), PDouble("min_severity", 0), PAsOf()), + ["get_analysis_findings"] = R(CatAnalysis, "Persisted analysis findings for a server.", PServer(), PHours(24), PAsOf()), /* ── sessions (DarlingMcpSessionTools) ── */ - ["get_active_queries"] = R(CatSessions, "Currently-active queries, optionally blocking-only.", PServer(), PHours(1), PText("database_name"), PBool("blocking_only", false), PLimit(50)), + ["get_active_queries"] = R(CatSessions, "Currently-active queries, optionally blocking-only.", PServer(), PHours(1), PText("database_name"), PBool("blocking_only", false), PLimit(50), PAsOf()), ["get_session_stats"] = R(CatSessions, "Session-level summary counters for a server.", PServer()), - ["get_waiting_tasks"] = R(CatSessions, "Tasks currently waiting, with wait type and duration.", PServer(), PHours(1), PLimit(30)), + ["get_waiting_tasks"] = R(CatSessions, "Tasks currently waiting, with wait type and duration.", PServer(), PHours(1), PLimit(30), PAsOf()), /* ── alerts / mute rules (DarlingMcpAlertTools) ── */ - ["get_alert_history"] = R(CatAlerts, "Recent fired-alert history for a server.", PServer(), PHours(24), PLimit(50)), + ["get_alert_history"] = R(CatAlerts, "Recent fired-alert history for a server.", PServer(), PHours(24), PLimit(50), PAsOf()), ["get_alert_settings"] = R(CatAlerts, "The current alert-settings configuration."), ["get_mute_rules"] = R(CatAlerts, "The alert mute rules (enabled-only by default).", PBool("enabled_only", true)), /* ── blocking / deadlocks (DarlingMcpBlockingTools) ── */ - ["get_blocked_process_xml"] = R(CatBlocking, "Blocked-process-report XML captures.", PServer(), PHours(24), PLimit(5)), - ["get_blocking"] = R(CatBlocking, "Blocking chains observed in the window.", PServer(), PHours(24), PLimit(30)), - ["get_blocking_trend"] = R(CatBlocking, "Blocking-event counts over time.", PServer(), PHours(24)), - ["get_deadlock_detail"] = R(CatBlocking, "Deadlock graph detail for recent deadlocks.", PServer(), PHours(24), PLimit(5)), - ["get_deadlock_trend"] = R(CatBlocking, "Deadlock counts over time.", PServer(), PHours(24)), - ["get_deadlocks"] = R(CatBlocking, "Recent deadlocks with victim/resource summary.", PServer(), PHours(24), PLimit(20)), + ["get_blocked_process_xml"] = R(CatBlocking, "Blocked-process-report XML captures.", PServer(), PHours(24), PLimit(5), PAsOf()), + ["get_blocking"] = R(CatBlocking, "Blocking chains observed in the window.", PServer(), PHours(24), PLimit(30), PAsOf()), + ["get_blocking_trend"] = R(CatBlocking, "Blocking-event counts over time.", PServer(), PHours(24), PAsOf()), + ["get_deadlock_detail"] = R(CatBlocking, "Deadlock graph detail for recent deadlocks.", PServer(), PHours(24), PLimit(5), PAsOf()), + ["get_deadlock_trend"] = R(CatBlocking, "Deadlock counts over time.", PServer(), PHours(24), PAsOf()), + ["get_deadlocks"] = R(CatBlocking, "Recent deadlocks with victim/resource summary.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_lock_wait_trend"] = R(CatBlocking, "Every LCK% wait type's wait ms/sec over time - the aggregate lock-wait lane.", PServer(), PHours(24), PAsOf()), /* ── automatic plan correction (DarlingMcpPlanCorrectionTools, #2028) ── */ - ["get_plan_corrections"] = R(CatAnalysis, "Automatic plan correction activity + per-database FORCE_LAST_GOOD_PLAN state.", PServer(), PHours(24), PLimit(50)), + ["get_plan_corrections"] = R(CatAnalysis, "Automatic plan correction activity + per-database FORCE_LAST_GOOD_PLAN state.", PServer(), PHours(24), PLimit(50), PAsOf()), /* ── config: current + history (DarlingMcpConfigTools / DarlingMcpConfigHistoryTools) ── */ ["get_database_config"] = R(CatConfig, "Database-level configuration for a server.", PServer(), PText("database_name")), ["get_server_config"] = R(CatConfig, "Server-level configuration (sp_configure) for a server.", PServer()), ["get_trace_flags"] = R(CatConfig, "Active trace flags for a server.", PServer()), - ["get_database_config_changes"] = R(CatConfig, "Database-configuration changes over time.", PServer(), PHours(168)), + ["get_database_config_changes"] = R(CatConfig, "Database-configuration changes over time.", PServer(), PHours(168), PAsOf()), ["get_database_scoped_config"] = R(CatConfig, "Database-scoped configuration for a database.", PServer(), PText("database_name")), ["get_query_store_health"] = R(CatConfig, "Per-database Query Store health: actual vs desired state, readonly_reason, storage vs cap.", PServer(), PText("database_name")), - ["get_server_config_changes"] = R(CatConfig, "Server-configuration changes over time.", PServer(), PHours(168)), - ["get_trace_flag_changes"] = R(CatConfig, "Trace-flag changes over time.", PServer(), PHours(168)), + ["get_server_config_changes"] = R(CatConfig, "Server-configuration changes over time.", PServer(), PHours(168), PAsOf()), + ["get_trace_flag_changes"] = R(CatConfig, "Trace-flag changes over time.", PServer(), PHours(168), PAsOf()), /* ── core data reads (DarlingMcpDataTools + long-query / fleet tools) ── */ ["get_collection_health"] = R(CatData, "Per-collector collection health for a server.", PServer()), - ["get_cpu_utilization"] = R(CatData, "CPU utilization over time.", PServer(), PHours(4)), + ["get_collection_log"] = R(CatData, "Raw per-run collector log for a server, newest first.", PServer(), PHours(24), PLimit(200), PAsOf()), + ["get_current_waits_trend"] = R(CatData, "Waiting-task and blocked-session series over time.", PServer(), PHours(4), PText("database_name"), PAsOf()), + ["get_blocking_stats"] = R(CatData, "Blocking duration and deadlock severity per minute.", PServer(), PHours(24), PAsOf()), + ["get_cpu_utilization"] = R(CatData, "CPU utilization over time.", PServer(), PHours(4), PAsOf()), ["get_file_io_stats"] = R(CatData, "Per-file IO stall/throughput stats.", PServer()), ["get_memory_clerks"] = R(CatData, "Top memory clerks by allocation.", PServer()), ["get_memory_stats"] = R(CatData, "Server memory summary counters.", PServer()), ["get_perfmon_stats"] = R(CatData, "Perfmon counter values, filtered by counter/instance.", PServer(), PText("counter_name"), PText("instance_name")), - ["get_query_store_top"] = R(CatData, "Top Query Store queries in the window.", PServer(), PHours(24), PTop(20), PText("database_name")), - ["get_long_query_completions"] = R(CatData, "Completed long-running queries captured by the XE trace.", PServer(), PHours(24), PLimit(30)), + ["get_query_heatmap"] = R(CatData, "Query counts per (time bin x log-magnitude bucket) - the viewer's Query Heatmap as a table.", PServer(), PHours(24), PText("metric"), PText("database_name"), PInt("bucket_minutes", 5), PLimit(500), PAsOf()), + ["get_query_store_regressions"] = R(CatData, "Queries whose Query Store performance got WORSE vs their baseline.", PServer(), PHours(24), PText("database_name"), PLimit(50), PAsOf()), + ["get_query_store_top"] = R(CatData, "Top Query Store queries in the window.", PServer(), PHours(24), PTop(20), PText("database_name"), PAsOf()), + ["get_long_query_completions"] = R(CatData, "Completed long-running queries captured by the XE trace.", PServer(), PHours(24), PLimit(30), PAsOf()), ["get_server_properties"] = R(CatData, "Server properties/inventory for a server.", PServer()), - ["get_tempdb_trend"] = R(CatData, "tempdb space usage over time.", PServer(), PHours(24)), - ["get_top_procedures_by_cpu"] = R(CatData, "Top stored procedures by CPU.", PServer(), PHours(24), PTop(20), PText("database_name")), - ["get_top_queries_by_cpu"] = R(CatData, "Top queries by CPU, optionally parallel-only / min-DOP.", PServer(), PHours(24), PTop(20), PText("database_name"), PBool("parallel_only", false), PInt("min_dop", 0)), - ["get_pg_top_queries"] = R(CatData, "Top PostgreSQL query shapes by total execution time (Aurora targets).", PServer(), PHours(24), PLimit(20)), - ["get_pg_wraparound_risk"] = R(CatData, "PostgreSQL XID/MultiXact freeze headroom per database.", PServer(), PHours(24)), - ["get_pg_xmin_horizon"] = R(CatData, "What is holding back the PostgreSQL xmin horizon, by cause.", PServer(), PHours(24)), - ["get_pg_replication_slots"] = R(CatData, "PostgreSQL replication slot health, including whether retained WAL is still growing.", PServer(), PHours(24)), - ["get_pg_autovacuum_health"] = R(CatData, "PostgreSQL tables behind on vacuum or analyze, ranked by how far past each table's own threshold.", PServer(), PHours(24), PLimit(20)), - ["get_pg_io_stats"] = R(CatData, "PostgreSQL I/O by backend type, object and context, differenced across the window.", PServer(), PHours(24), PLimit(20)), - ["get_pg_wait_stats"] = R(CatData, "Top PostgreSQL wait events in the window (Aurora targets).", PServer(), PHours(24), PLimit(20)), - ["get_pg_blocking"] = R(CatData, "PostgreSQL blocking chains that were sampled, with the root blocker attributed. A sample, not an event log.", PServer(), PHours(24), PLimit(50)), - ["get_wait_stats"] = R(CatData, "Top wait statistics in the window.", PServer(), PHours(24), PLimit(20)), - ["get_wait_trend"] = R(CatData, "One wait type's totals over time (requires wait_type).", PReqText("wait_type"), PServer(), PHours(24)), - ["get_wait_types"] = R(CatData, "The wait types observed in the window.", PServer(), PHours(24)), + ["get_tempdb_trend"] = R(CatData, "tempdb space usage over time.", PServer(), PHours(24), PAsOf()), + ["get_top_procedures_by_cpu"] = R(CatData, "Top stored procedures by CPU.", PServer(), PHours(24), PTop(20), PText("database_name"), PAsOf()), + ["get_top_queries_by_cpu"] = R(CatData, "Top queries by CPU, optionally parallel-only / min-DOP.", PServer(), PHours(24), PTop(20), PText("database_name"), PBool("parallel_only", false), PInt("min_dop", 0), PAsOf()), + ["get_pg_top_queries"] = R(CatData, "Top PostgreSQL query shapes by total execution time (Aurora targets).", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_plans"] = R(CatData, "Captured PostgreSQL execution plans, grouped by shape. Plans are redacted at collection.", PServer(), PHours(24), PLimit(10), PText("query_id"), PAsOf()), + ["get_pg_wraparound_risk"] = R(CatData, "PostgreSQL XID/MultiXact freeze headroom per database.", PServer(), PHours(24), PAsOf()), + ["get_pg_xmin_horizon"] = R(CatData, "What is holding back the PostgreSQL xmin horizon, by cause.", PServer(), PHours(24), PAsOf()), + ["get_pg_replication_slots"] = R(CatData, "PostgreSQL replication slot health, including whether retained WAL is still growing.", PServer(), PHours(24), PAsOf()), + ["get_pg_autovacuum_health"] = R(CatData, "PostgreSQL tables behind on vacuum or analyze, ranked by how far past each table's own threshold.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_io_stats"] = R(CatData, "PostgreSQL I/O by backend type, object and context, differenced across the window.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_wait_stats"] = R(CatData, "Top PostgreSQL wait events in the window (Aurora targets).", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_wait_sampling"] = R(CatData, "Sampled PostgreSQL waits by query shape, from pg_wait_sampling - the stock-PostgreSQL counterpart of get_pg_wait_stats. Sample counts, not measured durations; event_type CPU means running rather than waiting.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_kernel_stats"] = R(CatData, "Per-query OS CPU (user and system), device bytes and major faults, from pg_stat_kcache. The CPU half of the elapsed time get_pg_top_queries reports.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_predicate_stats"] = R(CatData, "Which columns queries actually filter on and how selectively, from pg_qualstats. SAMPLED counts - the evidence behind an index recommendation.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_pg_index_bloat"] = R(CatData, "MEASURED PostgreSQL index bloat from pgstattuple: leaf density, fragmentation and reclaimable bytes. Rows carrying a skipped_reason were not measured.", PServer(), PHours(168), PLimit(25), PAsOf()), + ["get_pg_column_stats"] = R(CatData, "Per-column distribution statistics the PLANNER uses: n_distinct, null fraction, correlation and top-value frequency.", PServer(), PHours(168), PLimit(25), PAsOf()), + ["get_pg_buffer_usage"] = R(CatData, "What is resident in the shared buffer pool per relation, from pg_buffercache. Residency, not read volume.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_pg_extensions"] = R(CatData, "Which PostgreSQL extensions are installed, outdated, available or absent, per database. Usually the reason another read is empty.", PServer(), PHours(168), PLimit(50), PAsOf()), + ["get_pg_lock_stats"] = R(CatData, "Sampled PostgreSQL lock activity by type, mode and relation. A sample of pg_locks, not an event log; for who blocks whom use get_pg_blocking.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_pg_write_stats"] = R(CatData, "Checkpoint and WAL write activity across the window: timed versus requested checkpoints, buffers written by whom, and WAL volume.", PServer(), PHours(24), PAsOf()), + ["get_pg_server_config"] = R(CatData, "The PostgreSQL server's configuration from pg_settings, non-default first, saying where each value came from and whether changing it needs a restart. Reports pending_restart, where the file and the running server disagree.", PServer(), PLimit(100), PBool("include_defaults", false)), + ["get_pg_server_config_changes"] = R(CatData, "PostgreSQL configuration parameters whose value CHANGED in the window, old beside new. Nothing else can reconstruct this after the fact.", PServer(), PHours(168), PLimit(100), PAsOf()), + ["get_pg_deadlocks"] = R(CatData, "PostgreSQL deadlocks reported in the window, with the victim, the lock modes and resources, and the victim's statement. Needs nothing configured on the target.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_pg_deadlock_detail"] = R(CatData, "PostgreSQL deadlock graphs in full: the whole wait graph and every participant's statement, as the server wrote it. Newest first, or one by deadlock_hash.", PServer(), PText("deadlock_hash"), PLimit(5)), + ["get_pg_wait_trend"] = R(CatTrends, "One PostgreSQL wait event over time, per second. Omit wait_event to follow whichever dominates. Estimates from a sampling profiler, so the shape is the finding.", PServer(), PText("wait_event"), PHours(24), PAsOf()), + ["get_pg_query_duration_trend"] = R(CatTrends, "One PostgreSQL statement over time by queryid: what a single execution cost in each interval. Omit queryid for the busiest statement. The regression read.", PServer(), PText("queryid"), PHours(24), PAsOf()), + ["get_pg_io_trend"] = R(CatTrends, "One PostgreSQL (backend_type, context) pair over time: I/O rates per second, the hit ratio per interval, and latency where the server measures it. Omit both to follow whichever pair moved the most I/O.", PServer(), PText("backend_type"), PText("context"), PHours(24), PAsOf()), + ["get_pg_database_trend"] = R(CatTrends, "One PostgreSQL database over time: temp-file spills, the interval's own cache hit ratio, deadlocks and the rollback share. Omit database for the biggest spiller. The cumulative ratio is a lifetime average that hides a cliff.", PServer(), PText("database"), PHours(24), PAsOf()), + ["get_pg_replication_stats"] = R(CatData, "Health of CONNECTED replicas from pg_stat_replication, with the worst lag in the window beside the latest. Counterpart of get_pg_replication_slots.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_pg_blocking"] = R(CatData, "PostgreSQL blocking chains that were sampled, with the root blocker attributed. A sample, not an event log.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_pg_database_stats"] = R(CatData, "PostgreSQL per-database temp-file spills, cache hit ratio, deadlocks and commit/rollback split, differenced across the window.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_pg_index_usage"] = R(CatData, "PostgreSQL per-index scan counts and size, with the constraint, replica-identity and validity facts that decide whether an unused index can actually be dropped.", PServer(), PHours(168), PLimit(25), PAsOf()), + ["get_pg_table_bloat"] = R(CatData, "PostgreSQL per-table bloat ESTIMATE with its measured sizes and dead-tuple counts. The estimate is suppressed, not captioned, when its statistics cannot be trusted.", PServer(), PHours(168), PLimit(25), PAsOf()), + ["get_pg_session_states"] = R(CatData, "PostgreSQL sessions holding a transaction open, and whether each one actually pins the xmin horizon - which is not the same question as how long it has been idle in transaction.", PServer(), PHours(24), PLimit(25), PAsOf()), + ["get_wait_stats"] = R(CatData, "Top wait statistics in the window.", PServer(), PHours(24), PLimit(20), PAsOf()), + ["get_wait_trend"] = R(CatData, "One wait type's totals over time (requires wait_type).", PReqText("wait_type"), PServer(), PHours(24), PAsOf()), + ["get_wait_types"] = R(CatData, "The wait types observed in the window.", PServer(), PHours(24), PAsOf()), ["list_servers"] = R(CatData, "The monitored servers known to the store."), /* ── trends (DarlingMcpTrendTools) ── */ - ["get_file_io_trend"] = R(CatTrends, "File-IO throughput over time.", PServer(), PHours(24)), - ["get_memory_trend"] = R(CatTrends, "Memory usage over time.", PServer(), PHours(24)), - ["get_perfmon_trend"] = R(CatTrends, "One perfmon counter over time (requires counter_name).", PReqText("counter_name"), PServer(), PHours(24)), - ["get_query_duration_trend"] = R(CatTrends, "Query-duration percentiles over time.", PServer(), PHours(24)), - ["get_query_trend"] = R(CatTrends, "One query's metrics over time (requires query_hash + database_name).", PReqText("query_hash"), PReqText("database_name"), PServer(), PHours(24)), + ["get_file_io_trend"] = R(CatTrends, "File-IO throughput over time.", PServer(), PHours(24), PAsOf()), + ["get_memory_trend"] = R(CatTrends, "Memory usage over time.", PServer(), PHours(24), PAsOf()), + ["get_perfmon_trend"] = R(CatTrends, "One perfmon counter over time (requires counter_name).", PReqText("counter_name"), PServer(), PHours(24), PAsOf()), + ["get_procedure_duration_trend"] = R(CatTrends, "Stored-procedure elapsed ms/sec + executions/sec over time.", PServer(), PHours(24), PAsOf()), + ["get_query_duration_trend"] = R(CatTrends, "Query-duration percentiles over time.", PServer(), PHours(24), PAsOf()), + ["get_query_store_duration_trend"] = R(CatTrends, "Query Store duration ms/sec + executions/sec over time.", PServer(), PHours(24), PAsOf()), + ["get_query_trend"] = R(CatTrends, "One query's metrics over time (requires query_hash + database_name).", PReqText("query_hash"), PReqText("database_name"), PServer(), PHours(24), PAsOf()), /* ── health / overview (DarlingMcpHealthTools / DarlingMcpFleetTools) ── */ ["get_server_summary"] = R(CatOverview, "A one-shot health summary for a server.", PServer()), ["get_daily_summary"] = R(CatOverview, "The daily health summary (optionally for a specific date).", PServer(), PText("summary_date")), + ["get_daily_summary_range"] = R(CatOverview, "One daily health summary per collected day over a span of days - the Performance Calendar's month grid.", PServer(), PInt("days_back", 30), PAsOf()), ["get_fleet_overview"] = R(CatOverview, "The banded cross-server fleet roll-up.", PHours(DefaultFleetHours)), ["get_ag_health"] = R(CatOverview, "Availability Group topology: replicas and per-database secondary state.", PServer()), ["get_store_metrics"] = R(CatOverview, "The monitoring store's own size/compression/growth series (self-metrics).", PInt("days_back", 30)), /* ── latch / spinlock (DarlingMcpLatchSpinlockTools) ── */ - ["get_latch_stats"] = R(CatLatch, "Top latch waits in the window.", PServer(), PHours(24), PTop(10)), - ["get_spinlock_stats"] = R(CatLatch, "Top spinlock activity in the window.", PServer(), PHours(24), PTop(10)), + ["get_latch_stats"] = R(CatLatch, "Top latch waits in the window.", PServer(), PHours(24), PTop(10), PAsOf()), + ["get_spinlock_stats"] = R(CatLatch, "Top spinlock activity in the window.", PServer(), PHours(24), PTop(10), PAsOf()), /* ── memory grants (DarlingMcpMemoryGrantTools) ── */ - ["get_memory_grants"] = R(CatMemoryGrants, "Active/recent memory grants.", PServer(), PHours(1)), - ["get_memory_pressure_events"] = R(CatMemoryGrants, "Memory-pressure events in the window.", PServer(), PHours(24)), - ["get_resource_semaphore"] = R(CatMemoryGrants, "Resource-semaphore state over time.", PServer(), PHours(24)), + ["get_memory_grants"] = R(CatMemoryGrants, "Active/recent memory grants.", PServer(), PHours(1), PAsOf()), + ["get_memory_pressure_events"] = R(CatMemoryGrants, "Memory-pressure events in the window.", PServer(), PHours(24), PAsOf()), + ["get_resource_semaphore"] = R(CatMemoryGrants, "Resource-semaphore state over time.", PServer(), PHours(24), PAsOf()), /* ── object / index stats (DarlingMcpObjectStatsTools) ── */ ["get_database_sizes"] = R(CatObjects, "Per-database size breakdown.", PServer()), ["get_pvs_stats"] = R(CatObjects, "ADR persistent version store state per database, with an optional top-5 size trend.", PServer(), PInt("trend_hours_back", 0)), - ["get_index_usage"] = R(CatObjects, "Index usage (seeks/scans/updates) per index.", PServer()), + ["get_index_usage"] = R(CatObjects, "Index usage (seeks/scans/updates) per index. Unused-first, so pass database_name unless you want a server-wide sweep; the answer carries matching_index_count and truncated.", PServer(), PText("database_name"), PLimit(200)), ["get_object_locking"] = R(CatObjects, "Per-object locking/contention stats.", PServer()), ["get_table_index_sizes"] = R(CatObjects, "Per-table/index size breakdown.", PServer()), /* ── plan cache / scheduler (DarlingMcpPlanCacheSchedulerTools) ── */ ["get_cpu_scheduler_pressure"] = R(CatPlanCache, "CPU scheduler pressure indicators.", PServer()), - ["get_plan_cache_bloat"] = R(CatPlanCache, "Plan-cache bloat / single-use plan indicators.", PServer(), PHours(24)), + ["get_plan_cache_bloat"] = R(CatPlanCache, "Plan-cache bloat / single-use plan indicators.", PServer(), PHours(24), PAsOf()), /* ── jobs (DarlingMcpJobTools) ── */ ["get_running_jobs"] = R(CatJobs, "Currently-running SQL Agent jobs.", PServer()), @@ -1169,17 +1209,18 @@ private static CatalogRead R(string category, string description, params Catalog ["get_plan_xml"] = R(CatPlans, "The stored execution-plan XML for a query (requires query_hash).", PReqText("query_hash"), PServer(), PText("database_name")), /* ── default trace (DarlingMcpDefaultTraceTools) ── */ - ["get_default_trace_events"] = R(CatDefaultTrace, "Default-trace events (file growth, DDL, security).", PServer(), PHours(24), PLimit(100)), + ["get_default_trace_events"] = R(CatDefaultTrace, "Default-trace events (file growth, DDL, security).", PServer(), PHours(24), PLimit(100), PAsOf()), /* ── system_health parse-on-read family (DarlingMcpHealthParserTools) ── */ - ["get_health_parser_cpu_tasks"] = R(CatSystemHealth, "system_health: CPU-bound task snapshots.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_io_issues"] = R(CatSystemHealth, "system_health: IO-latency issues.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_memory_broker"] = R(CatSystemHealth, "system_health: memory-broker ring-buffer entries.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_memory_conditions"] = R(CatSystemHealth, "system_health: resource-monitor memory conditions.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_memory_node_oom"] = R(CatSystemHealth, "system_health: per-node out-of-memory events.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_scheduler_issues"] = R(CatSystemHealth, "system_health: non-yielding scheduler issues.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_severe_errors"] = R(CatSystemHealth, "system_health: severe (sev >= 17) errors.", PServer(), PHours(24), PLimit(50)), - ["get_health_parser_system_health"] = R(CatSystemHealth, "system_health: the raw parsed session records.", PServer(), PHours(24), PLimit(50)), + ["get_health_parser_cpu_tasks"] = R(CatSystemHealth, "system_health: CPU-bound task snapshots.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_io_issues"] = R(CatSystemHealth, "system_health: IO-latency issues.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_memory_broker"] = R(CatSystemHealth, "system_health: memory-broker ring-buffer entries.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_memory_conditions"] = R(CatSystemHealth, "system_health: resource-monitor memory conditions.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_memory_node_oom"] = R(CatSystemHealth, "system_health: per-node out-of-memory events.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_scheduler_issues"] = R(CatSystemHealth, "system_health: non-yielding scheduler issues.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_severe_errors"] = R(CatSystemHealth, "system_health: severe (sev >= 17) errors.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_significant_waits"] = R(CatSystemHealth, "system_health: individual 500 ms+ waits with their statement.", PServer(), PHours(24), PLimit(50), PAsOf()), + ["get_health_parser_system_health"] = R(CatSystemHealth, "system_health: the raw parsed session records.", PServer(), PHours(24), PLimit(50), PAsOf()), }; /// Builds the /api/catalog body: the reads (names taken from , @@ -1330,10 +1371,21 @@ dimensions list — surfaced here so the composer offers it on every measure. */ /* Representative target profiles a measure's collector AppliesTo gate is evaluated against (design D4). SqlMajorVersion is pinned to a supported major (16 = SQL 2022) so the version-gated collectors - (query_stats/query_store, 2016+) read as available; the msdb-less on-prem profile isolates the SQL-Agent - dependency for the needsMsdb flag. */ + (query_stats/query_store, 2016+) read as available. */ private static readonly CollectorTargetInfo s_onPremTarget = new() { SqlMajorVersion = 16, HasMsdbAccess = true }; - private static readonly CollectorTargetInfo s_onPremNoMsdbTarget = new() { SqlMajorVersion = 16, HasMsdbAccess = false }; + + /* The SQL-Agent collectors, NAMED rather than probed. This used to be detected by evaluating the gate + against an msdb-less profile - "available with msdb, gated without it" - which stopped working when + #2559 removed HasMsdbAccess from those gates. The dependency itself did not go away: these three read + msdb.dbo.sysjobs and friends, so a panel built on them still needs the grant. What changed is what its + absence produces - a PERMISSIONS row rather than no dispatch - which arguably makes the badge MORE + useful, since the data starts flowing the moment the grant lands with no reconnect. + + Honest limit: this cannot notice a NEW collector that reads msdb. DarlingComposeTests pins that the set + is exactly these three and that each one's query really does reference msdb, which catches a rename or + a collector that stopped reading it - not an addition. */ + private static readonly HashSet s_msdbBackedTables = + new(StringComparer.Ordinal) { "running_jobs", "job_history", "agent_status" }; private static readonly CollectorTargetInfo s_azureSqlDbTarget = new() { IsAzureSqlDb = true, SqlMajorVersion = 16, HasMsdbAccess = false }; private static readonly CollectorTargetInfo s_azureMiTarget = new() { IsAzureManagedInstance = true, SqlMajorVersion = 16, HasMsdbAccess = true }; private static readonly CollectorTargetInfo s_awsRdsTarget = new() { IsAwsRds = true, SqlMajorVersion = 16, HasMsdbAccess = true }; @@ -1341,8 +1393,9 @@ SqlMajorVersion is pinned to a supported major (16 = SQL 2022) so the version-ga /// The per-server-type availability of a measure (design D4), derived from its owning collector's /// gate — the single authoritative target gate — so the composer /// can label/grey a measure a given server type can't collect. needsMsdb is the SQL-Agent dependency - /// (job/agent measures), detected as "available with msdb, gated without it" on an otherwise-capable on-prem - /// target. Returns null only if a measure's source has no collector (impossible — pinned by test). + /// (job/agent measures), taken from the named SQL-Agent set rather than probed from the gate - see the + /// note there on why #2559 made probing impossible and why the badge still earns its place. + /// Returns null only if a measure's source has no collector (impossible — pinned by test). private static JsonObject? BuildAppliesToNode(string sourceTable) { if (!s_collectorByTable.TryGetValue(sourceTable, out var collector)) @@ -1356,7 +1409,7 @@ SqlMajorVersion is pinned to a supported major (16 = SQL 2022) so the version-ga ["azureSqlDb"] = collector.AppliesTo(s_azureSqlDbTarget), ["azureMi"] = collector.AppliesTo(s_azureMiTarget), ["awsRds"] = collector.AppliesTo(s_awsRdsTarget), - ["needsMsdb"] = collector.AppliesTo(s_onPremTarget) && !collector.AppliesTo(s_onPremNoMsdbTarget), + ["needsMsdb"] = s_msdbBackedTables.Contains(sourceTable), }; } @@ -1503,108 +1556,143 @@ internal static IReadOnlyDictionary BuildReadDispatch() { /* ── analysis reads (take the DarlingAnalysisService) ── */ ["audit_config"] = (c, pg, an) => DarlingMcpTools.AuditConfig(an, pg, Server(c)), - ["compare_analysis"] = (c, pg, an) => DarlingMcpTools.CompareAnalysis(an, pg, Server(c), Hours(c, 4), QueryInt(c, "baseline_hours_back", null, 28)), - ["get_analysis_facts"] = (c, pg, an) => DarlingMcpTools.GetAnalysisFacts(an, pg, Server(c), Hours(c, 4), Str(c, "source"), QueryDouble(c, "min_severity", 0)), - ["get_analysis_findings"] = (c, pg, an) => DarlingMcpTools.GetAnalysisFindings(an, pg, Server(c), Hours(c, 24), QueryBool(c, "include_drilldown", false)), + ["compare_analysis"] = (c, pg, an) => DarlingMcpTools.CompareAnalysis(an, pg, Server(c), Hours(c, 4), QueryInt(c, "baseline_hours_back", null, 28), as_of: AsOf(c)), + ["get_analysis_facts"] = (c, pg, an) => DarlingMcpTools.GetAnalysisFacts(an, pg, Server(c), Hours(c, 4), Str(c, "source"), QueryDouble(c, "min_severity", 0), as_of: AsOf(c)), + ["get_analysis_findings"] = (c, pg, an) => DarlingMcpTools.GetAnalysisFindings(an, pg, Server(c), Hours(c, 24), QueryBool(c, "include_drilldown", false), as_of: AsOf(c)), /* ── sessions ── */ - ["get_active_queries"] = (c, pg, an) => DarlingMcpSessionTools.GetActiveQueries(pg, Server(c), Hours(c, 1), Str(c, "database_name"), QueryBool(c, "blocking_only", false), Rows(c, "limit", 50)), + ["get_active_queries"] = (c, pg, an) => DarlingMcpSessionTools.GetActiveQueries(pg, Server(c), Hours(c, 1), Str(c, "database_name"), QueryBool(c, "blocking_only", false), Rows(c, "limit", 50), as_of: AsOf(c)), ["get_session_stats"] = (c, pg, an) => DarlingMcpSessionTools.GetSessionStats(pg, Server(c)), - ["get_waiting_tasks"] = (c, pg, an) => DarlingMcpSessionTools.GetWaitingTasks(pg, Server(c), Hours(c, 1), Rows(c, "limit", 30)), + ["get_waiting_tasks"] = (c, pg, an) => DarlingMcpSessionTools.GetWaitingTasks(pg, Server(c), Hours(c, 1), Rows(c, "limit", 30), as_of: AsOf(c)), /* ── alerts / mute rules ── */ - ["get_alert_history"] = (c, pg, an) => DarlingMcpAlertTools.GetAlertHistory(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), + ["get_alert_history"] = (c, pg, an) => DarlingMcpAlertTools.GetAlertHistory(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), ["get_alert_settings"] = (c, pg, an) => DarlingMcpAlertTools.GetAlertSettings(pg), ["get_mute_rules"] = (c, pg, an) => DarlingMcpAlertTools.GetMuteRules(pg, QueryBool(c, "enabled_only", true)), /* ── blocking / deadlocks ── */ - ["get_blocked_process_xml"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlockedProcessXml(pg, Server(c), Hours(c, 24), Rows(c, "limit", 5)), - ["get_blocking"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlocking(pg, Server(c), Hours(c, 24), Rows(c, "limit", 30)), - ["get_blocking_trend"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlockingTrend(pg, Server(c), Hours(c, 24)), - ["get_deadlock_detail"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlockDetail(pg, Server(c), Hours(c, 24), Rows(c, "limit", 5)), - ["get_deadlock_trend"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlockTrend(pg, Server(c), Hours(c, 24)), - ["get_deadlocks"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlocks(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), + ["get_blocked_process_xml"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlockedProcessXml(pg, Server(c), Hours(c, 24), Rows(c, "limit", 5), as_of: AsOf(c)), + ["get_blocking"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlocking(pg, Server(c), Hours(c, 24), Rows(c, "limit", 30), as_of: AsOf(c)), + ["get_blocking_trend"] = (c, pg, an) => DarlingMcpBlockingTools.GetBlockingTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_deadlock_detail"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlockDetail(pg, Server(c), Hours(c, 24), Rows(c, "limit", 5), as_of: AsOf(c)), + ["get_deadlock_trend"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlockTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_deadlocks"] = (c, pg, an) => DarlingMcpBlockingTools.GetDeadlocks(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_lock_wait_trend"] = (c, pg, an) => DarlingMcpBlockingTools.GetLockWaitTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), /* ── automatic plan correction (#2028) ── */ - ["get_plan_corrections"] = (c, pg, an) => DarlingMcpPlanCorrectionTools.GetPlanCorrections(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), + ["get_plan_corrections"] = (c, pg, an) => DarlingMcpPlanCorrectionTools.GetPlanCorrections(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), /* ── config (current + history) ── */ ["get_database_config"] = (c, pg, an) => DarlingMcpConfigTools.GetDatabaseConfig(pg, Server(c), Str(c, "database_name")), ["get_server_config"] = (c, pg, an) => DarlingMcpConfigTools.GetServerConfig(pg, Server(c)), ["get_trace_flags"] = (c, pg, an) => DarlingMcpConfigTools.GetTraceFlags(pg, Server(c)), - ["get_database_config_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetDatabaseConfigChanges(pg, Server(c), Hours(c, 168)), + ["get_database_config_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetDatabaseConfigChanges(pg, Server(c), Hours(c, 168), as_of: AsOf(c)), ["get_database_scoped_config"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetDatabaseScopedConfig(pg, Server(c), Str(c, "database_name")), ["get_query_store_health"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetQueryStoreHealth(pg, Server(c), Str(c, "database_name")), - ["get_server_config_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetServerConfigChanges(pg, Server(c), Hours(c, 168)), - ["get_trace_flag_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetTraceFlagChanges(pg, Server(c), Hours(c, 168)), + ["get_server_config_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetServerConfigChanges(pg, Server(c), Hours(c, 168), as_of: AsOf(c)), + ["get_trace_flag_changes"] = (c, pg, an) => DarlingMcpConfigHistoryTools.GetTraceFlagChanges(pg, Server(c), Hours(c, 168), as_of: AsOf(c)), /* ── core data reads ── */ ["get_collection_health"] = (c, pg, an) => DarlingMcpDataTools.GetCollectionHealth(pg, Server(c)), - ["get_cpu_utilization"] = (c, pg, an) => DarlingMcpDataTools.GetCpuUtilization(pg, Server(c), Hours(c, 4)), + ["get_collection_log"] = (c, pg, an) => DarlingMcpDataTools.GetCollectionLog(pg, Server(c), Hours(c, 24), Rows(c, "limit", 200), as_of: AsOf(c)), + ["get_current_waits_trend"] = (c, pg, an) => DarlingMcpDataTools.GetCurrentWaitsTrend(pg, Server(c), Hours(c, 4), Str(c, "database_name"), as_of: AsOf(c)), + ["get_blocking_stats"] = (c, pg, an) => DarlingMcpDataTools.GetBlockingStats(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_cpu_utilization"] = (c, pg, an) => DarlingMcpDataTools.GetCpuUtilization(pg, Server(c), Hours(c, 4), as_of: AsOf(c)), ["get_file_io_stats"] = (c, pg, an) => DarlingMcpDataTools.GetFileIoStats(pg, Server(c)), ["get_memory_clerks"] = (c, pg, an) => DarlingMcpDataTools.GetMemoryClerks(pg, Server(c)), ["get_memory_stats"] = (c, pg, an) => DarlingMcpDataTools.GetMemoryStats(pg, Server(c)), ["get_perfmon_stats"] = (c, pg, an) => DarlingMcpDataTools.GetPerfmonStats(pg, Server(c), Str(c, "counter_name"), Str(c, "instance_name")), - ["get_query_store_top"] = (c, pg, an) => DarlingMcpDataTools.GetQueryStoreTop(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name")), - ["get_long_query_completions"] = (c, pg, an) => DarlingMcpLongQueryTools.GetLongQueryCompletions(pg, Server(c), Hours(c, 24), Rows(c, "limit", 30)), + ["get_query_heatmap"] = (c, pg, an) => DarlingMcpQueryHeatmapTools.GetQueryHeatmap(pg, Server(c), Hours(c, 24), Str(c, "metric"), Str(c, "database_name"), QueryInt(c, "bucket_minutes", null, 5), Rows(c, "limit", 500), as_of: AsOf(c)), + ["get_query_store_regressions"] = (c, pg, an) => DarlingMcpQueryStoreRegressionTools.GetQueryStoreRegressions(pg, Server(c), Hours(c, 24), Str(c, "database_name"), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_query_store_top"] = (c, pg, an) => DarlingMcpDataTools.GetQueryStoreTop(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name"), as_of: AsOf(c)), + ["get_long_query_completions"] = (c, pg, an) => DarlingMcpLongQueryTools.GetLongQueryCompletions(pg, Server(c), Hours(c, 24), Rows(c, "limit", 30), as_of: AsOf(c)), ["get_server_properties"] = (c, pg, an) => DarlingMcpDataTools.GetServerProperties(pg, Server(c)), - ["get_tempdb_trend"] = (c, pg, an) => DarlingMcpDataTools.GetTempDbTrend(pg, Server(c), Hours(c, 24)), - ["get_top_procedures_by_cpu"] = (c, pg, an) => DarlingMcpDataTools.GetTopProceduresByCpu(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name")), - ["get_top_queries_by_cpu"] = (c, pg, an) => DarlingMcpDataTools.GetTopQueriesByCpu(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name"), QueryBool(c, "parallel_only", false), QueryInt(c, "min_dop", null, 0)), - ["get_pg_top_queries"] = (c, pg, an) => DarlingMcpPgStatementTools.GetPgTopQueries(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), - ["get_pg_wraparound_risk"] = (c, pg, an) => DarlingMcpPgWraparoundTools.GetPgWraparoundRisk(pg, Server(c), Hours(c, 24)), - ["get_pg_xmin_horizon"] = (c, pg, an) => DarlingMcpPgXminTools.GetPgXminHorizon(pg, Server(c), Hours(c, 24)), - ["get_pg_replication_slots"] = (c, pg, an) => DarlingMcpPgSlotTools.GetPgReplicationSlots(pg, Server(c), Hours(c, 24)), - ["get_pg_autovacuum_health"] = (c, pg, an) => DarlingMcpPgAutovacuumTools.GetPgAutovacuumHealth(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), - ["get_pg_io_stats"] = (c, pg, an) => DarlingMcpPgIoTools.GetPgIoStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), - ["get_pg_wait_stats"] = (c, pg, an) => DarlingMcpPgWaitTools.GetPgWaitStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), - ["get_pg_blocking"] = (c, pg, an) => DarlingMcpPgBlockingTools.GetPgBlocking(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_wait_stats"] = (c, pg, an) => DarlingMcpDataTools.GetWaitStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20)), + ["get_tempdb_trend"] = (c, pg, an) => DarlingMcpDataTools.GetTempDbTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_top_procedures_by_cpu"] = (c, pg, an) => DarlingMcpDataTools.GetTopProceduresByCpu(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name"), as_of: AsOf(c)), + ["get_top_queries_by_cpu"] = (c, pg, an) => DarlingMcpDataTools.GetTopQueriesByCpu(pg, Server(c), Hours(c, 24), Rows(c, "top", 20), Str(c, "database_name"), QueryBool(c, "parallel_only", false), QueryInt(c, "min_dop", null, 0), as_of: AsOf(c)), + ["get_pg_top_queries"] = (c, pg, an) => DarlingMcpPgStatementTools.GetPgTopQueries(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + /* query_id arrives as TEXT and is passed through as text (#2548): a queryid that made a + round trip through a JSON number has already been rounded, and the tool rejects one it + cannot parse exactly rather than silently matching nothing. */ + ["get_pg_plans"] = (c, pg, an) => DarlingMcpPgPlanTools.GetPgPlans(pg, Server(c), Hours(c, 24), Rows(c, "limit", 10), Str(c, "query_id"), AsOf(c)), + ["get_pg_wraparound_risk"] = (c, pg, an) => DarlingMcpPgWraparoundTools.GetPgWraparoundRisk(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_xmin_horizon"] = (c, pg, an) => DarlingMcpPgXminTools.GetPgXminHorizon(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_replication_slots"] = (c, pg, an) => DarlingMcpPgSlotTools.GetPgReplicationSlots(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_autovacuum_health"] = (c, pg, an) => DarlingMcpPgAutovacuumTools.GetPgAutovacuumHealth(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_io_stats"] = (c, pg, an) => DarlingMcpPgIoTools.GetPgIoStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_wait_stats"] = (c, pg, an) => DarlingMcpPgWaitTools.GetPgWaitStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_wait_sampling"] = (c, pg, an) => DarlingMcpPgWaitSamplingTools.GetPgWaitSampling(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_kernel_stats"] = (c, pg, an) => DarlingMcpPgKernelStatsTools.GetPgKernelStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_predicate_stats"] = (c, pg, an) => DarlingMcpPgPredicateTools.GetPgPredicateStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_index_bloat"] = (c, pg, an) => DarlingMcpPgIndexTools.GetPgIndexBloat(pg, Server(c), Hours(c, 168), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_column_stats"] = (c, pg, an) => DarlingMcpPgIndexTools.GetPgColumnStats(pg, Server(c), Hours(c, 168), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_buffer_usage"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgBufferUsage(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_extensions"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgExtensions(pg, Server(c), Hours(c, 168), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_pg_lock_stats"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgLockStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_write_stats"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgWriteStats(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_server_config"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgServerConfig(pg, Server(c), Rows(c, "limit", 100), QueryBool(c, "include_defaults", false)), + ["get_pg_server_config_changes"] = (c, pg, an) => DarlingMcpPgServerStateTools.GetPgServerConfigChanges(pg, Server(c), Hours(c, 168), Rows(c, "limit", 100), as_of: AsOf(c)), + ["get_pg_deadlocks"] = (c, pg, an) => DarlingMcpPgDeadlockTools.GetPgDeadlocks(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_deadlock_detail"] = (c, pg, an) => DarlingMcpPgDeadlockTools.GetPgDeadlockDetail(pg, Server(c), Str(c, "deadlock_hash"), Rows(c, "limit", 5)), + ["get_pg_wait_trend"] = (c, pg, an) => DarlingMcpPgTrendTools.GetPgWaitTrend(pg, Server(c), Str(c, "wait_event"), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_query_duration_trend"] = (c, pg, an) => DarlingMcpPgTrendTools.GetPgQueryDurationTrend(pg, Server(c), Str(c, "queryid"), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_io_trend"] = (c, pg, an) => DarlingMcpPgTrendTools.GetPgIoTrend(pg, Server(c), Str(c, "backend_type"), Str(c, "context"), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_database_trend"] = (c, pg, an) => DarlingMcpPgTrendTools.GetPgDatabaseTrend(pg, Server(c), Str(c, "database"), Hours(c, 24), as_of: AsOf(c)), + ["get_pg_replication_stats"] = (c, pg, an) => DarlingMcpPgReplicationStatsTools.GetPgReplicationStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_blocking"] = (c, pg, an) => DarlingMcpPgBlockingTools.GetPgBlocking(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_pg_database_stats"] = (c, pg, an) => DarlingMcpPgDatabaseTools.GetPgDatabaseStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), + ["get_pg_index_usage"] = (c, pg, an) => DarlingMcpPgIndexUsageTools.GetPgIndexUsage(pg, Server(c), Hours(c, 168), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_table_bloat"] = (c, pg, an) => DarlingMcpPgTableBloatTools.GetPgTableBloat(pg, Server(c), Hours(c, 168), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_pg_session_states"] = (c, pg, an) => DarlingMcpPgSessionStatesTools.GetPgSessionStates(pg, Server(c), Hours(c, 24), Rows(c, "limit", 25), as_of: AsOf(c)), + ["get_wait_stats"] = (c, pg, an) => DarlingMcpDataTools.GetWaitStats(pg, Server(c), Hours(c, 24), Rows(c, "limit", 20), as_of: AsOf(c)), ["get_wait_trend"] = (c, pg, an) => RequireText(c, "wait_type", out var waitType) - ? DarlingMcpDataTools.GetWaitTrend(pg, waitType, Server(c), Hours(c, 24)) + ? DarlingMcpDataTools.GetWaitTrend(pg, waitType, Server(c), Hours(c, 24), as_of: AsOf(c)) : MissingParam("wait_type"), - ["get_wait_types"] = (c, pg, an) => DarlingMcpDataTools.GetWaitTypes(pg, Server(c), Hours(c, 24)), + ["get_wait_types"] = (c, pg, an) => DarlingMcpDataTools.GetWaitTypes(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), ["list_servers"] = (c, pg, an) => DarlingMcpDataTools.ListServers(pg), /* ── trends ── */ - ["get_file_io_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetFileIoTrend(pg, Server(c), Hours(c, 24)), - ["get_memory_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetMemoryTrend(pg, Server(c), Hours(c, 24)), + ["get_file_io_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetFileIoTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_memory_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetMemoryTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), ["get_perfmon_trend"] = (c, pg, an) => RequireText(c, "counter_name", out var counter) - ? DarlingMcpTrendTools.GetPerfmonTrend(pg, counter, Server(c), Hours(c, 24)) + ? DarlingMcpTrendTools.GetPerfmonTrend(pg, counter, Server(c), Hours(c, 24), as_of: AsOf(c)) : MissingParam("counter_name"), - ["get_query_duration_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetQueryDurationTrend(pg, Server(c), Hours(c, 24)), + ["get_procedure_duration_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetProcedureDurationTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_query_duration_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetQueryDurationTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_query_store_duration_trend"] = (c, pg, an) => DarlingMcpTrendTools.GetQueryStoreDurationTrend(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), ["get_query_trend"] = (c, pg, an) => RequireText(c, "query_hash", out var queryHash) ? (RequireText(c, "database_name", out var db) - ? DarlingMcpTrendTools.GetQueryTrend(pg, queryHash, db, Server(c), Hours(c, 24)) + ? DarlingMcpTrendTools.GetQueryTrend(pg, queryHash, db, Server(c), Hours(c, 24), as_of: AsOf(c)) : MissingParam("database_name")) : MissingParam("query_hash"), /* ── health / overview ── */ ["get_server_summary"] = (c, pg, an) => DarlingMcpHealthTools.GetServerSummary(pg, Server(c)), ["get_daily_summary"] = (c, pg, an) => DarlingMcpHealthTools.GetDailySummary(pg, Server(c), Str(c, "summary_date")), + ["get_daily_summary_range"] = (c, pg, an) => DarlingMcpHealthTools.GetDailySummaryRange(pg, Server(c), QueryInt(c, "days_back", null, 30), AsOf(c)), ["get_fleet_overview"] = (c, pg, an) => DarlingMcpFleetTools.GetFleetOverview(pg, Hours(c, DefaultFleetHours)), ["get_ag_health"] = (c, pg, an) => DarlingMcpAgTools.GetAgHealth(pg, Server(c)), ["get_store_metrics"] = (c, pg, an) => DarlingMcpStoreMetricsTools.GetStoreMetrics(pg, QueryInt(c, "days_back", null, 30)), /* ── latch / spinlock ── */ - ["get_latch_stats"] = (c, pg, an) => DarlingMcpLatchSpinlockTools.GetLatchStats(pg, Server(c), Hours(c, 24), Rows(c, "top", 10)), - ["get_spinlock_stats"] = (c, pg, an) => DarlingMcpLatchSpinlockTools.GetSpinlockStats(pg, Server(c), Hours(c, 24), Rows(c, "top", 10)), + ["get_latch_stats"] = (c, pg, an) => DarlingMcpLatchSpinlockTools.GetLatchStats(pg, Server(c), Hours(c, 24), Rows(c, "top", 10), as_of: AsOf(c)), + ["get_spinlock_stats"] = (c, pg, an) => DarlingMcpLatchSpinlockTools.GetSpinlockStats(pg, Server(c), Hours(c, 24), Rows(c, "top", 10), as_of: AsOf(c)), /* ── memory grants ── */ - ["get_memory_grants"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetMemoryGrants(pg, Server(c), Hours(c, 1)), - ["get_memory_pressure_events"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetMemoryPressureEvents(pg, Server(c), Hours(c, 24)), - ["get_resource_semaphore"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetResourceSemaphore(pg, Server(c), Hours(c, 24)), + ["get_memory_grants"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetMemoryGrants(pg, Server(c), Hours(c, 1), as_of: AsOf(c)), + ["get_memory_pressure_events"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetMemoryPressureEvents(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), + ["get_resource_semaphore"] = (c, pg, an) => DarlingMcpMemoryGrantTools.GetResourceSemaphore(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), /* ── object / index stats ── */ ["get_database_sizes"] = (c, pg, an) => DarlingMcpObjectStatsTools.GetDatabaseSizes(pg, Server(c)), ["get_pvs_stats"] = (c, pg, an) => DarlingMcpPvsTools.GetPvsStats(pg, Server(c), QueryInt(c, "trend_hours_back", null, 0)), - ["get_index_usage"] = (c, pg, an) => DarlingMcpObjectStatsTools.GetIndexUsage(pg, Server(c)), + ["get_index_usage"] = (c, pg, an) => DarlingMcpObjectStatsTools.GetIndexUsage(pg, Server(c), Str(c, "database_name"), Rows(c, "limit", 200)), ["get_object_locking"] = (c, pg, an) => DarlingMcpObjectStatsTools.GetObjectLocking(pg, Server(c)), ["get_table_index_sizes"] = (c, pg, an) => DarlingMcpObjectStatsTools.GetTableIndexSizes(pg, Server(c)), /* ── plan cache / scheduler ── */ ["get_cpu_scheduler_pressure"] = (c, pg, an) => DarlingMcpPlanCacheSchedulerTools.GetCpuSchedulerPressure(pg, Server(c)), - ["get_plan_cache_bloat"] = (c, pg, an) => DarlingMcpPlanCacheSchedulerTools.GetPlanCacheBloat(pg, Server(c), Hours(c, 24)), + ["get_plan_cache_bloat"] = (c, pg, an) => DarlingMcpPlanCacheSchedulerTools.GetPlanCacheBloat(pg, Server(c), Hours(c, 24), as_of: AsOf(c)), /* ── jobs ── */ ["get_running_jobs"] = (c, pg, an) => DarlingMcpJobTools.GetRunningJobs(pg, Server(c)), @@ -1615,17 +1703,18 @@ internal static IReadOnlyDictionary BuildReadDispatch() : MissingParam("query_hash"), /* ── default trace ── */ - ["get_default_trace_events"] = (c, pg, an) => DarlingMcpDefaultTraceTools.GetDefaultTraceEvents(pg, Server(c), Hours(c, 24), Rows(c, "limit", 100)), + ["get_default_trace_events"] = (c, pg, an) => DarlingMcpDefaultTraceTools.GetDefaultTraceEvents(pg, Server(c), Hours(c, 24), Rows(c, "limit", 100), as_of: AsOf(c)), /* ── system_health parse-on-read family ── */ - ["get_health_parser_cpu_tasks"] = (c, pg, an) => DarlingMcpHealthParserTools.GetCPUTasks(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_io_issues"] = (c, pg, an) => DarlingMcpHealthParserTools.GetIOIssues(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_memory_broker"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryBroker(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_memory_conditions"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryConditions(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_memory_node_oom"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryNodeOOM(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_scheduler_issues"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSchedulerIssues(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_severe_errors"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSevereErrors(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), - ["get_health_parser_system_health"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSystemHealth(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50)), + ["get_health_parser_cpu_tasks"] = (c, pg, an) => DarlingMcpHealthParserTools.GetCPUTasks(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_io_issues"] = (c, pg, an) => DarlingMcpHealthParserTools.GetIOIssues(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_memory_broker"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryBroker(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_memory_conditions"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryConditions(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_memory_node_oom"] = (c, pg, an) => DarlingMcpHealthParserTools.GetMemoryNodeOOM(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_scheduler_issues"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSchedulerIssues(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_severe_errors"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSevereErrors(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_significant_waits"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSignificantWaits(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), + ["get_health_parser_system_health"] = (c, pg, an) => DarlingMcpHealthParserTools.GetSystemHealth(pg, Server(c), Hours(c, 24), Rows(c, "limit", 50), as_of: AsOf(c)), }; } @@ -1682,6 +1771,14 @@ internal static ToolResponseKind ClassifyToolResponse(string result) /// The hours-back window from ?hours= (or the tool's own ?hours_back=), else the tool's default. private static int Hours(HttpContext context, int def) => QueryInt(context, "hours", "hours_back", def); + /// + /// The window ANCHOR from ?as_of=; null when absent, which is what makes the window end at now. + /// Passed through UNVALIDATED on purpose — the tool owns the parse and the refusal message, so the web and + /// MCP surfaces cannot disagree about what a bad anchor means, and the bare string reaches the same + /// 400 mapping every other client-correctable message does. + /// + private static string? AsOf(HttpContext context) => First(context, "as_of"); + /// An optional text parameter; null when absent or empty (so the tool sees its own default). private static string? Str(HttpContext context, string key) => First(context, key); diff --git a/Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs b/Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs index b33fea452..55c494aad 100644 --- a/Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs +++ b/Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -276,7 +276,7 @@ the loop's data source is disposed at RunCollectionLoopAsync scope exit. The pas /// /// BYO mode (managed=false) with any postgres.network.* or mcp.network.* /// set — the fields are IGNORED; the operator's own PostgreSQL governs exposure. - /// Managed mode with an EXPOSED store whose network.role is admin — names the + /// Managed mode with an EXPOSED store whose network.role admits admin — names the /// config_command / config_monitored_servers / config_notification /// service-credential pivot a remote admin connection can reach (D7 — the operator's informed opt-in). /// @@ -314,8 +314,12 @@ the same breath as 'Starting MCP server on 0.0.0.0'. The postgres.network notice && DarlingNetwork.IsExposedListenAddress(network.Listen) && string.Equals(DarlingNetwork.NormalizeNetworkRole(network.Role), "admin", StringComparison.Ordinal)) { + /* "admits" rather than "is": since #2665 the field can name both roles, and NormalizeNetworkRole + answers 'admin' for that too — correctly, because admin IS reachable. Wording it as "is 'admin'" + would read as wrong to the operator who wrote "admin,viewer" and invite them to dismiss the one + warning that matters here. */ warnings.Add( - "postgres.network.role is 'admin' — a REMOTE admin connection can write config_command (the test_connect service-credential pivot), config_monitored_servers, and config_notification (webhook exfil). This is an explicit opt-in; the secure default is 'viewer' (read-only). Only expose admin on a trusted network."); + "postgres.network.role admits 'admin' — a REMOTE admin connection can write config_command (the test_connect service-credential pivot), config_monitored_servers, and config_notification (webhook exfil). This is an explicit opt-in; the secure default is 'viewer' (read-only). Only expose admin on a trusted network."); } return warnings; @@ -348,8 +352,42 @@ across collection I/O — so a long-running command never blocks the collection /* Server IDs whose scheduled analysis is currently running — prevents relaunching analysis for a server whose previous (possibly hung) pass has not finished - (Lite's CollectionBackgroundService in-flight guard). */ - private readonly ConcurrentDictionary _analysisInFlight = new(); + (Lite's CollectionBackgroundService in-flight guard). The value carries when the pass started + and how loudly it has been reported, because #2430's defect was that a pass which never + finishes leaves this marker set FOREVER and the server is then skipped in silence on every + later cycle. */ + private readonly ConcurrentDictionary _analysisInFlight = new(); + + /* How far past its budget an in-flight pass must be before the sweep starts reporting it. The + ordinary overrun already gets the "exceeded Ns" warning; this is for the pass that ignored the + cancellation raised at that budget — wedged inside one of the store reads that still take no + token (see the note on RunAnalysisPassAsync), which no token can reach. */ + private const int StuckAnalysisMultiple = 3; + + /* Reports back off by doubling. A fixed repeat interval would either be slower than the analysis + cadence (useless) or produce one line per cycle forever (the spam that makes a log unreadable). + Doubling gives a handful of lines in the first day and a handful per day after — loud enough to + be noticed, quiet enough to stay noticed. Capped so the shift cannot run away on a service that + stays up for months, which this one does. */ + private const int StuckAnalysisMaxBackoffDoublings = 20; + + /// + /// Bookkeeping for one in-flight analysis pass (#2430). A class, not a struct, so the sweep can + /// update the report counters in place without a read-modify-write race against the completion + /// continuation, which only ever removes the whole entry. + /// + private sealed class AnalysisPassState + { + public AnalysisPassState(DateTime startedUtc) => StartedUtc = startedUtc; + + public DateTime StartedUtc { get; } + + /// Cycles this server has lost to the pass still being in flight. + public int SkippedCycles { get; set; } + + /// How many times it has been reported, which is what the backoff doubles on. + public int ReportCount { get; set; } + } /* MinValue = the first sweep after startup runs the retention purge, then daily. */ private DateTime _nextPurgeUtc = DateTime.MinValue; @@ -1203,7 +1241,7 @@ collection loop stops so both drain cleanly on shutdown. */ /* #2022 phase 2: the Query Store backfill worker, on its own tick (see s_queryStoreBackfillInterval). Fills the two windows the live path discards by design — the 60-minute first-contact tail and - 24h-clamped outage holes — newest-first, byte-budgeted, strictly BELOW the live path's floor, and + clamp-bounded outage holes — newest-first, byte-budgeted, strictly BELOW the live path's floor, and never past the raw tier's horizon. Plan capture reads the same live provider the runner does. */ var queryStoreBackfill = new QueryStoreBackfill(postgres, runner, deltas, _logger, () => config.CapturePlans, () => StoreConfigProvider.ClampTextBudgetMb(config.QueryStoreTextBudgetMb)); @@ -2134,6 +2172,10 @@ losing their text is not worth failing the sweep over. */ /// Reads (queryid, query) from the monitored PostgreSQL server with showtext = true (#2219). /// Capped, and ordered by total execution time so a catalog larger than the cap keeps the text for the /// queries anyone would actually look at rather than an arbitrary slice. + /// #2651: the SOURCE is chosen by flavor. This read was Aurora-only, so off Aurora the text table + /// was never populated at all — which made get_pg_top_queries return a null query_text on every row + /// forever, and made test_hypothetical_index (#2612) unable to resolve a statement on the one platform + /// it can be tested against. /// private static async Task<(List QueryIds, List Texts)> ReadPgStatementTextAsync( ServerRuntime runtime, CancellationToken cancellationToken) @@ -2143,7 +2185,9 @@ losing their text is not worth failing the sweep over. */ await using var connection = new Npgsql.NpgsqlConnection(runtime.ConnectionString); await connection.OpenAsync(cancellationToken); - await using var command = new Npgsql.NpgsqlCommand(PgStatementText.FetchSql, connection) { CommandTimeout = 60 }; + await using var command = new Npgsql.NpgsqlCommand( + PgStatementText.FetchSqlFor(runtime.Target.IsAurora, runtime.Target.PostgresMajorVersion), + connection) { CommandTimeout = 60 }; command.Parameters.AddWithValue(PgStatementTextRowCap); await using var reader = await command.ExecuteReaderAsync(cancellationToken); while (await reader.ReadAsync(cancellationToken)) @@ -3053,23 +3097,61 @@ private async Task RunAnalysisPassAsync( bool notifyFindings, CancellationToken stoppingToken) { - if (!_analysisInFlight.TryAdd(serverId, 0)) + if (!_analysisInFlight.TryAdd(serverId, new AnalysisPassState(DateTime.UtcNow))) { + ReportStuckAnalysis(serverId, displayName); return new AnalysisPassResult(AnalysisPassStatus.Skipped, 0, "analysis is already running for this server"); } + CancellationTokenSource? passCts = null; + var passStarted = false; + try { var analysisService = new DarlingAnalysisService(_postgres!, planFetcher, _logger); - var analyzeTask = analysisService.AnalyzeAsync(serverId, storageName, hoursBack: 4, stoppingToken); + + /* #2430: the TOKEN is the budget now; the Task.Delay below is only this sweep's patience. + Before this, AnalyzeAsync received the STOPPING token and nothing else, so the timeout + abandoned the wait without cancelling any work — and since the marker below is released + only on true completion, a pass that never finished left this server skipped in silence + for the life of the process. + + Arming the CTS before the task exists means Cancel can never race the continuation that + disposes it. There is deliberately no Task.Run here, unlike the Lite twin: DuckDB + implements no async execution, so Lite's pass ran its whole collection phase inline and + the race could not fire at all, while Npgsql is genuinely async and hands this thread + back at the first read. That difference is also why this defect stayed invisible on + Darling — a slow pass never took the sweep down with it, it just went quiet. */ + passCts = CancellationTokenSource.CreateLinkedTokenSource(stoppingToken); + passCts.CancelAfter(s_analysisTimeout); + var cts = passCts; + + /* Both tokens, and they mean different things: the first is what the reads observe, the + second is the only one that still means "the host is stopping". Handing the armed token + to the classifier alone would log every ordinary overrun as "abandoned at shutdown". */ + var analyzeTask = analysisService.AnalyzeAsync( + serverId, storageName, hoursBack: 4, cts.Token, stoppingToken); /* Clear the in-flight marker only when the task truly finishes — not when the timeout below moves us on — so a hung server is not relaunched. */ _ = analyzeTask.ContinueWith( - completed => _analysisInFlight.TryRemove(serverId, out _), + completed => + { + _analysisInFlight.TryRemove(serverId, out _); + cts.Dispose(); + }, TaskScheduler.Default); - var finished = await Task.WhenAny(analyzeTask, Task.Delay(s_analysisTimeout, stoppingToken)); + /* From here the continuation owns the marker and the token source, and neither + ContinueWith nor the call above throws, so there is no window in which the pass exists + with nothing committed to cleaning up after it. */ + passStarted = true; + + /* Wait the budget PLUS the unwind grace, so a pass that honours its cancellation is seen + finishing here rather than racing this sweep's own timer. Losing that race now carries + real information: the pass was asked to stop and did not. */ + var finished = await Task.WhenAny( + analyzeTask, Task.Delay(s_analysisTimeout + s_analysisShutdownGrace, stoppingToken)); if (stoppingToken.IsCancellationRequested) { @@ -3102,7 +3184,7 @@ it would cancel this wait instantly and defeat the grace. */ if (finished != analyzeTask) { _logger.LogWarning( - "[{Server}] Analysis exceeded {Timeout}s — skipped this cycle", + "[{Server}] Analysis exceeded {Timeout}s and has not unwound the cancellation raised at that budget — skipped this cycle. This server stays skipped while the pass is in flight, and is reported again if it stays that way; every other server is unaffected.", displayName, (int)s_analysisTimeout.TotalSeconds); return new AnalysisPassResult(AnalysisPassStatus.TimedOut, 0, $"analysis exceeded {(int)s_analysisTimeout.TotalSeconds}s"); } @@ -3111,6 +3193,31 @@ it would cancel this wait instantly and defeat the grace. */ to the notification channels when delivery is on (Lite's D0 split: production unconditional, delivery gated). */ var findings = await analyzeTask; + + /* A pass that ended early unwound as asked and returned nothing, so there is nothing to + route — say so, rather than letting it read as a clean all-clear, which is what the old + code did for every timed-out pass that came back before the sweep gave up on it. + + READ the pass's own classification rather than re-deriving one here (review, #2430). "No + findings and the budget token has fired" is equally true of a genuine fault that landed + after the budget expired, and calling that a timeout would bury the pass's ERROR under a + Warning saying it merely ran out of time. The pass classified this once and logged the + single line for it, so this adds no second line of its own — it only turns the answer + into the terminal state analyze_now reports. */ + if (analysisService.EndedEarlyAs is AnalysisAbandonKind ending) + { + return ending switch + { + AnalysisAbandonKind.Shutdown => + new AnalysisPassResult(AnalysisPassStatus.Skipped, 0, "service is stopping"), + AnalysisAbandonKind.Timeout => + new AnalysisPassResult(AnalysisPassStatus.TimedOut, 0, + $"analysis exceeded {(int)s_analysisTimeout.TotalSeconds}s"), + _ => new AnalysisPassResult(AnalysisPassStatus.Error, 0, + "analysis failed — the pass logged the fault"), + }; + } + if (notifyFindings) { await notificationService.NotifyAsync(findings); @@ -3141,13 +3248,65 @@ await DarlingObservability.WriteAnalysisStateAsync( catch (Exception ex) { _logger.LogError("[{Server}] Analysis failed: {Message}", displayName, ex.Message); - /* If analyzeTask was never created (e.g. ctor threw), the continuation - never ran — clear the marker defensively. */ - _analysisInFlight.TryRemove(serverId, out _); + + /* If the pass was never launched (e.g. the service ctor threw), no continuation exists to + clear the marker or release the token source — do both here, or this server is skipped + forever, which is the very defect this method is being fixed for. Once the pass IS + running the continuation owns them, and clearing them here would pull the token out + from under a live pass and re-admit that server next cycle on top of it. The old code + cleared unconditionally; it got away with it only because every path that could reach + here after launch had already completed the task. */ + if (!passStarted) + { + _analysisInFlight.TryRemove(serverId, out _); + passCts?.Dispose(); + } + return new AnalysisPassResult(AnalysisPassStatus.Error, 0, ex.Message); } } + /// + /// #2430: reports a server whose analysis pass is still in flight from an earlier cycle. The + /// in-flight guard is deliberately released only on true completion — that is what stops a hung + /// server piling up passes — but it also means a pass that never completes leaves the marker set + /// for the life of the service, and every later cycle skipped that server with nothing said at all. + /// The cancellation now raised at the budget clears the great majority of those; what remains is + /// the pass wedged in a store read that takes no token, and this is what makes THAT visible rather + /// than silent. + /// + /// A permanently-skipped server that is loudly skipped is a far smaller bug than one silently + /// skipped: the first costs findings and says so, the second looks exactly like a server with + /// nothing wrong with it. + /// + private void ReportStuckAnalysis(int serverId, string displayName) + { + if (!_analysisInFlight.TryGetValue(serverId, out var state)) + { + /* Finished between the TryAdd above and this read — it was never stuck. */ + return; + } + + state.SkippedCycles++; + + var inFlightFor = DateTime.UtcNow - state.StartedUtc; + var reportAfter = + (s_analysisTimeout * StuckAnalysisMultiple) * + Math.Pow(2, Math.Min(state.ReportCount, StuckAnalysisMaxBackoffDoublings)); + + if (inFlightFor < reportAfter) + { + return; + } + + state.ReportCount++; + + _logger.LogError( + "[{Server}] Analysis has been in flight for {Minutes:F0} minutes — over {Multiple}x its {Timeout}s budget — and did not stop when cancelled at that budget. {Skipped} analysis cycle(s) have been skipped for this server since, and every later cycle is skipped too until the pass unwinds or the service restarts. Analysis for every other server is unaffected.", + displayName, inFlightFor.TotalMinutes, StuckAnalysisMultiple, + (int)s_analysisTimeout.TotalSeconds, state.SkippedCycles); + } + /// /// The analyze_now command handler (Recommendations "Generate now", control-plane form): forces /// an immediate analysis pass for one monitored server, bypassing its NextAnalysisDue wait, and maps the @@ -3306,6 +3465,9 @@ public Task ExecuteActualPlanAsync(int serverId, ActualPlanReque public Task FetchActiveQueriesLiveAsync(int serverId, CancellationToken cancellationToken) => _worker.RunFetchActiveQueriesLiveAsync(_servers, _runner, serverId, cancellationToken); + + public Task TestHypotheticalIndexAsync(int serverId, HypotheticalIndexRequest request, CancellationToken cancellationToken) + => _worker.RunTestHypotheticalIndexAsync(_servers, serverId, request, cancellationToken); } /// @@ -3356,9 +3518,14 @@ no msdb either. */ { throw; } - catch (SqlException ex) when (ex.Number is 229 or 297 or 300 or 916) + catch (SqlException ex) when (SqlServerPermissionErrors.IsPermissionDenied(ex.Number)) { - /* Expected for read-only monitoring accounts; hit every alert cycle, so Info. The named + /* #2512: routed through the shared set rather than a fourth copy of it — this one had + 916 but neither 262 nor 8189, which is the drift the set exists to end. Widening is safe + in the only direction that matters here: every number in it means the login cannot read + what it asked for, and the response is to return no jobs rather than fail the alert + cycle. A 262 naming msdb is exactly this case and used to fall through to the warning. + Expected for read-only monitoring accounts; hit every alert cycle, so Info. The named remedy is direct table SELECTs, NOT SQLAgentReaderRole: that role gates the sp_help_job* interface only and confers nothing on the base tables this query reads — a #1823 field box had the role and still landed here every cycle. */ @@ -3508,9 +3675,11 @@ error. Re-read from server.Runtime because a preceding RunOneAsync in this loop nulled it on a connection-level failure. */ if (server.Runtime is null || !CollectorCatalog.EngineMatches(name, server.Runtime.Target) - /* And the within-engine gate, PostgreSQL only — same reasoning as the scheduled sweep. */ - || (server.Runtime.Target.Engine == CollectorTargetEngine.PostgreSql - && !CollectorCatalog.AppliesTo(name, server.Runtime.Target))) + /* And the within-engine gate, on EVERY engine — same reasoning as the scheduled sweep, + and extended to SQL Server with it (#2579). This loop runs on every connect and + reconnect, so leaving it PostgreSQL-scoped would keep landing fake SUCCESS rows for + gated-off SQL Server collectors at exactly the moments an operator is watching. */ + || !CollectorCatalog.AppliesTo(name, server.Runtime.Target)) { continue; } @@ -3627,23 +3796,29 @@ engine in the catalog every target would log a fake success per foreign collecto continue; } - /* WITHIN-engine gates get the same treatment, on PostgreSQL targets only. + /* WITHIN-engine gates get the same treatment, on EVERY engine. EngineMatches above drops the wrong DIALECT; it says nothing about a collector that is - right-dialect but inapplicable to this particular target — pg_wait_stats and - pg_statement_stats read Aurora-only functions, so on stock PostgreSQL they dispatched, came - back with 0 rows, and RunOneAsync recorded SUCCESS. Two collectors at a 1-minute cadence is - ~2,880 fake successes a day per server, and the PR promised "a graceful skip with an - explanation" instead. Confirmed on the review's live stock-PostgreSQL run. - - Scoped to PostgreSQL deliberately rather than applied to the composed gate for everyone: on - SQL Server the same zero-row-SUCCESS path covers a long-established handful of Azure-gated - collectors, and silencing those is a change to a shipping SKU's log semantics that deserves - its own decision rather than riding along here. + right-dialect but inapplicable to this particular target — pg_wait_stats reads Aurora's own + wait instrumentation, so on stock PostgreSQL it dispatched, came back with 0 rows, and + RunOneAsync recorded SUCCESS. At a 1-minute cadence that is ~1,440 fake successes a day per + server, and the PR promised "a graceful skip with an explanation" instead. (pg_statement_stats + was the second such collector until #2625 gave it a vanilla pg_stat_statements path; it now + applies to every PostgreSQL target and reaches this gate on none of them.) + + #2579 EXTENDS THIS TO SQL SERVER, which the PostgreSQL change deliberately left alone as + "its own decision" because it changes a shipping SKU's log semantics for the Azure-gated + collectors. Here is that decision, and what settled it is that the cost turned out to be + the opposite of cosmetic. On an AWS RDS fleet the SQL Server gates are not a handful: 84 + instances x agent_status and running_jobs x a 5-minute cadence is ~24,000 rows a day that + say SUCCESS about collectors deliberately not running. A gated-off run recorded as SUCCESS + is byte-identical to a real one — same status, zero rows, no note — so nothing downstream + can tell them apart. That is not merely noise: it is the shape the whole miss vocabulary + exists to prevent, and it read as evidence of working collection convincingly enough to + produce an issue and a PR built on it before the 0ms durations gave it away. No log row is the honest outcome, and it is not silent: --test-connection names exactly which collectors do not apply to a target, and why, before the service ever runs. */ - if (runtime.Target.Engine == CollectorTargetEngine.PostgreSql - && !CollectorCatalog.AppliesTo(name, runtime.Target)) + if (!CollectorCatalog.AppliesTo(name, runtime.Target)) { continue; } @@ -3726,11 +3901,12 @@ enabled one NOW regardless of frequency or NextDue — that is what "snapshot" m /* The THIRD dispatch loop, and it got neither engine gate in the first round: an operator snapshot against a PostgreSQL target dispatched every SQL Server collector, whose AppliesTo early-return yields zero rows and lands a burst of fake SUCCESS in - collection_log — the phantom-success class the other two loops (on-load :3157, scheduled - sweep) were gated against. Same predicate, same PostgreSQL-only scoping. */ + collection_log — the phantom-success class the other two loops (on-load, scheduled + sweep) were gated against. Same predicate, and since #2579 the same every-engine scoping: + an operator-triggered snapshot against an RDS target would otherwise land its own burst of + fake successes for the msdb-gated collectors. */ if (!CollectorCatalog.EngineMatches(name, runtime.Target) - || (runtime.Target.Engine == CollectorTargetEngine.PostgreSql - && !CollectorCatalog.AppliesTo(name, runtime.Target))) + || !CollectorCatalog.AppliesTo(name, runtime.Target)) { continue; } @@ -3798,8 +3974,11 @@ private async Task RunFetchPlanAsync( JsonError($"server '{displayName}' is not currently connected — the live plan cache can only be read from a connected server")); } + /* #2443: both arms now take the SAME token. The by-sql_handle arm always had it; the + by-plan_handle arm did not, so a cancelled fetch_plan command kept a session open on the + monitored server for whichever key the caller happened to use. */ var planXml = request.UsePlanHandle - ? await planFetcher.FetchPlanXmlAsync(serverId, request.PlanHandle!) + ? await planFetcher.FetchPlanXmlAsync(serverId, request.PlanHandle!, cancellationToken) : await planFetcher.FetchPlanBySqlHandleAsync( serverId, request.DatabaseName, request.SqlHandle!, request.StatementStartOffset, request.StatementEndOffset, cancellationToken); @@ -3825,6 +4004,114 @@ from a failure. planXml is null so the viewer's parse hits the not-in-cache bran /// poll miss, and so a wedged read cannot pin the single-threaded command loop past the stale-command reaper. public const int ActiveQueriesFetchTimeoutSeconds = 30; + /// + /// test_hypothetical_index (#2612): plan one stored statement with and without a candidate index. + /// + /// The statement text is resolved HERE, from this product's own pg_statement_text store, + /// keyed by the queryid the caller named. It is never taken from the caller — the request carries an + /// identifier and a candidate, and nothing else reaches SQL. + /// + /// PostgreSQL only, and it says so rather than failing obscurely on a SQL Server target: hypopg + /// and EXPLAIN (GENERIC_PLAN) have no SQL Server equivalent, and the candidate this answers about + /// comes from a PostgreSQL-only collector. + /// + private async Task RunTestHypotheticalIndexAsync( + List servers, int serverId, HypotheticalIndexRequest request, CancellationToken cancellationToken) + { + ServerLoopState? server; + ServerRuntime? runtime; + string displayName; + lock (_serversLock) + { + server = servers.Find(s => s.Config.ServerId == serverId); + runtime = server?.Runtime; + displayName = server?.Config.DisplayName ?? serverId.ToString(CultureInfo.InvariantCulture); + } + + if (server is null) + { + return new CommandOutcome(false, "server not monitored", JsonError($"no monitored server with server_id {serverId}")); + } + + if (runtime is null) + { + return new CommandOutcome(false, "server not connected", + JsonError($"server '{displayName}' is not currently connected — a hypothetical index has to be tested against the server's own statistics, so there is nothing to answer from while it is unreachable")); + } + + if (runtime.Target.Engine != CollectorTargetEngine.PostgreSql) + { + return new CommandOutcome(false, "not a PostgreSQL target", + JsonError($"server '{displayName}' is not PostgreSQL. Hypothetical indexes come from the hypopg extension and the plan comparison needs EXPLAIN (GENERIC_PLAN); neither has a SQL Server equivalent, and the index candidate this answers about comes from a PostgreSQL-only collector.")); + } + + if (!request.TryGetQueryId(out var queryId)) + { + return new CommandOutcome(false, "invalid queryid", JsonError("queryid must be a signed 64-bit integer sent as a STRING")); + } + + string? statementText; + await using (var lookup = _postgres!.CreateCommand( + "SELECT query_text FROM collect.pg_statement_text WHERE server_id = $1 AND queryid = $2")) + { + lookup.Parameters.AddWithValue(serverId); + lookup.Parameters.AddWithValue(queryId); + statementText = (await lookup.ExecuteScalarAsync(cancellationToken)) as string; + } + + if (string.IsNullOrWhiteSpace(statementText)) + { + return new CommandOutcome(false, "statement text not captured", + JsonError($"no statement text is stored for queryid {queryId} on '{displayName}'. Text is refreshed on its own cadence and a major-version upgrade re-keys every queryid, so a statement first seen minutes ago genuinely has none yet — there is nothing to re-plan until it does.")); + } + + try + { + await using var connection = new NpgsqlConnection(runtime.ConnectionString); + await connection.OpenAsync(cancellationToken); + + var result = await Targets.HypotheticalIndexExperiment.RunAsync( + connection, statementText, request.BuildCreateIndexStatement(), + runtime.Target.PostgresMajorVersion, cancellationToken); + + _logger.LogInformation( + "[{Server}] test_hypothetical_index on {Schema}.{Table}: planner would {Verdict} it ({Before:N2} -> {After:N2})", + displayName, request.SchemaName, request.TableName, + result.PlannerWouldUseIt ? "USE" : "NOT use", result.CostBefore, result.CostAfter); + + return new CommandOutcome(true, "hypothetical index tested", JsonSerializer.Serialize(new + { + server = displayName, + queryid = request.QueryId, + candidate = request.BuildCreateIndexStatement(), + planner_would_use_it = result.PlannerWouldUseIt, + estimated_cost_before = result.CostBefore, + estimated_cost_after = result.CostAfter, + hypothetical_index_name = result.HypotheticalIndexName, + explanation = result.Explanation, + plan_before = result.PlanBeforeJson, + plan_after = result.PlanAfterJson, + })); + } + catch (OperationCanceledException) + { + throw; + } + catch (PostgresException ex) + { + /* Reported rather than thrown, and the SQLSTATE travels: 42P01 here means the table named in the + candidate does not exist, which is a caller mistake, and 42501 means the login cannot plan + against it — two different conversations that a bare failure would merge. */ + return new CommandOutcome(false, "experiment failed", + JsonError($"planning on '{displayName}' failed: {ex.MessageText} (SQLSTATE {ex.SqlState})")); + } + catch (Exception ex) + { + return new CommandOutcome(false, "experiment failed", + JsonError($"planning on '{displayName}' failed: {ex.Message}")); + } + } + /// /// The fetch_active_queries command handler (headless-plan live-snapshot wave): reads the LIVE /// running-request DMV snapshot from one monitored server on demand and returns the rows — the on-demand, @@ -4160,6 +4447,16 @@ private static void BindActualPlanResolveParameters(NpgsqlCommand command, int s _ => "row", }; + /// + /// Where the missing extension has to be created, named when we know it (#2638). + /// + private static string WhereToCreateIt(string? connectedDatabase) + => string.IsNullOrWhiteSpace(connectedDatabase) + ? "in the connected database (CREATE EXTENSION ...). " + : $"in database '{connectedDatabase}', which is the one this collector connects to — an " + + "extension installed in a DIFFERENT database on the same cluster is invisible from here, " + + $"so run CREATE EXTENSION in '{connectedDatabase}'. "; + /// /// Maps a PostgreSQL fault to a collection_log status plus the sentence an operator needs. /// The store has five statuses and none of them is "this feature is not installed", so the @@ -4168,7 +4465,7 @@ private static void BindActualPlanResolveParameters(NpgsqlCommand command, int s /// general handler have it", which keeps the genuinely unexpected loud. /// internal static (string Status, string Explanation) PostgresFaultOutcome( - PostgresException ex, string collectorName) + PostgresException ex, string collectorName, string? connectedDatabase = null) { var fault = PostgresTargetProvider.Instance.Classify( ex, CollectorCatalog.YieldsOnLockTimeout(collectorName)); @@ -4181,12 +4478,18 @@ internal static (string Status, string Explanation) PostgresFaultOutcome( /* 42P01 / 42883: the relation or function is not there. Overwhelmingly an extension that was never created in the connected database rather than anything to do with privileges. */ + /* #2638: the database is NAMED. Extensions are per-database, and on a real fleet one was + installed in a different database on the same cluster from the one this collector connects + to — so an operator who checked the obvious database found it already there and concluded + the collector was broken. "Create it somewhere" is not an instruction; naming the database + makes it one. Threaded from ServerRuntime.ConnectedDatabase; when that is unknown the + sentence degrades to what it always said rather than inventing a name. */ CollectorTargetFault.ObjectMissing => ("PERMISSIONS", $"{ex.MessageText} (SQLSTATE {ex.SqlState}) — the source object does not exist on this " - + "target. This is NOT a missing grant: it is normally an extension that was never " - + "created in the connected database (CREATE EXTENSION pg_stat_statements), so the " - + "collector will keep degrading until it is. Recorded as a non-fatal skip rather than an " - + "error so it does not fill the log every cycle."), + + "target. This is NOT a missing grant: it is normally an extension that was never created " + + WhereToCreateIt(connectedDatabase) + + "The collector will keep degrading until it is. Recorded as a non-fatal skip rather than " + + "an error so it does not fill the log every cycle."), /* 0A000 / 55000 / 55006: the server will not do this, permanently or by configuration — pg_stat_wal on Aurora, or an optimized-reads cache that is switched off. */ @@ -4276,7 +4579,8 @@ they used to be free to run against the SAME server at once — measured as ~128 databases, #1837). The status stays SUCCESS, and every health/band read keys on status rather than on error_message, so the note is inert outside the Collection Log detail grid. */ await DarlingObservability.LogCollectionAsync( - _postgres!, runtime, collectorName, "SUCCESS", result.Rows, result.SqlMs, result.StorageMs, result.Note, _logger, cancellationToken); + _postgres!, runtime, collectorName, "SUCCESS", result.Rows, result.SqlMs, result.StorageMs, result.Note, + result.Fanout, _logger, cancellationToken); /* #2219: statement TEXT rides alongside the statement stats, on its own hourly cadence. Hung off the stats collector's success rather than given its own loop because it is meaningless without those @@ -4307,7 +4611,31 @@ hits the permissions filter. */ server.Config.DisplayName, collectorName, ex.Message); await DarlingObservability.LogCollectionAsync( - _postgres!, runtime, collectorName, "SESSION_MISSING", 0, 0, 0, ex.Message, _logger, cancellationToken); + _postgres!, runtime, collectorName, "SESSION_MISSING", 0, 0, 0, ex.Message, fanout: null, _logger, cancellationToken); + return 0; + } + catch (RdsLogUnavailableException ex) when (ex.IsAuthorizationFailure) + { + /* #2633: the AWS call was DENIED, so nothing was read. Degraded to PERMISSIONS rather than + ERROR for the same reason a 42501 from the pg_read_file route is — a least-privilege + deployment is an expected state an operator can act on, and screaming every cycle about it + would bury real faults — but it must NOT be recorded as a successful empty read, which is + what returning zero rows used to make it. + + Only the authorization case lands here. A throttle, a failover or an endpoint that stopped + resolving falls through to the general handler and stays loud, because a permanent-sounding + status on a transient fault is how an outage gets read as a configuration choice. */ + _logger.LogWarning(" [{Server}] {Collector} => PERMISSIONS: the RDS log API refused the call", + server.Config.DisplayName, collectorName); + + await DarlingObservability.LogCollectionAsync( + _postgres!, runtime, collectorName, "PERMISSIONS", 0, 0, 0, + $"{ex.Message} — the MONITORING HOST's IAM role lacks a grant this source needs, which is " + + "not a database grant: plan capture on managed PostgreSQL reads the server log through " + + "the RDS API, so the role needs rds:DescribeDBLogFiles and rds:DownloadDBLogFilePortion " + + "on the target instance. Nothing was read this cycle — this is NOT 'no plans were " + + "captured'.", + fanout: null, _logger, cancellationToken); return 0; } catch (SqlException ex) when (ex.Number == 1222 && CollectorCatalog.YieldsOnLockTimeout(collectorName)) @@ -4328,10 +4656,10 @@ general Exception catch. */ await DarlingObservability.LogCollectionAsync( _postgres!, runtime, collectorName, "YIELDED", 0, 0, 0, $"Lock-timeout yield (SQL error #{ex.Number}): the 1-second LOCK_TIMEOUT guard fired rather than waiting in a blocking chain. One snapshot sweep skipped; evidence of lock contention on the monitored server, not a monitoring failure.", - _logger, cancellationToken); + fanout: null, _logger, cancellationToken); return 0; } - catch (SqlException ex) when (ex.Number is 229 or 297 or 300 or 8189) + catch (SqlException ex) when (SqlServerPermissionErrors.IsPermissionDenied(ex.Number)) { /* Same Azure explanation Lite appends (#1631): error 300 on Azure SQL Database is a service objective limit phrased as a permission denied on 'master', which reads as a missing GRANT @@ -4339,18 +4667,23 @@ await DarlingObservability.LogCollectionAsync( searchable. Parity is the point — a Darling operator gets the identical sentence Lite gives. 8189 is sys.traces' own denial ("You do not have permission to run 'SYS.TRACES'", ALTER TRACE missing): a legitimate least-privilege choice (#1823) — ALTER TRACE is not read-only — - so default_trace_events must degrade as PERMISSIONS, not scream ERROR every cycle. */ - var message = ex.Message + AzureDmvPermissionHint.For(ex.Number, server.Runtime?.Target.IsAzureSqlDb == true); + so default_trace_events must degrade as PERMISSIONS, not scream ERROR every cycle. + #2512: the number set moved to SqlServerPermissionErrors, shared with Lite's catch and + with SqlServerTargetProvider.Classify, and gained 262 — "permission denied in database + 'tempdb'", the #2150 denial that used to record ERROR every cycle and is the reason + tempdb_stats was gated off Azure SQL Database at all. */ + var message = ex.Message + AzureDmvPermissionHint.For( + ex.Number, server.Runtime?.Target.IsAzureSqlDb == true, ex.Message); _logger.LogWarning(" [{Server}] {Collector} => insufficient permissions ({Number}): {Message}", server.Config.DisplayName, collectorName, ex.Number, message); await DarlingObservability.LogCollectionAsync( - _postgres!, runtime, collectorName, "PERMISSIONS", 0, 0, 0, message, _logger, cancellationToken); + _postgres!, runtime, collectorName, "PERMISSIONS", 0, 0, 0, message, fanout: null, _logger, cancellationToken); return 0; } catch (PostgresException ex) when ( - PostgresFaultOutcome(ex, collectorName) is { Status: not "ERROR" } outcome) + PostgresFaultOutcome(ex, collectorName, runtime.ConnectedDatabase) is { Status: not "ERROR" } outcome) { /* PostgreSQL faults classified by SQLSTATE through the same ITargetProvider.Classify the engine seam already exposes, so the runner and the provider cannot disagree about what an @@ -4381,7 +4714,7 @@ The message says WHICH kind it is rather than leaving "PERMISSIONS" to imply a m } await DarlingObservability.LogCollectionAsync( - _postgres!, runtime, collectorName, status, 0, 0, 0, explanation, _logger, cancellationToken); + _postgres!, runtime, collectorName, status, 0, 0, 0, explanation, fanout: null, _logger, cancellationToken); return 0; } catch (Exception ex) @@ -4419,7 +4752,7 @@ dropping the connection over one would turn a tuning problem into a reconnect st try { await DarlingObservability.LogCollectionAsync( - _postgres!, runtime, collectorName, "ERROR", 0, 0, 0, ex.Message, _logger, cancellationToken); + _postgres!, runtime, collectorName, "ERROR", 0, 0, 0, ex.Message, fanout: null, _logger, cancellationToken); } catch { @@ -4490,11 +4823,37 @@ await DarlingObservability.LogCollectionAsync( ["pg_wait_stats"] = (r, s, ct) => r.RunAsync(PgWaitStatsCollector.Instance, s, ct), ["pg_statement_stats"] = (r, s, ct) => r.RunAsync(PgStatementStatsCollector.Instance, s, ct), ["pg_wraparound_stats"] = (r, s, ct) => r.RunAsync(PgWraparoundStatsCollector.Instance, s, ct), + ["pg_server_config"] = (r, s, ct) => r.RunAsync(PgServerConfigCollector.Instance, s, ct), + ["pg_deadlocks"] = (r, s, ct) => r.RunAsync(PgDeadlocksCollector.Instance, s, ct), ["pg_xmin_horizon"] = (r, s, ct) => r.RunAsync(PgXminHorizonCollector.Instance, s, ct), ["pg_replication_slots"] = (r, s, ct) => r.RunAsync(PgReplicationSlotsCollector.Instance, s, ct), ["pg_autovacuum_stats"] = (r, s, ct) => r.RunAsync(PgAutovacuumStatsCollector.Instance, s, ct), ["pg_io_stats"] = (r, s, ct) => r.RunAsync(PgIoStatsCollector.Instance, s, ct), ["pg_blocking"] = (r, s, ct) => r.RunAsync(PgBlockingCollector.Instance, s, ct), + ["pg_database_stats"] = (r, s, ct) => r.RunAsync(PgDatabaseStatsCollector.Instance, s, ct), + ["pg_index_usage_stats"] = (r, s, ct) => r.RunAsync(PgIndexUsageStatsCollector.Instance, s, ct), + ["pg_table_bloat_stats"] = (r, s, ct) => r.RunAsync(PgTableBloatStatsCollector.Instance, s, ct), + ["pg_session_states"] = (r, s, ct) => r.RunAsync(PgSessionStatesCollector.Instance, s, ct), + ["pg_plan_capture_readiness"] = (r, s, ct) => r.RunAsync(PgPlanCaptureReadinessCollector.Instance, s, ct), + ["pg_write_stats"] = (r, s, ct) => r.RunAsync(PgWriteStatsCollector.Instance, s, ct), + ["pg_extension_availability"] = (r, s, ct) => r.RunAsync(PgExtensionAvailabilityCollector.Instance, s, ct), + ["pg_lock_stats"] = (r, s, ct) => r.RunAsync(PgLockStatsCollector.Instance, s, ct), + ["pg_wait_sampling"] = (r, s, ct) => r.RunAsync(PgWaitSamplingCollector.Instance, s, ct), + ["pg_kernel_stats"] = (r, s, ct) => r.RunAsync(PgKernelStatsCollector.Instance, s, ct), + ["pg_predicate_stats"] = (r, s, ct) => r.RunAsync(PgPredicateStatsCollector.Instance, s, ct), + /* TWO TRANSPORTS, one table. Self-hosted reads the server log with pg_read_file; Aurora and RDS + have no filesystem and pg_read_server_files is not grantable, so those go through the AWS log + API instead (#2538). The collector's own AppliesTo excludes managed targets, so without this + branch they would simply never capture a plan - and would look like they had nothing to say + rather than like they were on a different road. */ + ["pg_plan_capture"] = (r, s, ct) => + s.Target.IsAurora || s.Target.IsAwsRds + ? r.IngestRdsPlansAsync(s, ct) + : r.RunAsync(PgPlanCaptureCollector.Instance, s, ct), + ["pg_column_stats"] = (r, s, ct) => r.RunAsync(PgColumnStatsCollector.Instance, s, ct), + ["pg_replication_stats"] = (r, s, ct) => r.RunAsync(PgReplicationStatsCollector.Instance, s, ct), + ["pg_buffer_usage"] = (r, s, ct) => r.RunAsync(PgBufferUsageCollector.Instance, s, ct), + ["pg_index_bloat"] = (r, s, ct) => r.RunAsync(PgIndexBloatCollector.Instance, s, ct), }; /// diff --git a/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHostBinding.cs b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHostBinding.cs index 62fa0db9b..0e6b6b517 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHostBinding.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHostBinding.cs @@ -7,6 +7,7 @@ */ using System; +using System.Collections.Generic; using System.Net; using System.Runtime.Versioning; using System.Security.Cryptography; @@ -230,4 +231,248 @@ internal static bool FixedTimeTokenEquals(string? presented, string? expected) var presentedHash = SHA256.HashData(Encoding.UTF8.GetBytes(presented)); return CryptographicOperations.FixedTimeEquals(expectedHash, presentedHash); } + + /* --------------------------------------------------------------------------------------------------- + WHICH PLANE decided this endpoint is on, and on what port (#2389). One mcp object in darling.json + has TWO owners: enabled/port are a first-run SEED that the store's config_service row + overrides forever after, while network.* (listen / allowFrom / the DPAPI token) is file-only and + restart-only. Nothing in the file distinguishes them, so an operator who edits mcp.enabled, restarts and + greps the log finds "Starting MCP server on http://:5152" -- true, and then contradicted five + seconds later when the worker's first publish arrives and the supervisor reconciles to the store. + + network.* deliberately STAYS file-only rather than being moved into the store to match: the DPAPI token + is LocalMachine-scoped, so a blob in config_service is undecryptable on any other host that reads that + store, and a bind address + CIDR + token living in config_service would let a REMOTE admin store + connection -- the pivot the postgres.network role warning already names -- re-point this listener onto a + LAN interface behind a credential of its own choosing. Exposure should require touching the host. So the + split is kept and made VISIBLE instead: the effective values carry their origin into the start line, and + a disagreement is reported at the point of override rather than left to be inferred from two INFO lines. + --------------------------------------------------------------------------------------------------- */ + + /// Which plane supplied the effective enable/port a host supervisor is acting on (#2389). + internal enum EndpointToggleOrigin + { + /// darling.json, because the worker has not published yet (still bootstrapping, or it never + /// reached the store). PROVISIONAL: the control plane can contradict it within one poll interval. + File, + + /// The store's config.config_service row, published by the worker. Authoritative. + ControlPlane, + } + + /// The effective (enabled, port) a supervisor acts on, WITH its provenance (#2389). + /// / are true only when the control plane + /// supplied a value that DIFFERS from darling.json's -- the reportable disagreement. + internal readonly record struct EndpointToggle( + bool Enabled, int Port, EndpointToggleOrigin Origin, bool EnabledOverridden, bool PortOverridden); + + /// + /// PURE: the store row wins whenever the worker has published one -- byte-for-byte the old + /// published?.Enabled ?? config.Mcp.Enabled pair -- but it also reports WHERE the values came from + /// and whether they contradict the file, which a null-coalesce structurally cannot. + /// + internal static EndpointToggle ResolveEndpointToggle((bool Enabled, int Port)? published, bool fileEnabled, int filePort) + => published is null + ? new EndpointToggle(fileEnabled, filePort, EndpointToggleOrigin.File, false, false) + : new EndpointToggle( + published.Value.Enabled, + published.Value.Port, + EndpointToggleOrigin.ControlPlane, + EnabledOverridden: published.Value.Enabled != fileEnabled, + PortOverridden: published.Value.Port != filePort); + + /// + /// PURE: the provenance clause the start line carries, so "Starting ... on http://..." says on whose + /// authority it is starting -- and admits when it is running on file values the control plane has not + /// weighed in on yet, which is the line the operator greps and stops reading. + /// + internal static string DescribeToggleOrigin(EndpointToggle toggle) + => toggle.Origin == EndpointToggleOrigin.ControlPlane + ? "the control plane (config.config_service)" + : "darling.json (PROVISIONAL - the control plane has not published yet and may stop or rebind this server)"; + + /// + /// PURE: the one-line report of a control-plane override, or null when the two planes agree (or nothing is + /// published yet, in which case there is nothing to disagree with). is the + /// darling.json object name, the config_service column prefix, and the CLI verb suffix at once ("mcp" -> + /// mcp.enabled / config_service.mcp_enabled / --enable-mcp), which is what keeps the MCP and web wordings + /// from drifting apart. The file values are passed in rather than re-derived so the message quotes what the + /// caller actually loaded. + /// + internal static string? DescribeToggleOverride( + EndpointToggle toggle, string section, string surface, bool fileEnabled, int filePort) + { + if (toggle.Origin != EndpointToggleOrigin.ControlPlane || (!toggle.EnabledOverridden && !toggle.PortOverridden)) + { + return null; + } + + var fields = new List(2); + if (toggle.EnabledOverridden) + { + fields.Add( + $"enabled is {(fileEnabled ? "true" : "false")} in darling.json ({section}.enabled) but " + + $"{(toggle.Enabled ? "true" : "false")} in config.config_service.{section}_enabled"); + } + + if (toggle.PortOverridden) + { + fields.Add($"port is {filePort} in darling.json ({section}.port) but {toggle.Port} in config.config_service.{section}_port"); + } + + /* The ownership disclosure is unconditional because it is true in every state; the CONSEQUENCE is not. + Review catch: a single closing sentence claiming "the control plane keeps this endpoint off" is wrong + whenever the control plane turned it ON despite the file, and wrong for a port-only mismatch where + both planes agree it runs -- which is the likelier real case. So the state-specific half is emitted + only in the state it describes. */ + var consequence = toggle.Enabled + ? " It is what this RUNNING endpoint is bound and gated by, and no store setting can move it." + : " It is not what is keeping this endpoint down, and it takes effect as written the moment the " + + "control plane enables it."; + + return $"{surface} configuration disagrees across the two planes and the CONTROL PLANE WINS: " + + string.Join("; ", fields) + + $". After the first run darling.json's {section}.enabled/{section}.port are only the SEED -- change them with " + + $"--enable-{section}/--disable-{section} or the Viewer's Settings, or the file values will keep being ignored. " + + $"The {section}.network block is the OPPOSITE: file-only, restart-only, no store equivalent -- the control " + + "plane cannot change where this endpoint binds or what token it requires." + + consequence; + } + + /* --------------------------------------------------------------------------------------------------- + #2414: the same split, one layer down -- which port a FIREWALL rule is named for. + + A scoped rule carries its port inside its DisplayName ("PerformanceMonitor Darling MCP (port 5152)"), + so the port is not a parameter of the rule, it IS the rule's identity. The endpoint binds the CONTROL + PLANE's port; every firewall surface used to derive both the name and the -LocalPort from darling.json. + Those two agree right up until somebody moves the port in the Viewer's Settings, and then the elevated + verb opens a rule for a port nothing serves while leaving the served port shut -- an unreachable + endpoint AND an inbound allow rule with no listener behind it, which is the exact inverse of what + scoping the rule to a port was for. + + So a firewall verb resolves its port through ResolveEndpointToggle, the same call the two supervisors + bind on, and then SAYS which plane answered. Saying it is not decoration: a verb that opens a port and + guesses at which one silently is precisely how this defect stayed invisible for as long as it did, and + an operator who has to be told to re-run something needs to know what the last run actually did. The + describer is pure so the wording pins in a test rather than being discovered in the field. + --------------------------------------------------------------------------------------------------- */ + + /// + /// PURE: what a firewall verb discloses about the port it just scoped a rule to (#2414). Three states, + /// because they call for three different things from the reader: + /// + /// the control plane answered and DISAGREES with darling.json -- the defect's own precondition, so + /// it names both ports and says what naming the rule for the file's one would have cost; + /// the control plane answered and agrees -- a one-line confirmation, so an operator can see the + /// authoritative read happened rather than having to infer it from silence; + /// the control plane could NOT be read -- the file's seed was used, and this says so, why, and what + /// makes it wrong, because a firewall verb that falls back without saying so re-creates the bug quietly. + /// + /// is the darling.json object name, the config_service column prefix and the + /// CLI verb suffix at once ("mcp" -> mcp.port / config_service.mcp_port / --enable-mcp), which is what + /// keeps the MCP and web wordings from drifting apart -- the same seam + /// uses. is passed in rather than re-derived so the message quotes the value + /// the caller actually loaded. + /// + internal static string DescribeFirewallPortAuthority( + EndpointToggle toggle, string section, string surface, int filePort, string? storeUnavailableReason) + { + if (toggle.Origin == EndpointToggleOrigin.ControlPlane) + { + return toggle.PortOverridden + ? $"Firewall: scoping the {surface} rule to port {toggle.Port} -- the CONTROL PLANE's port " + + $"(config.config_service.{section}_port), which is the port this endpoint actually binds. " + + $"darling.json says {section}.port = {filePort}, but after the first run that is only the seed: " + + $"a rule named for {filePort} would leave the served port closed to the LAN and hold an inbound " + + "allow rule open on a port nothing is listening on." + : $"Firewall: scoping the {surface} rule to port {toggle.Port}, confirmed against the control plane " + + $"(config.config_service.{section}_port) -- darling.json's {section}.port agrees."; + } + + return $"Firewall: could NOT read the control plane ({storeUnavailableReason ?? "reason unknown"}), so the " + + $"{surface} rule is scoped to darling.json's {section}.port = {filePort}. That value is the FIRST-RUN " + + "SEED: it is the right port on a box whose store has never been written -- the normal state at install " + + "time, which is why this verb uses it rather than refusing -- and it is the WRONG port the moment the " + + $"{surface} port has been changed in the Viewer's Settings or with --enable-{section}. If the endpoint " + + "is unreachable from the LAN after this, re-run --configure-firewall once the store is up: it resolves " + + "the effective port and moves the rule."; + } + + /* --------------------------------------------------------------------------------------------------- + WHERE the network block came from, and when it can change (#2479, item 6). The other half of the same + split #2389 made visible. A network block is loaded ONCE, at start-up, and held for the process + lifetime by design: exposure should require touching the host, and the DPAPI token is LocalMachine- + scoped so it could not live in the store anyway. That is the right design and it is not going to + change -- but it produces the "I enabled it and it is still loopback-bound" trap in another costume, + because nothing said the value was frozen the moment the process started. + + #2389/#2411 fixed the seed-versus-store confusion by making the split VISIBLE rather than by removing + it, and #2414 did the same one layer down for the firewall port. This is that treatment applied to + the third face of the same object: the endpoint states plainly, at every start, whether a network + block was read from the file, what it decided, and that editing darling.json now changes nothing + until a restart. Said in BOTH modes on purpose -- the loopback line was the one that carried no + mention of the block at all, and loopback-when-you-expected-LAN is precisely the state an operator + is trying to diagnose. + --------------------------------------------------------------------------------------------------- */ + + /// + /// PURE: the one-line start-up statement of where {section}.network came from and when it can + /// change. is the darling.json object name and the CLI verb suffix at once + /// ("mcp" -> mcp.network / --enable-mcp), which is what keeps the two hosts' wordings from drifting. + /// + /// Three states, three different consequences, and only the true one is emitted. A single line + /// claiming "restart to apply" would be wrong for an endpoint already serving the LAN exactly as + /// configured, which is the likeliest state and the one where a spurious call to action costs the most + /// credibility. + /// + /// "mcp" or "web" - the darling.json object name. + /// "MCP" or "Web dashboard" - how the log names this endpoint elsewhere. + /// A {section}.network block was present in the file. + /// This endpoint actually bound a LAN address as a result. + /// The bound address, when exposed. + /// The source CIDR, when exposed. + internal static string DescribeNetworkBlockLifetime( + string section, + string surface, + bool blockConfigured, + bool exposed, + string? listen, + string? allowFrom) + { + /* Common to all three: what CAN be changed without a restart, so naming what cannot does not read + as "this endpoint is immovable". The control plane really does stop and rebind this server + live -- it just cannot touch these three fields. */ + var storeHalf = + $" The control plane still owns whether this endpoint runs and on which port " + + $"(config.config_service.{section}_enabled / {section}_port, moved by --enable-{section} / " + + $"--disable-{section} or the Viewer's Settings, and applied without a restart) - it simply " + + "cannot move where this binds, who may reach it, or what token it requires."; + + if (!blockConfigured) + { + return $"{surface} network exposure: darling.json has no {section}.network block, so this endpoint is " + + "loopback-only. Adding one is a FILE edit that takes effect on the next service RESTART; there " + + "is no store setting and no Viewer control that can expose it." + + storeHalf; + } + + if (exposed) + { + return $"{surface} network exposure: {section}.network was read from darling.json at start-up and is " + + $"FIXED for the lifetime of this process - LAN-bound on {listen} for {allowFrom}, behind the " + + "configured token, until the service is RESTARTED. Editing darling.json now changes nothing " + + "until then." + + storeHalf; + } + + /* The trap, named. An operator who edited the block and restarted nothing sees this and has their + answer; one who DID restart is pointed at the fail-closed reason already logged above rather than + being told the same thing twice in different words. */ + return $"{surface} network exposure: darling.json DEFINES {section}.network, but this endpoint is bound " + + "loopback-only. That block was read at start-up and is FIXED for the lifetime of this process, so " + + "editing darling.json now changes nothing until the service is RESTARTED. If you just edited it to " + + $"expose this endpoint, restart the service; if you already restarted, the {section}.network line " + + "logged above says why the exposure was refused." + + storeHalf; + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHttpRefusalLog.cs b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHttpRefusalLog.cs new file mode 100644 index 000000000..a171a154b --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingHttpRefusalLog.cs @@ -0,0 +1,355 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Net; +using System.Text; +using Microsoft.Extensions.Logging; + +namespace PerformanceMonitor.Darling.Service.Hosting; + +/// Which gate turned a request away. Naming it is the whole point (#2479, item 5). +internal enum DarlingRefusalGate +{ + /// The Host-header allowlist / DNS-rebinding guard. 400, both bind modes. + HostAllowlist, + + /// The access token — the MCP bearer header, or the web dashboard's ?token=. + Token, + + /// The in-app source-address allowlist built from network.allowFrom. 403. + SourceCidr, +} + +/// +/// One rate-limited WARN per refused request, saying which gate refused it (#2479, item 5). +/// +/// The gap. A 401, 403 or 400 on either listener left NOTHING in the service log — every +/// refusal site was a bare StatusCode = …; return;. So "is my token wrong, or my CIDR wrong?" was +/// answerable only from the client, which sees one opaque status code and cannot tell a Host-allowlist 400 +/// from a malformed request, or a CIDR 403 from anything else. The person who can FIX either of those is +/// on the server, and the server said nothing. +/// +/// Why the rate limit is not optional. These listeners are LAN-exposed on purpose. An exposed +/// port meets a scanner eventually, and an unthrottled line per refusal turns that into a log-filling +/// denial of service against the very file an operator debugs from — the log stops being readable exactly +/// when it matters. So a refusal is logged at most once per per +/// (gate, source address), and the number suppressed since the last line rides along on the next one, so +/// the volume is visible without being written. +/// +/// Why per-SOURCE and not per-gate alone. Per-gate is a smaller key set and it is wrong: one +/// scanner's line would win the window and hide the tester's own refusal completely, which is the failure +/// this exists to fix, reintroduced by the fix for it. Keyed per source, the tester's FIRST refusal is +/// always logged immediately — first sighting never waits — and only their repeats are folded, which cost +/// nothing to lose because a second identical refusal carries no information the first did not. +/// +/// Why there is a cap, and what it costs. A per-source map is unbounded by construction on an +/// exposed port. Past distinct sources in a window, further sources +/// fold into one per-gate aggregate bucket that says so, which bounds both memory and log volume to +/// (cap + gate count) entries and one line per gate per window. The accepted cost: under an active scan +/// the per-source slots may be taken, and a tester arriving mid-scan lands in the aggregate. That is +/// visible rather than silent — the line says the per-source budget is full — and if the port is being +/// scanned, that IS the more urgent thing for the operator to know. +/// +/// What is never logged. The token, any prefix of it, and its length. A refusal reports +/// whether a credential was PRESENTED, never anything about its value. Attacker-controlled text that does +/// reach the log (the Host header) goes through first: a log file is a text file, +/// and a header carrying CR/LF would otherwise forge log lines. +/// +internal sealed class DarlingHttpRefusalLog +{ + /// + /// How long one (gate, source) stays folded after a line is written. + /// + /// Long on purpose. The first sighting of a (gate, source) is ALWAYS logged immediately, so a + /// tester never waits for their answer; the window only governs REPEATS, and a second identical + /// refusal from the same address through the same gate tells nobody anything. A tester who fixes their + /// token and now fails the CIDR check instead trips a different gate, which is a different key, which + /// logs at once — so iterating stays responsive while a fixed fault stays quiet. + /// + internal static readonly TimeSpan DefaultWindow = TimeSpan.FromMinutes(10); + + /// Distinct sources tracked before new ones fold into the per-gate aggregate. + internal const int DefaultMaxTrackedSources = 16; + + /// Longest a Host header is echoed into the log. + internal const int MaxEchoedLength = 64; + + /// Stands in for a remote address ASP.NET Core could not give us. It is a real state (a + /// Unix-socket or in-process transport) and the CIDR gate fails closed on it, so it needs a key. + internal const string UnknownSource = "unknown"; + + private readonly TimeSpan _window; + private readonly int _maxTrackedSources; + + /* A plain Dictionary under a lock rather than a ConcurrentDictionary: the decision reads AND writes + several fields together (last-logged, last-seen, the suppressed counter) and has to consult Count + against the cap in the same breath, which a concurrent map cannot make atomic without a lock anyway. + Contention is a non-issue because this is only ever reached on a REFUSED request. */ + private readonly object _sync = new(); + private readonly Dictionary _perSource = new(StringComparer.Ordinal); + private readonly Dictionary _perGate = new(); + + public DarlingHttpRefusalLog(TimeSpan? window = null, int? maxTrackedSources = null) + { + _window = window ?? DefaultWindow; + _maxTrackedSources = maxTrackedSources ?? DefaultMaxTrackedSources; + } + + /// Whether this refusal should be written, and what the line owes the reader. + /// Write a line for this refusal. + /// Refusals folded into silence since the last line for this key. + /// Reported ON the next line, so the volume is visible without being written. + /// This line speaks for the per-gate aggregate rather than one source, + /// because the per-source budget was full. + internal readonly record struct Decision(bool Log, int SuppressedSinceLastLog, bool Aggregated); + + /// + /// Records one refusal and decides whether it earns a log line. Pure with respect to the clock — the + /// caller passes — so the whole policy is testable without a host, a socket + /// or a wait. + /// + public Decision Observe(DarlingRefusalGate gate, string? source, DateTime nowUtc) + { + var key = string.Concat( + gate.ToString(), + "|", + string.IsNullOrWhiteSpace(source) ? UnknownSource : source); + + lock (_sync) + { + Evict(nowUtc); + + if (_perSource.TryGetValue(key, out var tracked)) + { + return Fold(tracked, nowUtc, aggregated: false); + } + + if (_perSource.Count < _maxTrackedSources) + { + /* First sighting: logged immediately and unconditionally. This is the line the operator is + waiting for, and making it wait for a window would be the whole feature failing. */ + _perSource[key] = new Entry { LastLoggedUtc = nowUtc, LastSeenUtc = nowUtc, Suppressed = 0 }; + return new Decision(Log: true, SuppressedSinceLastLog: 0, Aggregated: false); + } + + if (!_perGate.TryGetValue(gate, out var aggregate)) + { + _perGate[gate] = new Entry { LastLoggedUtc = nowUtc, LastSeenUtc = nowUtc, Suppressed = 0 }; + return new Decision(Log: true, SuppressedSinceLastLog: 0, Aggregated: true); + } + + return Fold(aggregate, nowUtc, aggregated: true); + } + } + + private Decision Fold(Entry entry, DateTime nowUtc, bool aggregated) + { + entry.LastSeenUtc = nowUtc; + + if (nowUtc - entry.LastLoggedUtc < _window) + { + entry.Suppressed++; + return new Decision(Log: false, SuppressedSinceLastLog: 0, Aggregated: aggregated); + } + + var suppressed = entry.Suppressed; + entry.Suppressed = 0; + entry.LastLoggedUtc = nowUtc; + return new Decision(Log: true, SuppressedSinceLastLog: suppressed, Aggregated: aggregated); + } + + /* Two windows of silence and a source is forgotten, which is what frees the per-source budget again + after a scan stops. Keyed on LAST SEEN rather than last logged: an address still being refused every + second must not be evicted mid-window and then re-log as a first sighting, which would turn the + rate limit into a rate multiplier. */ + private void Evict(DateTime nowUtc) + { + var horizon = nowUtc - (_window + _window); + + List? staleSources = null; + foreach (var pair in _perSource) + { + if (pair.Value.LastSeenUtc < horizon) + { + (staleSources ??= new List()).Add(pair.Key); + } + } + + if (staleSources is not null) + { + foreach (var stale in staleSources) + { + _perSource.Remove(stale); + } + } + + List? staleGates = null; + foreach (var pair in _perGate) + { + if (pair.Value.LastSeenUtc < horizon) + { + (staleGates ??= new List()).Add(pair.Key); + } + } + + if (staleGates is not null) + { + foreach (var stale in staleGates) + { + _perGate.Remove(stale); + } + } + } + + /* Mutable, and only ever touched under _sync. LastLoggedUtc drives the window; LastSeenUtc drives + eviction; Suppressed is what the next line owes the reader. */ + private sealed class Entry + { + public DateTime LastLoggedUtc; + public DateTime LastSeenUtc; + public int Suppressed; + } + + /// Entries tracked right now — for the tests, which assert that the cap and the eviction + /// actually bound this rather than trusting that they do. + internal int TrackedCount + { + get { lock (_sync) { return _perSource.Count + _perGate.Count; } } + } + + /// + /// What actually happened, so the verb agrees with the status code (review catch on #2479). + /// + /// Not every gate rejection ends in a 4xx. The web dashboard answers an in-CIDR request carrying + /// a WRONG ?token= with a 200 and the login page — deliberately, and it is exactly the state an + /// operator asks about when a token they pasted did not work, which is why it is logged at all. But + /// "refused a request … with 200" is self-contradictory on its face, and it would mislead anyone + /// reading the log cold or filtering it for denials on status. + /// + internal static string DescribeOutcome(int statusCode) => + statusCode >= 400 ? "refused" : "did not authorize"; + + /// The gate, named the way an operator would have to name it to fix it. + internal static string Describe(DarlingRefusalGate gate) => gate switch + { + DarlingRefusalGate.HostAllowlist => "Host header allowlist", + DarlingRefusalGate.Token => "access token", + DarlingRefusalGate.SourceCidr => "source address allowlist (network.allowFrom)", + _ => gate.ToString(), + }; + + /// The remote address as a log key and a log value, or . + internal static string DescribeSource(IPAddress? remote) + { + if (remote is null) + { + return UnknownSource; + } + + var ip = remote.IsIPv4MappedToIPv6 ? remote.MapToIPv4() : remote; + return ip.ToString(); + } + + /// + /// Makes attacker-supplied text safe to put in a log line: control characters become '.', and the + /// result is truncated. + /// + /// CR and LF are the ones that matter. A log file is a text file and a reader — human or + /// otherwise — splits it on newlines, so a Host header carrying \r\n could forge whole log + /// entries. Truncation bounds the other half of the problem: a megabyte header should not become a + /// megabyte of log. + /// + internal static string Sanitize(string? value, int maxLength = MaxEchoedLength) + { + if (string.IsNullOrEmpty(value)) + { + return "(none)"; + } + + var take = Math.Min(value.Length, maxLength); + var builder = new StringBuilder(take + 1); + for (var i = 0; i < take; i++) + { + var c = value[i]; + builder.Append(c < ' ' || c == (char)0x7F ? '.' : c); + } + + if (value.Length > maxLength) + { + builder.Append('…'); + } + + return builder.ToString(); + } + + /// + /// The "N more were suppressed" clause, or empty when this is the first or only line. + /// + /// The wording branches on because the two buckets fold + /// different things, and saying so wrongly is worse than saying nothing. A per-source entry folds + /// repeats from ONE address. The aggregate entry folds the cap overflow, which is by definition many + /// DIFFERENT addresses through one gate — so "N further refusals from this source" would be actively + /// false there, and false in the direction that matters: an operator reading it would go looking for + /// one busy client when what is happening is a broad scan. Review catch on #2479. + /// + internal static string DescribeSuppression(Decision decision, TimeSpan window) + { + if (decision.SuppressedSinceLastLog <= 0) + { + return string.Empty; + } + + return string.Format( + CultureInfo.InvariantCulture, + decision.Aggregated + ? " {0} further refusal(s) through this gate, from other sources, were not logged in the last {1} minute(s)." + : " {0} further refusal(s) from this source were not logged in the last {1} minute(s).", + decision.SuppressedSinceLastLog, + (int)Math.Round(window.TotalMinutes)); + } + + /// + /// Observes one refusal and writes the WARN if it earns one. Both hosts call this, so the wording, + /// the rate limit and the redaction rules cannot drift between them. + /// + public void Report( + ILogger logger, + string surface, + DarlingRefusalGate gate, + int statusCode, + IPAddress? remote, + string reason, + DateTime nowUtc) + { + var source = DescribeSource(remote); + var decision = Observe(gate, source, nowUtc); + if (!decision.Log) + { + return; + } + + var scope = decision.Aggregated + ? " (speaking for several sources — the per-source log budget is full, which usually means this " + + "port is being scanned)" + : string.Empty; + + logger.LogWarning( + "{Surface} {Outcome} a request from {Source} (HTTP {Status}): the {Gate} gate rejected it — {Reason}.{Suppressed}{Scope}", + surface, + DescribeOutcome(statusCode), + source, + statusCode, + Describe(gate), + reason, + DescribeSuppression(decision, _window), + scope); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingWebTls.cs b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingWebTls.cs new file mode 100644 index 000000000..b8a93fe28 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Hosting/DarlingWebTls.cs @@ -0,0 +1,465 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Security.Cryptography.X509Certificates; + +namespace PerformanceMonitor.Darling.Service.Hosting; + +/// +/// Certificate resolution for the web dashboard's optional HTTPS listener (#2562). The dashboard shipped +/// plain-HTTP in every mode, so a LAN-exposed dashboard's access token and its HMAC session cookie crossed +/// the segment in the clear; this is the surface's own TLS rather than the reverse proxy that used to be the +/// only named MITM control. +/// +/// Web only, deliberately not MCP. The MCP endpoint stays plain HTTP on the older rationale that +/// a self-signed certificate breaks real MCP clients. That rationale is about MCP clients, not about the +/// wire, and it does not transfer: a browser is the web dashboard's only client, browsers have a +/// well-understood story for an internal CA, and the operator here supplies a certificate rather than the +/// product minting one. If MCP ever gets TLS it is a separate decision with a separate blast radius. +/// +/// Split on purpose. , and +/// are PURE — they decide shape and validity with no file, no clock, and no +/// logger, so the whole matrix pins in a unit test with no certificate on disk. is the one +/// effectful member. That is the same split the bind ladder uses in . +/// +/// Fail closed, never downgrade. Every failure here is reported to the caller as an exception or +/// a refusal string, and the web host answers it exactly as it answers an undecryptable token: refuse to +/// expose, bind loopback-only, log Critical. An operator who configured TLS and got plain HTTP on the LAN +/// anyway would have the one outcome this feature exists to prevent, so a broken certificate must never +/// resolve to "carry on without it". +/// +internal static class DarlingWebTls +{ + /// How far ahead of expiry the startup log begins warning. + internal const int ExpiryWarningDays = 30; + + /// What the web.network.tls block asks for, decided before any file is opened. + internal enum TlsShape + { + /// No tls block, or one whose every field is blank — plain HTTP (the caller warns when exposed). + NotConfigured, + + /// A PKCS#12 bundle: pfxPath, optionally with a password. + Pfx, + + /// A PEM pair: certPath + keyPath. + Pem, + + /// Configured but not usable as written — ambiguous or incomplete. Fail closed, never guess. + Invalid, + } + + /// The verdict from . is non-null exactly when + /// is . is independent of both: a + /// usable block that still says something the operator probably did not mean. + internal readonly record struct TlsPlan(TlsShape Shape, string? Problem, string? Warning = null); + + /// + /// A loaded certificate and the intermediates that must travel with it. Separate fields rather than one + /// collection because Kestrel wants them separately (ServerCertificate + + /// ServerCertificateChain), and because conflating "the certificate we present" with "the certs + /// that prove it" is the confusion that produced the bug this type exists to fix. + /// + internal readonly record struct LoadedCertificate(X509Certificate2 Leaf, X509Certificate2Collection Chain) + : IDisposable + { + /// Disposes the leaf AND every intermediate — on Windows each holds machine key-store state. + public void Dispose() + { + Leaf.Dispose(); + foreach (var extra in Chain) + { + extra.Dispose(); + } + } + } + + /// + /// PURE classification of the tls block. Never touches the filesystem: "the config names a PFX" and + /// "that PFX is loadable" are different failures at different times, and separating them keeps the whole + /// decision table testable and the error messages specific. + /// + /// Both forms configured is , not a precedence rule. A precedence rule + /// would silently serve one certificate while the operator watched the other one expire — and picking the + /// wrong one is indistinguishable from working until the day it is not. + /// + internal static TlsPlan Describe(WebTlsConfig? tls) + { + if (tls is null) + { + return new TlsPlan(TlsShape.NotConfigured, null); + } + + var hasPfx = !string.IsNullOrWhiteSpace(tls.PfxPath); + var hasCert = !string.IsNullOrWhiteSpace(tls.CertPath); + var hasKey = !string.IsNullOrWhiteSpace(tls.KeyPath); + var hasPassword = + !string.IsNullOrWhiteSpace(tls.PfxPassword) || !string.IsNullOrWhiteSpace(tls.EncryptedPfxPassword); + + if (hasPfx && (hasCert || hasKey)) + { + return new TlsPlan( + TlsShape.Invalid, + "web.network.tls names BOTH a PKCS#12 bundle (pfxPath) and a PEM pair (certPath/keyPath) — " + + "set one form or the other, never both, so the certificate actually served is the one you meant."); + } + + if (hasPfx) + { + return new TlsPlan(TlsShape.Pfx, null); + } + + if (hasCert && hasKey) + { + /* A PKCS#12 password left behind on a PEM deployment is inert — the certificate served is + unambiguous — so this WARNS rather than refusing, unlike the password-with-no-bundle case + below. The difference is what the operator believes: there, nothing is configured and they + think TLS is on, so refusing is the only thing that prevents plain HTTP; here TLS genuinely + is on, and taking a working dashboard down over a stale key would be the worse outcome. The + case that makes it worth saying anything at all is a half-finished PEM->PFX migration, where + the password landed before the pfxPath and the operator is watching the wrong certificate. */ + return new TlsPlan( + TlsShape.Pem, + null, + hasPassword + ? "web.network.tls sets a PKCS#12 password alongside a PEM pair — the PEM pair is being served " + + "and the password is ignored. Remove it, or finish setting pfxPath if the bundle was the one you meant." + : null); + } + + if (hasCert || hasKey) + { + /* A PEM certificate is not a keypair. Half a pair reads like a typo, and loading the certificate + without its key would produce a listener that completes no handshake. */ + return new TlsPlan( + TlsShape.Invalid, + hasCert + ? "web.network.tls sets certPath with no keyPath — a PEM certificate cannot serve TLS without its private key." + : "web.network.tls sets keyPath with no certPath — name the PEM certificate that key belongs to."); + } + + if (hasPassword) + { + /* A password with nothing to unlock is the shape of a half-finished edit, and the operator who + wrote it believes TLS is on. Refusing is louder than ignoring it. */ + return new TlsPlan( + TlsShape.Invalid, + "web.network.tls sets a PKCS#12 password but no pfxPath — there is no bundle for it to open."); + } + + return new TlsPlan(TlsShape.NotConfigured, null); + } + + /// + /// PURE lifetime gate: the reason this certificate cannot be served AT ALL, or null when it is usable. + /// Expired and not-yet-valid both refuse, because a listener that presents either one fails every + /// handshake — the dashboard is down whether we refuse here or the browser refuses there, and refusing + /// here says why in the service log instead of leaving it to a certificate warning nobody reads. + /// + /// Not-yet-valid is worth its own arm: it is the signature of a clock skew or a certificate issued + /// for a future rotation, and "expired" would be an actively misleading thing to log for it. + /// + internal static string? LifetimeRefusal(DateTimeOffset notBefore, DateTimeOffset notAfter, DateTimeOffset nowUtc) + { + if (nowUtc >= notAfter) + { + return $"the certificate expired on {notAfter.UtcDateTime:u} — TLS cannot be served with it"; + } + + if (nowUtc < notBefore) + { + return $"the certificate is not valid until {notBefore.UtcDateTime:u} (check the system clock) — TLS cannot be served with it yet"; + } + + return null; + } + + /// + /// PURE advance warning for a certificate that is usable today and expires within + /// ; null otherwise. A certificate that expires takes the dashboard down + /// with it, and this is a headless service — nobody is watching the padlock. The startup log is the only + /// place an operator can learn this before the outage. + /// + internal static string? ExpiryWarning(DateTimeOffset notAfter, DateTimeOffset nowUtc) + { + var remaining = notAfter - nowUtc; + if (remaining <= TimeSpan.Zero || remaining > TimeSpan.FromDays(ExpiryWarningDays)) + { + return null; + } + + /* Whole days, rounded UP, so the last 23 hours read "1 day" rather than "0 days". */ + var days = (int)Math.Ceiling(remaining.TotalDays); + return $"expires in {days} day{(days == 1 ? string.Empty : "s")} ({notAfter.UtcDateTime:u})"; + } + + /// + /// EFFECTFUL load of the configured certificate. Throws naming the + /// setting and the path on any failure — the caller turns that into the fail-closed degrade. + /// + /// The block, already classified by . + /// The classification, so this never re-decides what the config meant. + internal static LoadedCertificate Load(WebTlsConfig tls, TlsShape shape) + { + ArgumentNullException.ThrowIfNull(tls); + + return shape switch + { + TlsShape.Pfx => LoadPfx(tls), + TlsShape.Pem => LoadPem(tls), + _ => throw new InvalidOperationException( + $"web.network.tls cannot be loaded in shape {shape} — Describe() must be consulted first."), + }; + } + + private static LoadedCertificate LoadPfx(WebTlsConfig tls) + { + var path = tls.PfxPath!.Trim(); + RequireFile(path, "web.network.tls.pfxPath"); + + string? password; + try + { + password = tls.ResolvePfxPassword(out _); + } + catch (Exception ex) + { + /* An env:/file: reference that does not resolve, or a DPAPI blob from another machine. Naming the + setting matters more than the exception type: the operator has three password slots. */ + throw new InvalidOperationException( + $"web.network.tls: the PKCS#12 password could not be resolved ({ex.Message})", ex); + } + + X509Certificate2Collection bundle; + try + { + /* The COLLECTION loader, not LoadPkcs12FromFile: a PKCS#12 bundle routinely carries the issuing + chain beside the leaf, and the single-certificate loader returns only one of them — silently, + so the listener comes up and serves an incomplete chain. */ + bundle = X509CertificateLoader.LoadPkcs12CollectionFromFile(path, password, KeyStorageFlags()); + } + catch (Exception ex) + { + throw new InvalidOperationException( + $"web.network.tls.pfxPath '{path}' could not be loaded ({ex.Message}) — " + + "check the password slot too: Windows reports a wrong PKCS#12 password as unreadable data " + + "rather than as a bad password.", + ex); + } + + var leaf = FindLeaf(bundle); + if (leaf is null) + { + foreach (var candidate in bundle) + { + candidate.Dispose(); + } + + throw new InvalidOperationException( + $"web.network.tls.pfxPath '{path}' contains no certificate with a private key — a TLS server " + + "certificate must carry its key (export the bundle with the key included)."); + } + + return new LoadedCertificate(leaf, IntermediatesOf(bundle, leaf)); + } + + private static LoadedCertificate LoadPem(WebTlsConfig tls) + { + var certPath = tls.CertPath!.Trim(); + var keyPath = tls.KeyPath!.Trim(); + RequireFile(certPath, "web.network.tls.certPath"); + RequireFile(keyPath, "web.network.tls.keyPath"); + + /* Read the WHOLE file, not just the leaf. CreateFromPemFile below materializes only the FIRST + certificate, so on its own it drops every intermediate the operator appended — which is exactly what + the config doc tells them to do, and exactly the incomplete-chain handshake failure that produces on + any client that has not independently cached the intermediate. Measured before this was fixed: a PEM + holding leaf + intermediate served one certificate. */ + var bundle = new X509Certificate2Collection(); + X509Certificate2 fromPem; + try + { + bundle.ImportFromPemFile(certPath); + fromPem = X509Certificate2.CreateFromPemFile(certPath, keyPath); + } + catch (Exception ex) + { + foreach (var loaded in bundle) + { + loaded.Dispose(); + } + + throw new InvalidOperationException( + $"web.network.tls: the PEM pair '{certPath}' / '{keyPath}' could not be loaded ({ex.Message})", ex); + } + + /* FOOTGUN (load-bearing): a certificate built from PEM carries an EPHEMERAL private key, and Windows' + SslStream cannot use an ephemeral key for server authentication — Kestrel accepts the certificate at + configuration time and then fails every handshake at runtime, which is the worst possible place for + this to surface. Round-tripping through an in-memory PKCS#12 re-associates the key through the + platform's own key store and is the standard fix. Costs one export/import at startup, once. */ + X509Certificate2 leaf; + using (fromPem) + { + var pkcs12 = fromPem.Export(X509ContentType.Pkcs12); + try + { + leaf = X509CertificateLoader.LoadPkcs12(pkcs12, password: null, KeyStorageFlags()); + } + catch (Exception ex) + { + foreach (var loaded in bundle) + { + loaded.Dispose(); + } + + throw new InvalidOperationException( + $"web.network.tls: the PEM pair '{certPath}' / '{keyPath}' loaded but could not be prepared " + + $"for the TLS listener ({ex.Message})", + ex); + } + } + + return new LoadedCertificate(leaf, IntermediatesOf(bundle, leaf)); + } + + /// + /// The end-entity certificate in a PKCS#12 bundle, or null when there is none to serve. + /// + /// "The first one with a private key" is not good enough, which a test caught rather than a + /// review: a bundle exported wholesale from a CA machine can carry keys for the intermediate and even the + /// root, and that rule then happily serves the ROOT as the server certificate. Position is no help either + /// — PKCS#12 ordering is not a contract. + /// + /// So the leaf is identified by what actually makes it a leaf: it holds a private key AND it is the + /// terminal node — no other certificate in the bundle was issued by it. Compared on the raw encoded + /// names, never the rendered strings, because X.500 name formatting is not canonical. + /// + /// The fallback to "first with a key" exists for a bundle that defeats the rule (a cross-signed + /// oddity, or a chain whose links are not all present). Serving something the operator supplied and + /// letting the handshake or the SAN warning report it beats refusing to start over a shape we did not + /// anticipate. + /// + private static X509Certificate2? FindLeaf(X509Certificate2Collection bundle) + { + X509Certificate2? firstWithKey = null; + + foreach (var candidate in bundle) + { + if (!candidate.HasPrivateKey) + { + continue; + } + + firstWithKey ??= candidate; + + var issuedSomething = false; + foreach (var other in bundle) + { + if (ReferenceEquals(other, candidate)) + { + continue; + } + + if (other.IssuerName.RawData.AsSpan().SequenceEqual(candidate.SubjectName.RawData)) + { + issuedSomething = true; + break; + } + } + + if (!issuedSomething) + { + return candidate; + } + } + + return firstWithKey; + } + + /// + /// The certificates from that must accompany on the + /// wire, disposing the ones that must not. Two exclusions, each deliberate: + /// + /// + /// The leaf itself, matched by thumbprint — Kestrel is handed it separately, and sending it + /// twice is a malformed chain. + /// Any self-issued root. A root a client does not already trust is not made trustworthy by + /// our sending it, and a root it does trust it already has; either way it is handshake bytes that buy + /// nothing. Bundles routinely include one because that is what "export the whole chain" produces. + /// + /// + private static X509Certificate2Collection IntermediatesOf(X509Certificate2Collection bundle, X509Certificate2 leaf) + { + var chain = new X509Certificate2Collection(); + foreach (var candidate in bundle) + { + /* Self-issued is decided on the RAW encoded names, matching FindLeaf above and for the same + reason: X.500 name formatting is not canonical, so two encodings of the same DN can render + differently (different string types, different attribute ordering) and still be one name. + Getting it wrong here sends a root down the wire that buys nothing. */ + if (string.Equals(candidate.Thumbprint, leaf.Thumbprint, StringComparison.OrdinalIgnoreCase) + || candidate.SubjectName.RawData.AsSpan().SequenceEqual(candidate.IssuerName.RawData)) + { + /* Not the caller's to release: the leaf is returned separately and lives on. */ + if (!ReferenceEquals(candidate, leaf)) + { + candidate.Dispose(); + } + + continue; + } + + chain.Add(candidate); + } + + return chain; + } + + /// + /// Key storage flags per platform. Neither arm is arbitrary, and the one that looks most obviously + /// correct — everywhere, keeping the key out of any + /// store — is wrong on BOTH platforms, in two different ways: + /// + /// + /// Windows accepts an ephemeral key and then cannot use it for TLS SERVER authentication. + /// The load succeeds, Kestrel configures happily, and every handshake fails at runtime — the same + /// limitation the PEM round-trip above exists to work around, arriving at the worst possible moment. + /// macOS refuses the flag outright ("This platform does not support loading with + /// EphemeralKeySet"), so the load throws and the dashboard fail-closes to loopback — caught only because + /// the shipped loader was run against a real certificate on a Mac rather than reasoned about. + /// + /// + /// Windows: . The service runs as a virtual + /// account (NT SERVICE\PerformanceMonitor Darling) with no loaded user profile, so the default user + /// key set is not available to it. Deliberately WITHOUT + /// : without that flag the key material is removed from the + /// machine key store when the certificate is disposed, which is why every bail path in the web host + /// disposes it. + /// + /// Everywhere else (Linux containers, macOS dev): the default key set. There is no machine + /// key store to opt into, and on Linux — the compose distribution's platform — the key is held in process + /// memory regardless of the flag. + /// + private static X509KeyStorageFlags KeyStorageFlags() + => OperatingSystem.IsWindows() + ? X509KeyStorageFlags.MachineKeySet + : X509KeyStorageFlags.DefaultKeySet; + + /// Existence check that names the SETTING as well as the path — the caller has four path slots + /// and "file not found" alone sends them to the wrong one. + private static void RequireFile(string path, string settingName) + { + if (!File.Exists(path)) + { + throw new InvalidOperationException($"{settingName} '{path}' does not exist or is not readable"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/HypotheticalIndexRequest.cs b/Darling/PerformanceMonitor.Darling.Service/HypotheticalIndexRequest.cs new file mode 100644 index 000000000..d08f86797 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/HypotheticalIndexRequest.cs @@ -0,0 +1,146 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Text.Json; +using System.Text.Json.Serialization; + +namespace PerformanceMonitor.Darling.Service; + +/// +/// The subject of a hypothetical-index experiment (#2612): one candidate index, and one stored statement +/// to re-plan with and without it. +/// +/// +/// On demand only, never scheduled — the decision recorded on the issue before any code was written. +/// There is no collector, no cadence and no sweep: this is a tool invocation with a caller and a subject, +/// driven from a pg_predicate_stats row a human is already looking at. That collector separates the +/// two reasons a column looks interesting, and only one of them is this experiment's business: poor +/// selectivity, where an index might help. A large estimate error means the planner does not understand +/// the column, and an index will not fix a plan built on a wrong row count. +/// +/// +/// +/// Which server. The registered server the caller names, and no other. The statistics that made the +/// column a candidate are that server's statistics, and a hypothetical index tested against different ones +/// answers a different question. An operator who wants the experiment run on a replica registers the +/// replica — which is a decision they can see, rather than a substitution the product made quietly. +/// +/// +/// +/// Property names are camelCase to match the viewer's serialization (parsed case-insensitively). Pure args +/// model plus validation; never throws. +/// +/// +public sealed record HypotheticalIndexRequest( + [property: JsonPropertyName("queryid")] string? QueryId, + [property: JsonPropertyName("schemaName")] string? SchemaName, + [property: JsonPropertyName("tableName")] string? TableName, + [property: JsonPropertyName("columns")] IReadOnlyList? Columns, + [property: JsonPropertyName("databaseName")] string? DatabaseName) +{ + /// + /// How many columns one candidate may carry. A bound rather than a judgment about index design: the + /// column list is spliced into DDL text handed to hypopg_create_index, and an unbounded list is + /// an unbounded statement. + /// + public const int MaxColumns = 8; + + /// + /// queryid travels as a STRING, both directions, like every other queryid on this surface: it is + /// a signed 64-bit value and a JSON number would round it in any double-decoding parser, producing an + /// id that resolves to no stored statement. + /// + public bool TryGetQueryId(out long queryId) + => long.TryParse(QueryId, System.Globalization.NumberStyles.Integer, + System.Globalization.CultureInfo.InvariantCulture, out queryId); + + /// + /// True when the request names a statement AND a candidate. Both halves are required and neither has a + /// sensible default: without the statement there is nothing to re-plan, and without the candidate there + /// is no experiment — only an EXPLAIN of somebody's query, which this is not for. + /// + public bool IsComplete => + TryGetQueryId(out _) + && IsSafeIdentifier(SchemaName) + && IsSafeIdentifier(TableName) + && Columns is { Count: > 0 and <= MaxColumns } + && Columns.All(IsSafeIdentifier); + + /// + /// Identifier acceptance, and it is deliberately narrow. + /// + /// These names are spliced into a DDL string that hypopg_create_index parses, so they are + /// the one place in this feature where caller-supplied text reaches SQL text. They cannot be passed as + /// parameters — the function takes a whole CREATE INDEX statement as a string — so the defence is that + /// nothing but an unqualified, unquoted identifier is accepted at all. A name needing quoting is + /// refused rather than escaped: refusing is auditable, and escaping is the thing that gets one case + /// wrong three years later. + /// + public static bool IsSafeIdentifier(string? value) + => !string.IsNullOrWhiteSpace(value) + && value.Length <= 63 + && (char.IsLetter(value[0]) || value[0] == '_') + && value.All(c => char.IsLetterOrDigit(c) || c == '_'); + + private static readonly JsonSerializerOptions s_options = new() + { + PropertyNameCaseInsensitive = true, + ReadCommentHandling = JsonCommentHandling.Skip, + AllowTrailingCommas = true, + }; + + /// + /// Parses args_json, returning false for null, blank, malformed, or incomplete input. Never + /// throws — the dispatch turns a false here into a command failure with a message, which is the only + /// way a caller finds out they sent something unusable. + /// + public static bool TryParse(string? argsJson, out HypotheticalIndexRequest request) + { + request = new HypotheticalIndexRequest(null, null, null, null, null); + + if (string.IsNullOrWhiteSpace(argsJson)) + { + return false; + } + + try + { + var parsed = JsonSerializer.Deserialize(argsJson, s_options); + + if (parsed is null || !parsed.IsComplete) + { + return false; + } + + request = parsed; + return true; + } + catch (JsonException) + { + return false; + } + } + + /// + /// The CREATE INDEX text handed to hypopg_create_index. Composed only from identifiers + /// that already passed , and asserted again here so a future caller that + /// skips cannot compose one. + /// + public string BuildCreateIndexStatement() + { + if (!IsComplete) + { + throw new InvalidOperationException("An incomplete hypothetical-index request cannot compose DDL."); + } + + return $"CREATE INDEX ON {SchemaName}.{TableName} ({string.Join(", ", Columns!)})"; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingAlertReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingAlertReader.cs index acf784f47..845ed3f4b 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingAlertReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingAlertReader.cs @@ -52,26 +52,32 @@ public sealed record AlertHistoryReadRow( muted, detail_text"; - /// Per-server alert history — the viewer's AlertHistorySql. $1 window start, $2 server_id, - /// $3 limit (naive UTC / int / int). + /// Per-server alert history — the viewer's AlertHistorySql. $1 window start, $2 window + /// end, $3 server_id, $4 limit (naive UTC / naive UTC / int / int). + /// + /// The upper edge is bounded rather than open (#2495): the row cap is applied by the database, so + /// trimming after the read would spend the whole LIMIT on rows newer than the anchor and hand back an + /// empty window that looks like a quiet one. public const string AlertHistorySql = @" SELECT" + AlertHistorySelectColumns + @" FROM config_alert_log WHERE alert_time >= $1 -AND server_id = $2 +AND alert_time <= $2 +AND server_id = $3 AND dismissed = FALSE ORDER BY alert_time DESC -LIMIT $3"; +LIMIT $4"; /// All-servers alert history (the fleet default) — the viewer's AlertHistoryAllServersSql. - /// $1 window start, $2 limit (naive UTC / int). + /// $1 window start, $2 window end, $3 limit (naive UTC / naive UTC / int). public const string AlertHistoryAllServersSql = @" SELECT" + AlertHistorySelectColumns + @" FROM config_alert_log WHERE alert_time >= $1 +AND alert_time <= $2 AND dismissed = FALSE ORDER BY alert_time DESC -LIMIT $2"; +LIMIT $3"; /// /// Recent alerts newest first, excluding dismissed rows — the Alert History read. With no @@ -79,12 +85,13 @@ ORDER BY alert_time DESC /// server. Mirrors the viewer's optional-serverId GetAlertHistoryAsync. /// public static async Task> GetAlertHistoryAsync( - NpgsqlDataSource postgres, DateTime sinceUtc, int? serverId, int limit, CancellationToken cancellationToken = default) + NpgsqlDataSource postgres, DateTime sinceUtc, DateTime untilUtc, int? serverId, int limit, CancellationToken cancellationToken = default) { var rows = new List(); await using var command = postgres.CreateCommand(serverId.HasValue ? AlertHistorySql : AlertHistoryAllServersSql); DarlingMcpReadParameters.AddTimestamp(command, sinceUtc); + DarlingMcpReadParameters.AddTimestamp(command, untilUtc); if (serverId.HasValue) { DarlingMcpReadParameters.AddInt(command, serverId.Value); @@ -144,10 +151,16 @@ public sealed record AlertSettingsReadRow( int DiskCriticalFreePercent, int DiskCriticalFreeGb, int AnalysisNotifyCooldownMinutes, - int StoreJobCadenceWarnPercent); + int StoreJobCadenceWarnPercent, + /* #2391 (V79, #2349's knobs): APPENDED, never inserted — every field above is positional and read + by ordinal, so placing these anywhere but the end would silently re-map all of them. */ + bool FileGrowthEnabled, + int FileGrowthRiseMb, + int FileGrowthVolumePercent, + int FileGrowthLookbackMinutes); /// The single global alert-settings row (id=1) — the viewer's AlertSettingsSelectSql. The - /// 47 columns are read in the SAME order the service reads them (StoreConfigProvider). This had + /// 58 columns are read in the SAME order the service reads them (StoreConfigProvider). This had /// stopped at 36, so get_alert_settings reported a store whose newest five knobs did not exist: /// an MCP client could not see the V33 connection opt-ins or the V35 Availability Group family at all. public const string AlertSettingsSelectSql = @" @@ -167,7 +180,8 @@ public sealed record AlertSettingsReadRow( pvs_floor_gb, database_state_enabled, self_disk_free_warn_percent, collection_stale_minutes, collection_failure_threshold, disk_critical_free_percent, disk_critical_free_gb, analysis_notify_cooldown_minutes, - store_job_cadence_warn_percent + store_job_cadence_warn_percent, + file_growth_enabled, file_growth_rise_mb, file_growth_volume_percent, file_growth_lookback_minutes FROM config_alert_settings WHERE id = 1"; @@ -205,6 +219,8 @@ FROM config_alert_settings /* #2107 threshold knobs (V55) at 47–52; #2136 cadence-warn knob (V57) at 53. */ reader.GetInt32(47), reader.GetInt32(48), reader.GetInt32(49), reader.GetInt32(50), reader.GetInt32(51), reader.GetInt32(52), - reader.GetInt32(53)); + reader.GetInt32(53), + /* #2391: V79 file-growth knobs at 54–57. */ + reader.GetBoolean(54), reader.GetInt32(55), reader.GetInt32(56), reader.GetInt32(57)); } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingBlockingTrendReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingBlockingTrendReader.cs index 0931fbe6d..290c57de2 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingBlockingTrendReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingBlockingTrendReader.cs @@ -15,14 +15,16 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// -/// Service-side reads for the blocking-incident and deadlock per-minute trend MCP tools -/// ( get_blocking_trend / get_deadlock_trend). The SQL is reproduced -/// verbatim from the viewer's Blocking-Trends charts (ViewerDataService.BlockingTrends.cs), which are -/// Lite's GetBlockingTrendAsync / GetDeadlockTrendAsync ported to Postgres. Both are STORED -/// reads (no live monitored-server hit) sharing a (bucket timestamp, COUNT(*)) shape, so one reader -/// maps both; COUNT(*) is bigint in Postgres, read via GetInt64 and narrowed to the point's int. +/// Service-side reads for the three Blocking-Trends MCP tools ( +/// get_blocking_trend / get_deadlock_trend / get_lock_wait_trend). The SQL is reproduced verbatim from the +/// viewer's Blocking-Trends charts (ViewerDataService.BlockingTrends.cs), which are Lite's +/// GetBlockingTrendAsync / GetDeadlockTrendAsync / GetLockWaitTrendAsync ported to +/// Postgres. All are STORED reads (no live monitored-server hit). The first two share a +/// (bucket timestamp, COUNT(*)) shape, so one reader maps both; COUNT(*) is bigint in Postgres, +/// read via GetInt64 and narrowed to the point's int. The lock-wait lane has its own mapper — it is a +/// fractional RATE per (collection, wait type) rather than a count. /// Public-const SQL so Darling.Tests pin the dialect (the XE-preferred + DMV-fallback union, the deadlock -/// bucket-on-deadlock_time) without a live Postgres. +/// bucket-on-deadlock_time, the lock-wait LAG interval) without a live Postgres. /// internal static class DarlingBlockingTrendReader { @@ -77,6 +79,43 @@ GROUP BY DATE_TRUNC('minute', deadlock_time) ORDER BY bucket """; + /// One LCK% wait type's per-second wait rate at one collection (mirror of the viewer's + /// LockWaitTrendPoint). + public sealed record LockWaitTrendReadPoint(DateTime CollectionTime, string WaitType, double WaitTimeMsPerSecond); + + /// + /// LCK% wait per-second rates — the viewer's LockWaitTrendSql, VERBATIM. Reads + /// v_wait_stats filtered to lock waits, derives each row's collection interval from the LAG of + /// the prior collection_time (partitioned by wait type), and divides the delta wait time by that + /// interval for a per-second rate. The delta is CAST to double precision before the division so a + /// sub-one-millisecond-per-second rate does not truncate to zero — the same defect #2507 found in the + /// execution-count trend, where a quiet server reported as an idle one. Negative deltas are dropped: + /// those are the counter reset across a SQL Server restart, not a negative wait. $1 server_id, + /// $2 window start, $3 window end (naive UTC). + /// + public const string LockWaitTrendSql = """ + WITH raw AS + ( + SELECT + collection_time, + wait_type, + delta_wait_time_ms, + extract(epoch FROM (date_trunc('second', collection_time) - date_trunc('second', LAG(collection_time) OVER (PARTITION BY wait_type ORDER BY collection_time)))) AS interval_seconds + FROM v_wait_stats + WHERE server_id = $1 + AND wait_type LIKE 'LCK%' + AND collection_time >= $2 + AND collection_time <= $3 + ) + SELECT + collection_time, + wait_type, + CASE WHEN interval_seconds > 0 THEN CAST(delta_wait_time_ms AS double precision) / interval_seconds ELSE 0 END AS wait_time_ms_per_second + FROM raw + WHERE delta_wait_time_ms >= 0 + ORDER BY collection_time, wait_type + """; + /// Blocking-incident-per-minute buckets for one server over the window. public static Task> GetBlockingTrendAsync( NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) @@ -87,6 +126,29 @@ public static Task> GetDeadlockTrendAsync( NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => ReadCountTrendAsync(postgres, DeadlockTrendSql, serverId, startUtc, endUtc, cancellationToken); + /// + /// LCK% wait per-second rates for one server over the window — one row per (collection, lock wait type). + /// Its own mapper rather than the count-trend one above: this returns three columns of two different + /// types, and the rate is a double the whole point of which is that it is fractional. + /// + public static async Task> GetLockWaitTrendAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + { + var items = new List(); + await using var command = postgres.CreateCommand(LockWaitTrendSql); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + items.Add(new LockWaitTrendReadPoint( + reader.GetDateTime(0), + reader.IsDBNull(1) ? string.Empty : reader.GetString(1), + reader.IsDBNull(2) ? 0 : reader.GetDouble(2))); + } + + return items; + } + /// The blocking and deadlock trends share a (bucket timestamp, COUNT(*)) shape, so one reader /// maps both. COUNT(*) is bigint in Postgres, read via GetInt64 and narrowed to the point's int. private static async Task> ReadCountTrendAsync( @@ -105,4 +167,145 @@ private static async Task> ReadCountTrendAsync( return items; } + + /* ───────────────── the denominator an empty trend needs ───────────────── */ + + /// + /// One collector's SUCCESSFUL run count inside a window, with the first and last of those runs. + /// Both trends read EDGE tables — rows exist only where an event happened — so "no rows" is a + /// capture that found nothing and a capture that never ran, wearing the same face. Neither table can + /// tell them apart; collection_log can, because a collector that ran and stored nothing still + /// records a SUCCESS with zero rows. Same reasoning as + /// DarlingPgBlockingReader.PgBlockingCaptureCounts, applied to the SQL Server side. + /// + public sealed record CaptureCount(string CollectorName, long Runs, DateTime? FirstRunAt, DateTime? LastRunAt); + + /// + /// Successful runs per blocking collector inside the window. BOTH capture paths, deliberately: the trend + /// above unions v_blocked_process_reports with v_dmv_blocking_snapshots, so counting one of + /// them would report "never captured" for a server capturing perfectly well through the other — the wrong + /// branch in exactly the case this exists to get right. Only SUCCESS counts as having looked; a PERMISSIONS + /// or ERROR row is a collector that did not see the window either. $1 server_id, $2/$3 window (naive UTC). + /// + public const string BlockingCaptureCountsSql = """ + SELECT + collector_name, + COUNT(*), + MIN(collection_time), + MAX(collection_time) + FROM v_collection_log + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND status = 'SUCCESS' + AND collector_name IN ('blocked_process_report', 'dmv_blocking_snapshot') + GROUP BY collector_name + ORDER BY collector_name + """; + + /// + /// Successful runs of the deadlock collector inside the window. One capture path here, not two — deadlocks + /// come only from the deadlocks collector's system_health read, and there is no DMV fallback to + /// count. $1 server_id, $2/$3 window (naive UTC). + /// + public const string DeadlockCaptureCountsSql = """ + SELECT + collector_name, + COUNT(*), + MIN(collection_time), + MAX(collection_time) + FROM v_collection_log + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND status = 'SUCCESS' + AND collector_name = 'deadlocks' + GROUP BY collector_name + ORDER BY collector_name + """; + + /// + /// Whether either blocking collector has EVER run successfully for this server, ignoring any window. + /// Asked ONLY when the window count came back zero, and only to pick which sentence is true: a + /// server whose collectors have run before has a GAP in this window (widen it, or go look at collection + /// health), while one that has never run them is not collecting blocking at all. Both are "not an + /// all-clear" and they want different next moves. LIMIT 1, so it stops at the first row. + /// NOT the same question as asking whether an EVENT was ever captured. A server collected + /// perfectly for months that simply never blocked has no event rows, and an event-existence probe + /// reports it as never captured — the reassuring-answer failure inverted, sending someone to fix + /// collection that is working. Hence "collector run", not "capture": the denominator is whether we + /// LOOKED, not whether we found something. + /// This applies to get_blocking_stats too, which originally used an event-existence probe on the + /// grounds that its verdict is about severity. That reasoning does not survive contact with a healthy + /// server: zero severity is the NORMAL state, so the same false alarm fires. It uses these. + /// + public const string HasAnyBlockingCollectorRunSql = """ + SELECT 1 + FROM v_collection_log + WHERE server_id = $1 + AND status = 'SUCCESS' + AND collector_name IN ('blocked_process_report', 'dmv_blocking_snapshot') + LIMIT 1 + """; + + /// Whether the deadlock collector has EVER run successfully for this server. See + /// for why the question is asked at all. + public const string HasAnyDeadlockCollectorRunSql = """ + SELECT 1 + FROM v_collection_log + WHERE server_id = $1 + AND status = 'SUCCESS' + AND collector_name = 'deadlocks' + LIMIT 1 + """; + + /// Runs . + public static Task> GetBlockingCaptureCountsAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + => ReadCaptureCountsAsync(postgres, BlockingCaptureCountsSql, serverId, startUtc, endUtc, cancellationToken); + + /// Runs . + public static Task> GetDeadlockCaptureCountsAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + => ReadCaptureCountsAsync(postgres, DeadlockCaptureCountsSql, serverId, startUtc, endUtc, cancellationToken); + + /// Runs . + public static Task HasAnyBlockingCollectorRunAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnyCaptureAsync(postgres, HasAnyBlockingCollectorRunSql, serverId, cancellationToken); + + /// Runs . + public static Task HasAnyDeadlockCollectorRunAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnyCaptureAsync(postgres, HasAnyDeadlockCollectorRunSql, serverId, cancellationToken); + + /// Both capture-count reads share a (collector_name, COUNT(*), MIN, MAX) shape, so one mapper + /// serves them. COUNT(*) is bigint in Postgres. + private static async Task> ReadCaptureCountsAsync( + NpgsqlDataSource postgres, string sql, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken) + { + var items = new List(); + await using var command = postgres.CreateCommand(sql); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + items.Add(new CaptureCount( + reader.GetString(0), + reader.IsDBNull(1) ? 0 : reader.GetInt64(1), + reader.IsDBNull(2) ? null : reader.GetDateTime(2), + reader.IsDBNull(3) ? null : reader.GetDateTime(3))); + } + + return items; + } + + /// Both existence probes share one shape: a scalar that is null when no row qualifies. + private static async Task HasAnyCaptureAsync( + NpgsqlDataSource postgres, string sql, int serverId, CancellationToken cancellationToken) + { + await using var command = postgres.CreateCommand(sql); + DarlingMcpReadParameters.AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingDataReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingDataReader.cs index 432d4f53c..b6fc77cfa 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingDataReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingDataReader.cs @@ -65,6 +65,32 @@ public sealed record MemoryStatsRow( /// One memory clerk's footprint at the latest snapshot. public sealed record MemoryClerkRow(string ClerkType, double MemoryMb); + /* #2484: the two series behind the viewer's Current Waits tab. Both read waiting_tasks, which the + snapshot reads already expose row-by-row -- get_waiting_tasks answers "what is waiting now" and + never "was it worse an hour ago", which is the question that decides whether anything is wrong. */ + public sealed record WaitingTaskTrendRow(DateTime CollectionTime, string WaitType, long TotalWaitMs); + + public sealed record BlockedSessionTrendRow(DateTime CollectionTime, string DatabaseName, long BlockedCount); + + /* + One row of the RAW per-run collection log, as opposed to the CollectorHealth rollup above. + The rollup answers "is this collector healthy over seven days"; this answers "what happened + on each run", which is the question an operator actually has when collection looks wrong and + the rollup says HEALTHY. Durations are split the way the collectors report them -- total, the + part spent on the monitored server, and the part spent writing to the store -- because a + collector that is slow because the target is slow needs a different fix from one that is slow + because the store is. + */ + public sealed record CollectionLogEntry( + string CollectorName, + DateTime CollectionTime, + double? DurationMs, + double? SqlDurationMs, + double? StoreDurationMs, + long? RowsCollected, + string? Status, + string? ErrorMessage); + /// One database file's latest I/O snapshot; avg latency is computed by the tool. public sealed record FileIoRow( string DatabaseName, string FileName, string FileType, string PhysicalName, double SizeMb, @@ -138,7 +164,7 @@ public sealed record ServerPropertiesReadRow( /// naive UTC by subtracting the per-batch UTC offset (#1262). Windows on collection_time (the /// reliable naive-UTC clock, not the server-local sample_time). Reads the base /// cpu_utilization_stats table (the de-skew window function needs collection_time alongside - /// sample_time). $1 server_id, $2 window start (naive UTC). + /// sample_time). $1 server_id, $2 window start, $3 window end (naive UTC). /// public const string CpuUtilizationSql = """ SELECT @@ -152,16 +178,18 @@ public sealed record ServerPropertiesReadRow( FROM cpu_utilization_stats WHERE server_id = $1 AND collection_time >= $2 + AND collection_time <= $3 ORDER BY sample_time """; public static async Task> GetCpuUtilizationAsync( - NpgsqlDataSource postgres, int serverId, DateTime startUtc, CancellationToken cancellationToken = default) + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) { var samples = new List(); await using var command = postgres.CreateCommand(CpuUtilizationSql); AddInt(command, serverId); AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); await using var reader = await command.ExecuteReaderAsync(cancellationToken); while (await reader.ReadAsync(cancellationToken)) { @@ -286,6 +314,30 @@ public static async Task> GetDistinctWaitTypesAsync( return items; } + /// + /// Whether this server has EVER recorded a wait sample, ignoring any window. + /// Lets an empty get_wait_types say WHICH kind of nothing it found. "No wait types in the last N + /// hours" is true both of a quiet window and of a server nothing has been stored for, and those want + /// opposite responses — widen the window, versus go find out why collection is not running. Reads + /// v_wait_stats, the same source reads, so it can never + /// report "collected" for rows the read cannot see. LIMIT 1, so it stops at the first row. + /// + public const string HasAnyWaitStatSql = """ + SELECT 1 + FROM v_wait_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// Runs . + public static async Task HasAnyWaitStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HasAnyWaitStatSql); + AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } + /// /// A single wait type's per-second trend — Lite's GetWaitStatsTrendAsync: the interval rate is /// this collection's delta divided by the seconds since the previous collection (a LAG over the @@ -471,7 +523,7 @@ public static async Task> GetLatestFileIoStatsAsync( /// /// tempdb space-usage samples over the window — the viewer's TempDbTrendSql. MB columns are /// numeric(18,2) → double precision; total_sessions_using_tempdb is bigint, top_session_id is - /// integer. $1 server_id, $2 window start (naive UTC). + /// integer. $1 server_id, $2 window start, $3 window end (naive UTC). /// public const string TempDbTrendSql = """ SELECT @@ -487,16 +539,18 @@ public static async Task> GetLatestFileIoStatsAsync( FROM v_tempdb_stats WHERE server_id = $1 AND collection_time >= $2 + AND collection_time <= $3 ORDER BY collection_time """; public static async Task> GetTempDbTrendAsync( - NpgsqlDataSource postgres, int serverId, DateTime startUtc, CancellationToken cancellationToken = default) + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) { var samples = new List(); await using var command = postgres.CreateCommand(TempDbTrendSql); AddInt(command, serverId); AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); await using var reader = await command.ExecuteReaderAsync(cancellationToken); while (await reader.ReadAsync(cancellationToken)) { @@ -1096,10 +1150,15 @@ public static async Task> GetServerListAsync( } /// - /// Per-collector 7-day health aggregate — the viewer's CollectionHealthSql (Lite's - /// GetCollectionHealthAsync): one row per collector with run/success/error counts, average - /// duration, last success/run/error timestamps, and the permission-denied count for the banding. - /// SKIPPED counts as a healthy run. $1 server_id, $2 window start (naive UTC — the trailing 7 days). + /// Per-collector 7-day health aggregate — Lite's GetCollectionHealthAsync: one row per + /// collector with run/success/error counts, the average / maximum / p95 single-run duration, last + /// success/run/error timestamps, and the permission-denied count for the banding. SKIPPED counts as + /// a healthy run. $1 server_id, $2 window start (naive UTC — the trailing 7 days). + /// + /// 16 columns since #2460, and no longer column-identical to the WPF viewer's own + /// CollectionHealthSql: the two duration statistics feed the MCP tool's sweep-pressure + /// arithmetic, which the viewer's health grid does not serve. Lite's DuckDB read carries them at + /// the SAME ordinals, which is the parity that matters here — both MCP surfaces read positionally. /// public const string CollectionHealthSql = """ SELECT @@ -1108,6 +1167,24 @@ public static async Task> GetServerListAsync( SUM(CASE WHEN status = 'SUCCESS' THEN 1 ELSE 0 END) AS success_count, SUM(CASE WHEN status = 'ERROR' THEN 1 ELSE 0 END) AS error_count, AVG(duration_ms) AS avg_duration_ms, + -- #2460: the mean above describes a collector whose runs all cost about the same, and + -- says nothing true about one whose runs come in two sizes. query_store on a dense shard + -- reported a 13,834 ms average over 1,155 runs where 958 of them yielded nothing and cost + -- ~36 ms, which puts the other 197 at ~80,900 ms EACH — each one on its own larger than + -- the whole 60,000 ms sweep budget. duration_ms has been written per run since V2; + -- nothing had ever read it as anything but a mean. + -- + -- p95 rather than the max for the number a decision is made from: a max is one run, so a + -- single pathological cycle would make a collector look permanently terrible for the rest + -- of the window. p95 also scales itself to the sample — over 3,500 runs it discards the + -- one bad cycle, and over the six runs a daily collector gets in a week it lands on the + -- max, which is right, because with six samples there is no outlier anyone can afford to + -- throw away. DISC rather than CONT so the answer is a duration some run actually took + -- instead of an interpolation between the two modes, which would be a number describing + -- no run at all — the exact defect this column exists to end. Both engines ignore NULL + -- duration_ms here, as AVG already does. Byte-identical to Lite's DuckDB read. + MAX(duration_ms) AS max_duration_ms, + PERCENTILE_DISC(0.95) WITHIN GROUP (ORDER BY duration_ms) AS p95_duration_ms, MAX(CASE WHEN status IN ('SUCCESS', 'SKIPPED') THEN collection_time END) AS last_success_time, MAX(collection_time) AS last_run_time, -- #1855: the message from the NEWEST failing run, not MAX()'s lexicographically greatest @@ -1154,7 +1231,21 @@ AND database_id > 4 ) THEN 1 ELSE 0 - END AS has_user_databases + END AS has_user_databases, + -- #2472: the per-database fan-out, described for ONE run — the dearest one in the window. + -- Five collectors run once per database and the run writes a single blended duration_ms, so + -- "eight databases at 10.1s" and "one at 62s beside seven at 2.7s" are the same 80,900 ms and + -- want opposite fixes. These four are that one run's parts, and their ratio + -- (slowest_item_ms * fanout_items / slowest_run_duration_ms) is 1.0 for an even fan-out and + -- 6.1 for the dominated example. All four come from the SAME row via slowest_rank rather than + -- four independent aggregates, because a slowest item taken from one run and a width taken + -- from another compose into a ratio describing no run that ever happened — the same defect + -- PERCENTILE_CONT was rejected for above. NULL throughout for a collector that never fans + -- out, which is most of them. + MAX(CASE WHEN slowest_rank = 1 THEN fanout_item_count END) AS fanout_items, + MAX(CASE WHEN slowest_rank = 1 THEN slowest_item END) AS slowest_item, + MAX(CASE WHEN slowest_rank = 1 THEN slowest_item_ms END) AS slowest_item_ms, + MAX(CASE WHEN slowest_rank = 1 THEN duration_ms END) AS slowest_run_duration_ms FROM ( -- #1855: rank each class of message newest-first so the two exemplar columns above can take @@ -1170,6 +1261,13 @@ END AS has_user_databases duration_ms, status, error_message, + -- #2472: projected here because this subquery enumerates its columns rather than + -- SELECT *-ing them, so an aggregate outside that names a column the inner query does + -- not carry fails at the STORE and nowhere earlier — no compiler, no text assertion and + -- no local build can see it. + fanout_item_count, + slowest_item, + slowest_item_ms, ROW_NUMBER() OVER ( PARTITION BY collector_name @@ -1183,7 +1281,18 @@ PARTITION BY collector_name ORDER BY (CASE WHEN status IN ('ERROR', 'PERMISSIONS') THEN error_message END) IS NULL, collection_time DESC, error_message DESC - ) AS error_rank + ) AS error_rank, + -- #2472: the window's dearest single ITEM, and with it the run that carried it. Ranked on + -- slowest_item_ms rather than duration_ms because the question is which database is + -- expensive, not which cycle was: a run whose total is the largest only because every + -- database was busy is exactly the shape a per-database override should NOT be aimed at. + ROW_NUMBER() OVER + ( + PARTITION BY collector_name + ORDER BY slowest_item_ms IS NULL, + slowest_item_ms DESC, + collection_time DESC + ) AS slowest_rank FROM v_collection_log WHERE server_id = $1 AND collection_time >= $2 @@ -1209,15 +1318,22 @@ public static async Task> GetCollectionHealthAsync( SuccessCount = reader.IsDBNull(2) ? 0 : Convert.ToInt64(reader.GetValue(2)), ErrorCount = reader.IsDBNull(3) ? 0 : Convert.ToInt64(reader.GetValue(3)), AvgDurationMs = reader.IsDBNull(4) ? 0 : Convert.ToDouble(reader.GetValue(4)), - LastSuccessTime = reader.IsDBNull(5) ? null : reader.GetDateTime(5), - LastRunTime = reader.IsDBNull(6) ? null : reader.GetDateTime(6), - LastError = reader.IsDBNull(7) ? null : reader.GetString(7), - LastErrorTime = reader.IsDBNull(8) ? null : reader.GetDateTime(8), - PermissionDeniedCount = reader.IsDBNull(9) ? 0 : Convert.ToInt64(reader.GetValue(9)), - YieldCount = reader.IsDBNull(10) ? 0 : Convert.ToInt64(reader.GetValue(10)), - LastNote = reader.IsDBNull(11) ? null : reader.GetString(11), - NoteCount = reader.IsDBNull(12) ? 0 : Convert.ToInt64(reader.GetValue(12)), - TargetHasUserDatabases = !reader.IsDBNull(13) && Convert.ToInt64(reader.GetValue(13)) != 0, + MaxDurationMs = reader.IsDBNull(5) ? 0 : Convert.ToDouble(reader.GetValue(5)), + P95DurationMs = reader.IsDBNull(6) ? 0 : Convert.ToDouble(reader.GetValue(6)), + LastSuccessTime = reader.IsDBNull(7) ? null : reader.GetDateTime(7), + LastRunTime = reader.IsDBNull(8) ? null : reader.GetDateTime(8), + LastError = reader.IsDBNull(9) ? null : reader.GetString(9), + LastErrorTime = reader.IsDBNull(10) ? null : reader.GetDateTime(10), + PermissionDeniedCount = reader.IsDBNull(11) ? 0 : Convert.ToInt64(reader.GetValue(11)), + YieldCount = reader.IsDBNull(12) ? 0 : Convert.ToInt64(reader.GetValue(12)), + LastNote = reader.IsDBNull(13) ? null : reader.GetString(13), + NoteCount = reader.IsDBNull(14) ? 0 : Convert.ToInt64(reader.GetValue(14)), + TargetHasUserDatabases = !reader.IsDBNull(15) && Convert.ToInt64(reader.GetValue(15)) != 0, + /* Ordinals are positional and these four were APPENDED (#2472) — never inserted. */ + FanoutItems = reader.IsDBNull(16) ? null : Convert.ToInt32(reader.GetValue(16)), + SlowestItem = reader.IsDBNull(17) ? null : reader.GetString(17), + SlowestItemMs = reader.IsDBNull(18) ? null : Convert.ToInt32(reader.GetValue(18)), + SlowestRunDurationMs = reader.IsDBNull(19) ? null : Convert.ToInt32(reader.GetValue(19)), }); } @@ -1290,6 +1406,337 @@ private static void AddWindow(NpgsqlCommand command, int serverId, DateTime star AddTimestamp(command, endUtc); } + /// + /// The raw per-run collection log for one server inside an explicit window, newest first, capped. + /// Bounded on BOTH sides rather than by a single now-relative lower bound, matching how the + /// viewer's Collection Log tab windows this read: a caller asking about a past incident wants the + /// rows from THEN, and an hours-back-from-now span cannot express that. + /// Reads v_collection_log, the same view the viewer uses, so the web dashboard and the + /// MCP surface cannot drift from what the desktop shows. $1 server_id, $2 window start, $3 window + /// end (naive UTC), $4 row cap. + /// + public const string CollectionLogSql = """ + SELECT + collector_name, + collection_time, + duration_ms, + sql_duration_ms, + duckdb_duration_ms, + rows_collected, + status, + error_message + FROM v_collection_log + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + ORDER BY collection_time DESC + LIMIT $4 + """; + + /// + /// Whether this server has EVER recorded a collector run, ignoring any window. + /// Exists so an empty log read can say which kind of nothing it found. "No runs in the last + /// 24 hours" is true both of a quiet window and of a server that has never collected, and those + /// need opposite responses from the caller -- widen the window, versus go find out why collection + /// is not running. LIMIT 1 with no ordering, so it stops at the first row rather than scanning. + /// + public const string HasAnyCollectionLogSql = """ + SELECT 1 + FROM v_collection_log + WHERE server_id = $1 + LIMIT 1 + """; + + /// Runs . + public static async Task HasAnyCollectionLogAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HasAnyCollectionLogSql); + AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } + + /// Runs . See it for the window semantics. + public static async Task> GetCollectionLogAsync( + NpgsqlDataSource postgres, + int serverId, + DateTime windowStartUtc, + DateTime windowEndUtc, + int maxRows, + CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(CollectionLogSql); + AddInt(command, serverId); + AddTimestamp(command, windowStartUtc); + AddTimestamp(command, windowEndUtc); + AddInt(command, maxRows); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new CollectionLogEntry( + reader.GetString(0), + reader.GetDateTime(1), + reader.IsDBNull(2) ? null : Convert.ToDouble(reader.GetValue(2)), + reader.IsDBNull(3) ? null : Convert.ToDouble(reader.GetValue(3)), + reader.IsDBNull(4) ? null : Convert.ToDouble(reader.GetValue(4)), + reader.IsDBNull(5) ? null : Convert.ToInt64(reader.GetValue(5)), + reader.IsDBNull(6) ? null : reader.GetString(6), + reader.IsDBNull(7) ? null : reader.GetString(7))); + } + + return rows; + } + + /// + /// Waiting-task total wait duration per wait type per collection, for one server over an explicit + /// window. The viewer's Current Waits reader verbatim. $1 server_id, $2 start, $3 end (naive UTC). + /// + public const string WaitingTaskTrendSql = """ + SELECT + collection_time, + wait_type, + CAST(SUM(wait_duration_ms) AS bigint) AS total_wait_ms + FROM waiting_tasks + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND wait_type IS NOT NULL + GROUP BY + collection_time, + wait_type + ORDER BY + collection_time, + wait_type + """; + + /// + /// Whether the waiting-task collector has EVER sampled this server, ignoring any window. + /// Separates an all-clear from missing data. Of the two, the wrong answer here is the + /// REASSURING one: "nothing was waiting" stops a caller looking, where "never collected" sends them + /// to check the collector. LIMIT 1, so it stops at the first row. + /// + public const string HasAnyWaitingTaskSampleSql = """ + SELECT 1 + FROM waiting_tasks + WHERE server_id = $1 + LIMIT 1 + """; + + /// Runs . + public static async Task HasAnyWaitingTaskSampleAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HasAnyWaitingTaskSampleSql); + AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } + + /// Runs . + public static async Task> GetWaitingTaskTrendAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(WaitingTaskTrendSql); + AddInt(command, serverId); + AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new WaitingTaskTrendRow( + reader.GetDateTime(0), + reader.IsDBNull(1) ? "" : reader.GetString(1), + reader.IsDBNull(2) ? 0 : reader.GetInt64(2))); + } + + return rows; + } + + /// + /// Blocked-session count per database per collection, for one server over an explicit window. + /// blocking_session_id > 0 is the blocked-ness test, so this counts sessions WAITING ON another + /// session rather than every waiting task. The optional database filter is kept from the viewer's read + /// rather than dropped for a simpler signature: on a busy instance one database usually owns the + /// blocking, and a series that cannot be narrowed to it answers a different question from the one the + /// viewer answers. $1 server_id, $2 start, $3 end (naive UTC), $4 database name or NULL. + /// + public const string BlockedSessionTrendSql = """ + SELECT + collection_time, + database_name, + COUNT(*) AS blocked_count + FROM waiting_tasks + WHERE server_id = $1 + AND blocking_session_id > 0 + AND collection_time >= $2 + AND collection_time <= $3 + AND database_name IS NOT NULL + AND ($4::text IS NULL OR database_name = $4) + GROUP BY + collection_time, + database_name + ORDER BY + collection_time, + database_name + """; + + /// Runs . + public static async Task> GetBlockedSessionTrendAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + string? databaseName = null, CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(BlockedSessionTrendSql); + AddInt(command, serverId); + AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); + command.Parameters.Add(new NpgsqlParameter + { + Value = string.IsNullOrWhiteSpace(databaseName) ? DBNull.Value : databaseName, + NpgsqlDbType = NpgsqlTypes.NpgsqlDbType.Text, + }); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new BlockedSessionTrendRow( + reader.GetDateTime(0), + reader.IsDBNull(1) ? "" : reader.GetString(1), + reader.IsDBNull(2) ? 0 : reader.GetInt64(2))); + } + + return rows; + } + + /// + /// Blocking-duration aggregate per minute, the viewer's Blocking Stats read verbatim. + /// XE blocked-process reports are the primary source and the DMV snapshot is the fallback, and + /// the fallback contributes ONLY when the XE source has no rows in the window at all. Mixing them + /// would double-count the same incident from two captures, so it is a fallback and never a union. + /// $1 server_id, $2 start, $3 end (naive UTC). + /// + public const string BlockingDurationStatsSql = """ + WITH bpr AS ( + SELECT + DATE_TRUNC('minute', event_time) AS bucket, + COUNT(*) AS event_count, + CAST(SUM(wait_time_ms) AS bigint) AS total_duration_ms, + MAX(wait_time_ms) AS max_duration_ms, + CAST(AVG(wait_time_ms) AS double precision) AS avg_duration_ms + FROM v_blocked_process_reports + WHERE server_id = $1 AND event_time >= $2 AND event_time <= $3 + GROUP BY DATE_TRUNC('minute', event_time) + ), + dmv AS ( + SELECT + DATE_TRUNC('minute', event_time) AS bucket, + COUNT(*) AS event_count, + CAST(SUM(wait_time_ms) AS bigint) AS total_duration_ms, + MAX(wait_time_ms) AS max_duration_ms, + CAST(AVG(wait_time_ms) AS double precision) AS avg_duration_ms + FROM v_dmv_blocking_snapshots + WHERE server_id = $1 AND event_time >= $2 AND event_time <= $3 + GROUP BY DATE_TRUNC('minute', event_time) + ) + SELECT bucket, event_count, total_duration_ms, max_duration_ms, avg_duration_ms FROM bpr + UNION ALL + SELECT bucket, event_count, total_duration_ms, max_duration_ms, avg_duration_ms FROM dmv WHERE NOT EXISTS (SELECT 1 FROM bpr) + ORDER BY bucket + """; + + public sealed record BlockingDurationStatsRow( + DateTime Time, long EventCount, long TotalDurationMs, long MaxDurationMs, double AvgDurationMs); + + /// + /// Whether ANY of the three capture paths behind the blocking-severity read has ever produced a row. + /// Each can be off independently: the XE blocked-process report needs its session running, the + /// DMV snapshot needs its collector enabled, and deadlock capture is separate from both. Probing one + /// would report "never captured" for a server capturing fine through another; probing neither would + /// let a silent capture gap read as a clean bill of health. + /// Deadlocks are in here because the verdict gates on the blocking series AND the deadlock + /// series both being empty. Probing only the two blocking sources would call a server 'genuinely + /// clear' on the strength of blocking capture alone, while deadlock capture had never run -- the + /// reassuring-wrong answer this probe exists to prevent, missed for the deadlock half. + /// + public const string HasAnyBlockingCaptureSql = """ + SELECT 1 + WHERE EXISTS (SELECT 1 FROM v_blocked_process_reports WHERE server_id = $1) + OR EXISTS (SELECT 1 FROM v_dmv_blocking_snapshots WHERE server_id = $1) + OR EXISTS (SELECT 1 FROM v_deadlocks WHERE server_id = $1) + """; + + /// Runs . + public static async Task HasAnyBlockingCaptureAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HasAnyBlockingCaptureSql); + AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } + + /// Runs . + public static async Task> GetBlockingDurationStatsAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(BlockingDurationStatsSql); + AddInt(command, serverId); + AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new BlockingDurationStatsRow( + reader.GetDateTime(0), + reader.IsDBNull(1) ? 0 : Convert.ToInt64(reader.GetValue(1)), + reader.IsDBNull(2) ? 0 : Convert.ToInt64(reader.GetValue(2)), + reader.IsDBNull(3) ? 0 : Convert.ToInt64(reader.GetValue(3)), + reader.IsDBNull(4) ? 0 : Convert.ToDouble(reader.GetValue(4)))); + } + + return rows; + } + + /// + /// The raw deadlock graphs in the window, for severity aggregation. + /// Windowed on collection_time but ORDERED and bucketed by deadlock_time -- the graph carries + /// when the deadlock happened, while collection_time is only when we picked it up, and the two differ + /// by up to a collection interval. $1 server_id, $2 start, $3 end (naive UTC). + /// + public const string DeadlockSeverityGraphsSql = """ + SELECT + deadlock_time, + deadlock_graph_xml + FROM v_deadlocks + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + ORDER BY deadlock_time + """; + + /// Runs , returning graphs for the shared aggregator. + public static async Task> GetDeadlockGraphsAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken = default) + { + var rows = new List<(DateTime? DeadlockTime, string? Xml)>(); + await using var command = postgres.CreateCommand(DeadlockSeverityGraphsSql); + AddInt(command, serverId); + AddTimestamp(command, startUtc); + AddTimestamp(command, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(( + reader.IsDBNull(0) ? null : reader.GetDateTime(0), + reader.IsDBNull(1) ? null : reader.GetString(1))); + } + + return rows; + } + private static void AddInt(NpgsqlCommand command, int value) => command.Parameters.Add(new NpgsqlParameter { TypedValue = value }); @@ -1325,6 +1772,23 @@ internal sealed class CollectorHealth public long SuccessCount { get; set; } public long ErrorCount { get; set; } public double AvgDurationMs { get; set; } + + /// + /// The single worst run in the window (#2460). A FACT, never a decision input: one pathological + /// cycle would otherwise make a collector read as permanently terrible for seven days. Its job is + /// to sit beside — when the two agree the tail is routine, and when the + /// max towers over the p95 the max was a one-off. + /// + public double MaxDurationMs { get; set; } + + /// + /// The 95th-percentile run in the window (#2460) — what a HEAVY run of this collector costs, as + /// opposed to what its runs cost on average. The number the sweep's peak-cycle arithmetic is built + /// from (via ), because a mean over a bimodal + /// collector describes neither of its populations. + /// + public double P95DurationMs { get; set; } + public DateTime? LastSuccessTime { get; set; } public DateTime? LastRunTime { get; set; } public string? LastError { get; set; } @@ -1352,6 +1816,48 @@ internal sealed class CollectorHealth /// public bool TargetHasUserDatabases { get; set; } + /* ── The per-database fan-out rollup (#2472) ───────────────────────────────────────────────────── + Four parts of ONE run — the window's dearest single item and the run that carried it. NULL on + every collector that does not fan out, which is 36 of the 41: the columns are only written by a + productive per-database run. + + They exist because the tail statistics above cannot answer this. MaxDurationMs and P95DurationMs + aggregate over RUNS, and each run is one blended row, so the two shapes an operator has to tell + apart — an even fan-out and one dominated by a single database — produce identical values in + both. Worse on this fleet: query_store's runs are 84% empty enumerations on a busy shard and + 100% on a quiet one, so its p95/avg ratio is already saturated by empty-versus-productive and + says nothing at all about which database is expensive. */ + + /// How many items that run fanned out over, or null when it did not fan out. The + /// denominator of : without it a slowest-item duration is a number with + /// no baseline, because "62 seconds" means something different across 2 databases and across 12. + public int? FanoutItems { get; set; } + + /// The dearest item in the window — a database name, for every fan-out that exists today. + public string? SlowestItem { get; set; } + + /// What that item cost, SQL plus storage. + public int? SlowestItemMs { get; set; } + + /// The whole run that item came from, so the share is against the number the operator sees + /// on the collection_log row rather than against a sum of item slices. + public int? SlowestRunDurationMs { get; set; } + + /// + /// The answer, as one number: 1.0 is a perfectly even fan-out and it rises with concentration. Eight + /// databases at 10.1s each gives 1.0; one at 62s beside seven at 2.7s gives 6.1 — the same 80,900 ms + /// run either way. Roughly 2.0 or above is the shape a per-database schedule override or a stagger + /// can actually target; near 1.0 says the cost is the fan-out's WIDTH and bounded parallelism is the + /// lever instead (#2468). + /// + /// Null when the collector does not fan out, and also when the run's duration is zero — a + /// ratio against nothing is not a smaller answer, it is a wrong one. + /// + public double? FanoutDominance => + FanoutItems is > 0 && SlowestItemMs.HasValue && SlowestRunDurationMs is > 0 + ? (double)SlowestItemMs.Value * FanoutItems.Value / SlowestRunDurationMs.Value + : null; + public double FailureRatePercent => TotalRuns > 0 ? (double)ErrorCount / TotalRuns * 100 : 0; public double HoursSinceLastSuccess => LastSuccessTime.HasValue diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingEngineCapability.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingEngineCapability.cs new file mode 100644 index 000000000..0d9670f10 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingEngineCapability.cs @@ -0,0 +1,156 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The store half of the #2511 engine-capability answer: reads this server's probed engine edition and, when +/// the collector serving a read cannot run on that engine, returns the not_collected envelope both +/// SKUs emit. The DECISION and the WORDS both come from +/// — nothing about which collectors are gated, and nothing about how +/// the gap is described, lives in this file or in its Lite twin (McpEngineCapability). +/// +/// Why not_collected rather than unavailable. The miss vocabulary already has the +/// right word: not_collected is "the input names something this server does not collect", which is +/// exactly true and final here. unavailable means "supported here, just not retrievable now", and it +/// sends an operator hunting for a collector to restart — which is the defect #2511 was filed about, not the +/// fix for it. +/// +/// Called only on the miss path. Every call site checks capability after its read came back +/// empty, never before it. That costs nothing in the common case, and — more importantly — a server whose +/// registry row says one engine while its collected rows say another (a re-registration, a restored +/// database) still gets its DATA rather than a confident explanation of why it cannot have any. +/// +internal static class DarlingEngineCapability +{ + /// + /// The two engine facts the registration upsert stamps on every connect + /// (): the probed SERVERPROPERTY('EngineEdition') and, since + /// V82 (#2530), the target's engine KIND. Exposed as a const so Darling.Tests can pin the dialect without + /// a live store. $1 server_id. + /// + /// ONE round trip for both, deliberately: they are read together on every miss, and two reads would + /// also make it possible to answer the kind axis from one server's row and the edition axis from a stale + /// copy of another's. + /// + public const string ServerEngineSql = @" +SELECT sql_engine_edition, engine_kind +FROM servers +WHERE server_id = $1"; + + /// + /// The not_collected envelope when cannot run on this server's + /// engine, or null when it can — in which case the caller falls through to its own + /// empty/unavailable miss, unchanged. + /// + /// A registry read that FAILS answers null, deliberately. This runs on a path that has already + /// found no data; turning a capability probe into a read error would replace one honest miss with a + /// worse one. + /// + public static async Task NotCollectedStatusAsync( + NpgsqlDataSource postgres, + int serverId, + string serverName, + string collectorName, + CancellationToken cancellationToken = default) + { + int engineEdition; + string? engineKind; + try + { + (engineEdition, engineKind) = await ReadServerEngineAsync(postgres, serverId, cancellationToken); + } + catch (Exception) + { + return null; + } + + var message = CollectorEngineCapability.NotCollectedMessage(serverName, engineEdition, engineKind, collectorName); + return message is null ? null : McpHelpers.Status("not_collected", message); + } + + /// + /// The server's probed engine edition and engine kind, defaulting to + /// and null when the registry has no + /// row or a NULL. + /// + /// The two NULLs mean different things and are both correct. A PostgreSQL target always has edition + /// 0 — SERVERPROPERTY does not exist there — so the edition is unknown for it permanently, and it + /// is the KIND that carries the fact. A NULL kind is a row no connect has stamped since V82 landed, which + /// makes no claim on that axis and leaves the edition axis answering exactly as it did before (#2530). + /// + private static async Task<(int EngineEdition, string? EngineKind)> ReadServerEngineAsync( + NpgsqlDataSource postgres, + int serverId, + CancellationToken cancellationToken) + { + await using var command = postgres.CreateCommand(ServerEngineSql); + DarlingMcpReadParameters.AddInt(command, serverId); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + + if (!await reader.ReadAsync(cancellationToken)) + { + return (CollectorEngineCapability.UnknownEngineEdition, null); + } + + var edition = reader.IsDBNull(0) ? CollectorEngineCapability.UnknownEngineEdition : reader.GetInt32(0); + var kind = reader.IsDBNull(1) ? null : reader.GetString(1); + return (edition, kind); + } + + /// The registry's PostgreSQL major for one server (V100, #2653). $1 server_id. + public const string PostgresMajorVersionSql = @" +SELECT postgres_major_version +FROM servers +WHERE server_id = $1"; + + /// + /// The target's probed PostgreSQL major, or null when the registry makes no claim — no row, a + /// server no connect has stamped since V100 landed, or a SQL Server target, where it is not a fact about + /// that server at all. + /// + /// Null is not an error and must not be rendered as a version. Callers use this to decide + /// whether they may state that a column is absent on this server's version; with no claim they say + /// nothing about the version instead of guessing, because a wrong version in an explanation is worse + /// than an unexplained NULL. + /// + /// A registry read that FAILS answers null for the same reason + /// does: this runs on a path that already has its data, and a + /// capability probe must never turn a good answer into a read error. + /// + public static async Task PostgresMajorVersionAsync( + NpgsqlDataSource postgres, + int serverId, + CancellationToken cancellationToken = default) + { + try + { + await using var command = postgres.CreateCommand(PostgresMajorVersionSql); + DarlingMcpReadParameters.AddInt(command, serverId); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + + if (!await reader.ReadAsync(cancellationToken) || reader.IsDBNull(0)) + { + return null; + } + + return reader.GetInt32(0); + } + catch (Exception) + { + return null; + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingFleetReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingFleetReader.cs index 85ddfe5b3..8b7737140 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingFleetReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingFleetReader.cs @@ -14,6 +14,7 @@ using System.Threading; using System.Threading.Tasks; using Npgsql; +using PerformanceMonitor.Collectors; using PerformanceMonitor.Common; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -60,6 +61,11 @@ internal static class DarlingFleetReader /// editions otherwise), the reliable per-server platform signal the composer's D4 auto-greying keys on; /// nullable when a server has not yet connected. /// + /// engine_kind (V82, #2530) is the other engine axis, and the one the edition cannot carry: + /// a PostgreSQL target has no SERVERPROPERTY, so it lands at edition 0 exactly like a SQL Server + /// that has never connected. Riding on the SAME registry row costs no extra round-trip, which is what + /// keeps this reader's bounded fan-out bounded. + /// /// The is_silenced column (#2031) is the SQL mirror of the Viewer's /// ViewerDataService.IsWholeServerSilence predicate — an enabled, unexpired mute rule scoped to the /// server (matched case-insensitively on the same COALESCE(display, storage) name the card shows, which is @@ -67,7 +73,7 @@ internal static class DarlingFleetReader /// web seat has no silence action; this exists so a dataless-quiet server and a silenced one stop looking /// identical on the fleet cards and to get_fleet_overview. $ none. public const string FleetServersSql = @" -SELECT s.server_id, COALESCE(s.display_name, s.server_name) AS display_name, s.server_name, s.sql_engine_edition, +SELECT s.server_id, COALESCE(s.display_name, s.server_name) AS display_name, s.server_name, s.sql_engine_edition, s.engine_kind, EXISTS ( SELECT 1 @@ -317,23 +323,14 @@ private static FleetServerCard BuildCard( FailedCollectorCount = collectors.Failing, }; - /* Freshness -> the card's connection state, exactly as the WPF card's ApplyFreshness. */ + /* Freshness -> the card's collection state, through the SAME mapping the WPF card and the sidebar + row use (#2473). It was a hand-written copy of ApplyFreshness that happened to agree; the copy on + the sidebar row happened not to, which is the argument for none of them writing it out. */ var freshness = ServerHealthClassifier.ClassifyFreshness(lastCollection, now); - bool? isOnline; - bool awaitingFirstCollection; - bool hasCollectorErrors; - if (freshness == ServerFreshness.NeverCollected) - { - isOnline = null; - awaitingFirstCollection = true; - hasCollectorErrors = false; - } - else - { - isOnline = freshness != ServerFreshness.Offline; - awaitingFirstCollection = false; - hasCollectorErrors = freshness == ServerFreshness.Stale; - } + var flags = ServerCollectionStatusRules.FlagsFor(freshness); + var isOnline = flags.IsOnline; + var awaitingFirstCollection = flags.AwaitingFirstCollection; + var hasCollectorErrors = flags.HasCollectorErrors; var overall = ServerHealthClassifier.OverallMetricSeverity(metrics); var band = ServerHealthClassifier.ClassifyBand(isOnline, awaitingFirstCollection, hasCollectorErrors, overall); @@ -343,12 +340,19 @@ private static FleetServerCard BuildCard( deliberately not surfaced. */ var (isAzureSqlDb, isAzureManagedInstance) = ClassifyPlatform(server.EngineEdition); + /* Per-server target ENGINE (#2530), the axis the platform flags above cannot express: they are all + derived from a SQL Server SERVERPROPERTY, which a PostgreSQL target does not have. */ + var (isPostgres, isAurora) = ClassifyEngineKind(server.EngineKind); + return new FleetServerCard { ServerId = server.ServerId, DisplayName = server.DisplayName, ServerName = server.ServerName, EngineEdition = server.EngineEdition, + EngineKind = server.EngineKind, + IsPostgres = isPostgres, + IsAurora = isAurora, IsAzureSqlDb = isAzureSqlDb, IsAzureManagedInstance = isAzureManagedInstance, IsSilenced = server.IsSilenced, @@ -478,7 +482,9 @@ private static string BuildReason(FleetServerCard c) if (c.AwaitingFirstCollection) { - return "Awaiting first collection"; + /* The word itself, not a copy of it — this was one of five spellings of the phrase across four + files, which is the duplication #2473's pin now forbids. */ + return ServerCollectionStatus.AwaitingFirstCollection.Word(); } var parts = new List(); @@ -523,13 +529,11 @@ private static string BuildReason(FleetServerCard c) return parts.Count > 0 ? string.Join(", ", parts) : "Needs attention"; } - private static string StatusLabel(bool? isOnline, bool awaitingFirstCollection, bool hasCollectorErrors) => isOnline switch - { - true when hasCollectorErrors => "Warning", - true => "Online", - false => "Offline", - _ => awaitingFirstCollection ? "Awaiting first collection" : "Unknown", - }; + /// The card's status word. Delegates to the one ladder every Darling surface renders (#2473): + /// this file's own copy agreed with the WPF card, but the WPF sidebar row's copy did not, and three + /// agreeing copies plus one that does not is still four places where the answer is decided. + private static string StatusLabel(bool? isOnline, bool awaitingFirstCollection, bool hasCollectorErrors) => + ServerCollectionStatusRules.Classify(isOnline, hasCollectorErrors, awaitingFirstCollection).Word(); /// /// Classifies a server's raw SERVERPROPERTY('EngineEdition') into the RELIABLE per-server platform flags the @@ -545,6 +549,20 @@ private static string BuildReason(FleetServerCard c) internal static (bool IsAzureSqlDb, bool IsAzureManagedInstance) ClassifyPlatform(int? engineEdition) => (engineEdition == 5, engineEdition == 8); + /// + /// Classifies the stored engine-kind token (V82, #2530) into the two booleans a browser actually branches + /// on, so no consumer has to know the vocabulary's spelling. The raw token still rides on the card beside + /// them — a UI that wants to LABEL the engine needs the word, and a UI that wants to choose a tab set + /// needs the boolean. + /// + /// null — a pre-V82 row, or a server that has not connected since the rung landed — is + /// neither, so a card with no signal renders exactly as it did before this column existed rather than + /// claiming SQL Server on the strength of an absence. Same discipline as + /// 's null edition. + /// + internal static (bool IsPostgres, bool IsAurora) ClassifyEngineKind(string? engineKind) => + (MonitoredEngineKind.IsPostgres(engineKind), MonitoredEngineKind.IsAurora(engineKind)); + /* ─────────────────────────── per-query readers ─────────────────────────── */ private static async Task> ReadServersAsync(NpgsqlDataSource postgres, CancellationToken cancellationToken) @@ -559,7 +577,8 @@ private static async Task> ReadServersAsync(NpgsqlDataSourc reader.IsDBNull(1) ? "" : reader.GetString(1), reader.IsDBNull(2) ? "" : reader.GetString(2), reader.IsDBNull(3) ? null : reader.GetInt32(3), - !reader.IsDBNull(4) && reader.GetBoolean(4))); + reader.IsDBNull(4) ? null : reader.GetString(4), + !reader.IsDBNull(5) && reader.GetBoolean(5))); } return rows; @@ -765,7 +784,7 @@ private static void AddTimestamp(NpgsqlCommand command, DateTime value) => /* ─────────────────────────── raw-read carriers (internal) ─────────────────────────── */ - private readonly record struct FleetServerRow(int ServerId, string DisplayName, string ServerName, int? EngineEdition, bool IsSilenced); + private readonly record struct FleetServerRow(int ServerId, string DisplayName, string ServerName, int? EngineEdition, string? EngineKind, bool IsSilenced); private readonly record struct CpuRow(double? SqlCpu, double? OtherCpu); private readonly record struct MemoryRow(double? MemoryMb, double? BufferPoolMb); private readonly record struct MemoryPressureRow(long WaiterCount, long TimeoutCount, long ForcedCount, double? GrantedMemoryMb); @@ -791,6 +810,56 @@ public sealed class FleetServerCard /// signal the composer's D4 measure auto-greying matches a measure's appliesTo against. [JsonPropertyName("engine_edition")] public int? EngineEdition { get; init; } + /// The target's engine KIND as the registry recorded it on its last connect (#2530) — + /// sqlserver, postgres, or aurora-postgres; null when no connect has stamped it (a + /// pre-V82 store, or a server that has never connected). + /// + /// This is the discriminator engine_edition cannot supply and never could: a PostgreSQL + /// target has no SERVERPROPERTY, so it lands with edition 0, which is also what a SQL Server that + /// has never connected lands with. Every surface that wants to show PostgreSQL panels to PostgreSQL + /// targets and SQL Server tabs to SQL Server targets branches on THIS. + [JsonPropertyName("engine_kind")] public string? EngineKind { get; init; } + + /// How reads in a sentence — "SQL Server", "PostgreSQL", "Aurora + /// PostgreSQL" — or null when the store has made no claim. + /// + /// On the card because is deliberately the + /// single copy of those words, and a browser mapping three tokens to three strings itself would be a + /// second table in a language the first cannot be shared with. The web server page had exactly that for + /// the length of one review round. + /// + /// DERIVED rather than assigned, so it cannot be forgotten: every card is built by an object + /// initializer, and a settable field would be null on any card whose builder did not think of it — + /// including one added later on a path nobody re-reads. + /// + /// Three answers, and the third is the interesting one. An ABSENT kind is null: a surface has + /// nothing to say about a server whose engine was never stamped, and no badge is better than a badge + /// describing the store's silence as a property of the server. A RECOGNISED token gets + /// 's words. A token this build has never heard of — + /// a store written by a NEWER build — gets the token back verbatim, NOT the describer's + /// "an unrecognised engine": that phrase is worded to sit mid-sentence in the capability messages, and as + /// a label beside "SQL Server" and "Aurora PostgreSQL" it reads as the wrong part of speech. The raw token + /// is also the more useful of the two, being the string an operator would search their own store for. It + /// is deliberately not mapped onto a default, which is the whole reason the describer refuses to guess in + /// the first place. + /// + [JsonPropertyName("engine_description")] + public string? EngineDescription => + string.IsNullOrWhiteSpace(EngineKind) ? null + : MonitoredEngineKind.IsKnown(EngineKind) ? MonitoredEngineKind.DescribeEngineKind(EngineKind) + : EngineKind.Trim(); + + /// True when this server is PostgreSQL (Aurora or stock) — derived from + /// , so a consumer never has to know the token vocabulary. False when the kind is + /// unknown: absence of a claim, not a claim of SQL Server. + [JsonPropertyName("is_postgres")] public bool IsPostgres { get; init; } + + /// True when this server is Amazon Aurora PostgreSQL specifically — derived from + /// . Separate from because a large proprietary surface + /// (the aurora_stat_* functions) exists only there, so a panel fed by an Aurora-only collector has + /// to be able to tell the two apart. + [JsonPropertyName("is_aurora")] public bool IsAurora { get; init; } + /// True when this server is Azure SQL Database (engine edition 5) — reliable, derived from /// . [JsonPropertyName("is_azure_sql_db")] public bool IsAzureSqlDb { get; init; } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAgTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAgTools.cs index 62e79cc51..d7dd7a1fa 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAgTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAgTools.cs @@ -64,6 +64,19 @@ would silently narrow the fleet view on a one-server store. */ if (result.AvailabilityGroupCount == 0) { + /* #2511: only when the caller SCOPED to one server, because engine edition is per server and + the fleet-wide read has no single engine to speak for. A fleet-wide miss keeps the general + sentence below, which already says the AG collectors do not run on Azure SQL Database. */ + if (serverIdFilter is int scopedServerId && resolvedName is not null) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, scopedServerId, resolvedName, "ag_replica_states"); + if (gated != null) + { + return gated; + } + } + /* The message names the RESOLVED storage name, not the caller's spelling — one condition, and a partial or differently-cased argument reads back as the server it actually matched. */ return McpHelpers.Status( diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAlertTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAlertTools.cs index e5f5406d8..0175259de 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAlertTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpAlertTools.cs @@ -68,9 +68,10 @@ public static async Task GetAlertHistory( NpgsqlDataSource postgres, [Description("Server name or display name. Omit to return alerts across all servers (the fleet default).")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, - [Description("Maximum rows. Default 50.")] int limit = 50) + [Description("Maximum rows. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { - var hoursError = McpHelpers.ValidateHoursBack(hours_back); + var hoursError = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (hoursError != null) return hoursError; var limitError = McpHelpers.ValidateTop(limit); if (limitError != null) return limitError; @@ -89,8 +90,8 @@ public static async Task GetAlertHistory( try { - var since = DateTime.UtcNow.AddHours(-hours_back); - var rows = await DarlingAlertReader.GetAlertHistoryAsync(postgres, since, serverId, limit); + var since = windowEnd.AddHours(-hours_back); + var rows = await DarlingAlertReader.GetAlertHistoryAsync(postgres, since, windowEnd, serverId, limit); if (rows.Count == 0) return McpHelpers.Status("empty", "No alerts found in the specified time range."); @@ -123,7 +124,7 @@ public static async Task GetAlertHistory( } } - [McpServerTool(Name = "get_alert_settings"), Description("Gets the current alert configuration the service is using: which alerts are enabled and their thresholds (CPU, blocking, deadlocks, poison waits, long-running queries/jobs, tempdb, low disk, failed jobs, database state), the cooldown, excluded databases, the deadlock/blocking delivery mode, and the scheduled-analysis cadence. This is the single global settings row the service hot-swaps in. SMTP/webhook delivery credentials are managed separately and are not reported here.")] + [McpServerTool(Name = "get_alert_settings"), Description("Gets the current alert configuration the service is using: which alerts are enabled and their thresholds (CPU, blocking, deadlocks, poison waits, long-running queries/jobs, tempdb, low disk, failed jobs, database state, Availability Group health, connection loss), the cooldown, excluded databases, the deadlock/blocking delivery mode, and the scheduled-analysis cadence. This is the single global settings row the service hot-swaps in. SMTP/webhook delivery credentials are managed separately and are not reported here.")] public static async Task GetAlertSettings( NpgsqlDataSource postgres) { @@ -144,11 +145,26 @@ public static async Task GetAlertSettings( } /// The nested JSON shape get_alert_settings returns AND update_alert_settings echoes back — the same - /// field names update_alert_settings accepts on the way in, so a read → modify → write round-trips. + /// field names update_alert_settings accepts on the way in, so a read → modify → write round-trips. + /// + /// That last clause is an INVARIANT, not a habit: #2417 found it broken in both directions at once + /// (six columns read and emitted to nobody, one key emitted the writer refused). It is now asserted as a + /// set equality between this payload's keys and the columns AlertSettingsSelectSql reads — + /// EveryColumnRead_IsEmittedByThePayload_AndAcceptedByTheWriter. Adding a key here without the + /// matching arm in (or the reverse) fails that test rather than + /// shipping. private static object BuildAlertSettingsPayload(DarlingAlertReader.AlertSettingsReadRow s) => new { alerts_enabled = s.Enabled, notify_connection_changes = s.NotifyConnectionChanges, + /* #2417: the connection family's two sub-settings, kept TOP-LEVEL beside their master switch + rather than folded into a `connection` group. The master shipped as a top-level key and Lite + emits it there too, so a group could only ever hold two of the three -- splitting one family + across two levels of the document is worse for a reader than two extra top-level keys. The + spellings are the store's own column names, which are also the keys Lite already reads out of + settings.json (App.LoadAlertSettings), so Lite's payload can adopt them verbatim. */ + notify_connection_down_at_startup = s.NotifyConnectionDownAtStartup, + connection_refire_minutes = s.ConnectionRefireMinutes, cpu = new { enabled = s.CpuEnabled, threshold_percent = s.CpuThresholdPercent, mode = s.CpuMode }, blocking = new { @@ -191,9 +207,34 @@ public static async Task GetAlertSettings( store_job_cadence_warn_percent = s.StoreJobCadenceWarnPercent }, pvs = new { enabled = s.PvsEnabled, threshold_percent = s.PvsThresholdPercent, floor_gb = s.PvsFloorGb }, + /* #2391: #2349's knobs reached 3.5.0 with the store plane only, so an alert that ships OFF could + be enabled only by UPDATEing config_alert_settings by hand. Reported by @gotqn. */ + file_growth = new + { + enabled = s.FileGrowthEnabled, + rise_mb = s.FileGrowthRiseMb, + volume_percent = s.FileGrowthVolumePercent, + lookback_minutes = s.FileGrowthLookbackMinutes + }, long_running_job = new { enabled = s.LongRunningJobEnabled, multiplier = s.LongRunningJobMultiplier }, failed_job = new { enabled = s.FailedJobEnabled, lookback_minutes = s.FailedJobLookbackMinutes }, database_state = new { enabled = s.DatabaseStateEnabled }, + /* #2417: the AG family, read out of the store since V35/V37 and emitted to nobody until now -- so + an agent asked why nothing alerted when a replica fell behind could not see the lag threshold, + could not see whether AG notification was on at all, and had no way to tell "configured not to + alert" from "failed to alert". Grouped like every other alert family, and the master switch is + named `enabled` for the same reason cpu/blocking/pvs/database_state are: inside this payload + `enabled` always means "this alert family is on", and notify_ag_health IS that switch rather + than a second opt-in behind one. The thresholds take the house _threshold_ spelling + (blocking.wait_threshold_seconds, poison_wait.threshold_ms, low_disk.threshold_gb) instead of + transliterating the columns' older _alert_ suffix, which appears nowhere else on the wire. */ + ag = new + { + enabled = s.NotifyAgHealth, + lag_threshold_seconds = s.AgLagAlertSeconds, + redo_queue_threshold_kb = s.AgRedoQueueAlertKb, + disconnect_refire_minutes = s.AgDisconnectRefireMinutes + }, cooldown_minutes = s.CooldownMinutes, excluded_databases = s.ExcludedDatabases, delivery = new { mode = s.DeliveryMode, per_event_max = s.PerEventMax }, @@ -215,11 +256,32 @@ public static async Task GetMuteRules( { try { - var rules = (await new PgMuteRuleStore(postgres).LoadAllAsync()).AsEnumerable(); + var all = await new PgMuteRuleStore(postgres).LoadAllAsync(); + var rules = all.AsEnumerable(); if (enabled_only) rules = rules.Where(r => r.Enabled && (r.ExpiresAtUtc == null || r.ExpiresAtUtc > DateTime.UtcNow)); var list = rules.ToList(); + + if (list.Count == 0) + { + /* + This tool exists so an agent can tell a genuinely healthy-quiet server from one whose + alerts are being suppressed, and an empty array answered that question with silence. + Both kinds of empty are true negatives -- nothing is being suppressed either way, which + is why both are "empty" rather than one being "unavailable" -- but they are not the same + fact. No rule has ever been written is a different state from five rules that all lapsed, + and the second one is a mute somebody INTENDED that is no longer in force. + */ + var configured = all.Count; + return McpHelpers.Status( + "empty", + configured == 0 + ? "No mute rules are configured for this store, so no alert is being suppressed anywhere — a quiet alert history is genuine rather than muted." + : $"No mute rules are in force right now: {configured} rule(s) exist but every one of them is disabled or expired, so nothing is being suppressed. Pass enabled_only=false to list them — this is a lapsed mute, not an absent one.", + new { enabled_only, configured_count = configured, excluded_by_filter = configured - list.Count }); + } + return JsonSerializer.Serialize(new { total_count = list.Count, @@ -462,6 +524,26 @@ void AddInt(string column, JsonNode? node, string field, int min, int max) } } + /* ag_redo_queue_alert_kb is the one bigint on this row, and DarlingAlertSettings clamps it in + long arithmetic. Binding it through AddInt would work today only because the ceiling happens to + fit in an int; typing it here means a raised ceiling stays writable instead of silently + rejecting every value above int.MaxValue as "not an integer". */ + void AddLong(string column, JsonNode? node, string field, long min, long max) + { + if (error != null) return; + if (node is JsonValue v && v.TryGetValue(out var l)) + { + if (l < min || l > max) + error = $"'{field}' must be an integer between {min} and {max}."; + else + updates.Add((column, new NpgsqlParameter { TypedValue = l })); + } + else + { + error = $"'{field}' must be an integer between {min} and {max}."; + } + } + void AddDouble(string column, JsonNode? node, string field, double min, double max) { if (error != null) return; @@ -547,6 +629,11 @@ void Group(JsonNode? node, string group, Action handleKey) { case "alerts_enabled": AddBool("enabled", prop.Value, "alerts_enabled"); break; case "notify_connection_changes": AddBool("notify_connection_changes", prop.Value, "notify_connection_changes"); break; + /* #2417: bounds mirror DarlingAlertSettings' clamps EXACTLY -- Clamp(0, 1440) on the + refire, nothing to clamp on the at-startup opt-in. Zero is IN range because 0 is the + shipped configuration (one alert per outage, no re-fire), not an invalid one. */ + case "notify_connection_down_at_startup": AddBool("notify_connection_down_at_startup", prop.Value, "notify_connection_down_at_startup"); break; + case "connection_refire_minutes": AddInt("connection_refire_minutes", prop.Value, "connection_refire_minutes", 0, 1440); break; case "cooldown_minutes": AddInt("cooldown_minutes", prop.Value, "cooldown_minutes", 1, 120); break; case "excluded_databases": AddStringArray("excluded_databases", prop.Value, "excluded_databases"); break; @@ -570,6 +657,13 @@ void Group(JsonNode? node, string group, Action handleKey) { case "enabled": AddBool("blocking_enabled", n, "blocking.enabled"); break; case "count_threshold": AddInt("blocking_count_threshold", n, "blocking.count_threshold", 1, int.MaxValue); break; + /* #2417: get_alert_settings has emitted this key since #1839 and the writer + never took it, so handing a whole read payload back -- the round trip this + tool's own description tells the caller to perform -- was rejected with + "Unknown field 'blocking.wait_threshold_seconds'", naming a field the caller + did not choose to send. The bound is the engine's Math.Max(0, ...), so 0 + keeps disabling the second gate rather than becoming invalid. */ + case "wait_threshold_seconds": AddInt("blocking_wait_seconds_threshold", n, "blocking.wait_threshold_seconds", 0, int.MaxValue); break; default: error = $"Unknown field 'blocking.{k}'."; break; } }); @@ -675,6 +769,25 @@ the value the sweep uses. */ }); break; + /* #2391: bounds mirror DarlingAlertSettings' clamps EXACTLY — Max(0) on the rise, + [0,100] on the volume percent, [5,1440] on the lookback. If these drift apart the tool + accepts a value the engine then silently rewrites, which reads as the setting not + sticking. Zero on either gate disables that gate rather than being invalid (#2349), + which is why the rise floor is 0 and not 1. */ + case "file_growth": + Group(prop.Value, "file_growth", (k, n) => + { + switch (k) + { + case "enabled": AddBool("file_growth_enabled", n, "file_growth.enabled"); break; + case "rise_mb": AddInt("file_growth_rise_mb", n, "file_growth.rise_mb", 0, int.MaxValue); break; + case "volume_percent": AddInt("file_growth_volume_percent", n, "file_growth.volume_percent", 0, 100); break; + case "lookback_minutes": AddInt("file_growth_lookback_minutes", n, "file_growth.lookback_minutes", 5, 1440); break; + default: error = $"Unknown field 'file_growth.{k}'."; break; + } + }); + break; + case "long_running_job": Group(prop.Value, "long_running_job", (k, n) => { @@ -710,6 +823,24 @@ the value the sweep uses. */ }); break; + /* #2417: bounds mirror DarlingAlertSettings' clamps EXACTLY -- Clamp(lag, 0, 86400), + Clamp(redo, 0L, 1073741824L), Clamp(refire, 0, 1440). Zero is IN range on all three + and means OFF for that gate (the redo queue SHIPS at 0, because a healthy queue size + is workload-specific), so a floor of 1 would remove the shipped configuration. */ + case "ag": + Group(prop.Value, "ag", (k, n) => + { + switch (k) + { + case "enabled": AddBool("notify_ag_health", n, "ag.enabled"); break; + case "lag_threshold_seconds": AddInt("ag_lag_alert_seconds", n, "ag.lag_threshold_seconds", 0, 86400); break; + case "redo_queue_threshold_kb": AddLong("ag_redo_queue_alert_kb", n, "ag.redo_queue_threshold_kb", 0L, 1073741824L); break; + case "disconnect_refire_minutes": AddInt("ag_disconnect_refire_minutes", n, "ag.disconnect_refire_minutes", 0, 1440); break; + default: error = $"Unknown field 'ag.{k}'."; break; + } + }); + break; + case "delivery": Group(prop.Value, "delivery", (k, n) => { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpBlockingTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpBlockingTools.cs index 360a33d51..f9a36d017 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpBlockingTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpBlockingTools.cs @@ -7,6 +7,7 @@ */ using System; +using System.Collections.Generic; using System.ComponentModel; using System.Linq; using System.Text.Json; @@ -49,23 +50,25 @@ public static async Task GetBlocking( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, [Description("Maximum rows. Default 30.")] int limit = 30, - [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null) + [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveWithFingerprintNameAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingBlockingReader.GetRecentBlockedProcessReportsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No blocking events found in the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "blocked_process_report") + ?? McpHelpers.Status("empty", "No blocking events found in the specified time range."); /* #2159: fingerprint the WHOLE window, then filter, then cap. Capping first would let `limit` discard the very incident the key names — the caller asked for one specific incident, not for @@ -151,23 +154,30 @@ public static async Task GetDeadlocks( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, [Description("Maximum rows. Default 20.")] int limit = 20, - [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null) + [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveWithFingerprintNameAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingBlockingReader.GetRecentDeadlocksAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No deadlocks found in the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "deadlocks") + /* #2546: capability first (permanent), then the runtime precondition (fixable), then the + read's own miss. A deadlock capture whose XE session is gone records SESSION_MISSING + and then returns zero rows forever, which is byte-identical to a server that simply + did not deadlock — the one answer nobody should be given without being told. */ + ?? await DarlingRuntimePrecondition.StatusAsync(postgres, resolved.ServerId, resolved.ServerName, "deadlocks") + ?? McpHelpers.Status("empty", "No deadlocks found in the specified time range."); /* #2159: see get_blocking — fingerprint the window, filter, then cap. */ var examined = rows.Count; @@ -218,19 +228,20 @@ public static async Task GetDeadlockDetail( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, [Description("Maximum deadlocks to return. Default 5.")] int limit = 5, - [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null) + [Description("Optional #1140 alert fingerprint (the alert's Dedup Key). When supplied, returns only the incident with that key — paste it straight from an alert or ticket instead of scanning the window. The key is scoped to the server's display name and the incident's involved objects.")] string? dedup_key = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveWithFingerprintNameAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingBlockingReader.GetRecentDeadlocksAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); @@ -239,7 +250,8 @@ public static async Task GetDeadlockDetail( consume one of the `limit` slots the caller wanted spent on real graphs. */ var candidates = rows.Where(r => r.HasDeadlockXml).ToList(); if (candidates.Count == 0) - return McpHelpers.Status("empty", "No deadlock XML available in the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "deadlocks") + ?? McpHelpers.Status("empty", "No deadlock XML available in the specified time range."); var examined = candidates.Count; var keys = DarlingIncidentFingerprint.DeadlockKeys( @@ -287,24 +299,29 @@ public static async Task GetBlockedProcessXml( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, - [Description("Maximum reports to return. Default 5.")] int limit = 5) + [Description("Maximum reports to return. Default 5.")] int limit = 5, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingBlockingReader.GetRecentBlockedProcessReportsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); var withXml = rows.Where(r => r.HasReportXml).Take(limit).ToList(); if (withXml.Count == 0) - return McpHelpers.Status("empty", "No blocked process report XML available in the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "blocked_process_report") + /* #2546: same order and same reason as get_deadlocks — a blocked-process capture whose + session is gone is indistinguishable here from a server that never blocked. */ + ?? await DarlingRuntimePrecondition.StatusAsync(postgres, resolved.ServerId, resolved.ServerName, "blocked_process_report") + ?? McpHelpers.Status("empty", "No blocked process report XML available in the specified time range."); var result = withXml.Select(r => new { @@ -333,19 +350,44 @@ public static async Task GetBlockedProcessXml( public static async Task GetBlockingTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; + var start = now.AddHours(-hours_back); var points = await DarlingBlockingTrendReader.GetBlockingTrendAsync( - postgres, resolved.ServerId, now.AddHours(-hours_back), now); + postgres, resolved.ServerId, start, now); + + if (points.Count == 0) + { + /* + An empty trend is two facts, and the WRONG one is the reassuring one. "No blocking" + reads as an all-clear and a caller who believes it stops looking; "nothing collected" + means nothing at all is known about the window. The stored tables cannot tell them + apart -- both are an absence of rows in an EDGE table -- so the denominator has to come + from collection_log, which records a SUCCESS with zero rows for a collector that ran + and saw nothing. Probed only here, on the path that already found nothing. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "blocked_process_report"); + if (gated != null) + { + return gated; + } + + var captures = await DarlingBlockingTrendReader.GetBlockingCaptureCountsAsync( + postgres, resolved.ServerId, start, now); + return await EmptyTrend( + "blocking", resolved.ServerName, hours_back, captures, + () => DarlingBlockingTrendReader.HasAnyBlockingCollectorRunAsync(postgres, resolved.ServerId)); + } return JsonSerializer.Serialize(new { @@ -364,19 +406,40 @@ public static async Task GetBlockingTrend( public static async Task GetDeadlockTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; + var start = now.AddHours(-hours_back); var points = await DarlingBlockingTrendReader.GetDeadlockTrendAsync( - postgres, resolved.ServerId, now.AddHours(-hours_back), now); + postgres, resolved.ServerId, start, now); + + if (points.Count == 0) + { + /* Same two facts as the blocking trend above, same denominator, same reason. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "deadlocks"); + if (gated != null) + { + return gated; + } + + var captures = await DarlingBlockingTrendReader.GetDeadlockCaptureCountsAsync( + postgres, resolved.ServerId, start, now); + return await EmptyTrend( + /* SINGULAR: the subject lands in "No {subject} was recorded", and "no deadlocks + was recorded" is not a sentence. It also reads correctly in the other two, + where it modifies the collector rather than the event. */ + "deadlock", resolved.ServerName, hours_back, captures, + () => DarlingBlockingTrendReader.HasAnyDeadlockCollectorRunAsync(postgres, resolved.ServerId)); + } return JsonSerializer.Serialize(new { @@ -390,4 +453,131 @@ public static async Task GetDeadlockTrend( return McpHelpers.FormatError("get_deadlock_trend", ex); } } + + [McpServerTool(Name = "get_lock_wait_trend"), Description("Gets the AGGREGATE lock-wait rate over time for a server: every LCK% wait type's wait milliseconds per second at each collection, the viewer's Blocking Trends lock-wait chart. get_wait_trend charts ONE named wait type and get_blocking_trend counts incidents; this is the whole lock family at once, as a rate rather than a count. Use it to see whether a server's lock pressure is rising when no single wait type dominates, to tell a few long blocks from constant low-grade contention, and to pick which LCK type to hand to get_wait_trend next. The rate is delta wait time divided by the seconds since the previous collection, so it is comparable across servers collecting on different cadences.")] + public static async Task GetLockWaitTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + var points = await DarlingBlockingTrendReader.GetLockWaitTrendAsync( + postgres, resolved.ServerId, start, end); + + if (points.Count == 0) + { + /* + The denominator here is the DATA, not collection_log — the opposite of the two trends + above, and deliberately so. Those read EDGE tables, where a row exists only because + something went wrong, so a healthy server has none and a data probe would report it as + uncollected. wait_stats is PERIODIC: the collector writes a row every cycle for every + wait type it observes, whatever the server is doing, so the presence of ANY wait sample + is proof somebody looked. Probing v_wait_stats — the SAME relation this read walks, so it + can never report "collected" for rows the read cannot see — is both cheaper and more + precise than asking collection_log which collector ran. + + The LCK% filter is deliberately NOT applied to the probe. A server that has collected + wait stats for months and never taken a lock wait is the all-clear this branch is for, + and filtering the probe the same way would call that server uncollected. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "wait_stats"); + if (gated != null) + { + return gated; + } + + return await DarlingDataReader.HasAnyWaitStatAsync(postgres, resolved.ServerId) + ? McpHelpers.Status( + "empty", + $"No lock waits recorded for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS collected wait stats before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.") + : McpHelpers.Status( + "unavailable", + $"No wait stats have EVER been recorded for {resolved.ServerName}, so this is NOT a report of a server without lock contention — nothing has been stored for it at all. Delta wait stats need a SECOND collection cycle before the first row exists, so on a newly added server this clears itself; otherwise check that collection is running and that the server is enabled."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + /* + Rows, not a pre-pivoted series per wait type. The caller decides whether to sum the + family or chart the members, and a pivot here would have to pick a top-N and silently + drop the rest — on a read whose premise is that no single LCK type dominates. + */ + trend = points.Select(p => new + { + collection_time = p.CollectionTime.ToString("o"), + wait_type = p.WaitType, + wait_time_ms_per_second = Math.Round(p.WaitTimeMsPerSecond, 3), + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_lock_wait_trend", ex); + } + } + + /// + /// The empty answer both per-minute trends give, and the whole point of #2485: trend: [] is the + /// same bytes on a server that had no blocking and on one that collected nothing, and an agent holding + /// only the JSON cannot tell a clean bill of health from a hole in coverage. + /// + /// Two statuses, picked by the denominator rather than by the edge rows. empty means captures + /// ran in this window and none of them saw the event -- a real all-clear, bounded by the sampling + /// interval. unavailable means no capture ran, so the window says nothing either way; the + /// existence probe then separates a server that has never collected this at all from one with a GAP, + /// because "check that collection is running" and "widen the window" are different next moves. + /// + /// carries the per-collector run counts. A count of three events means + /// something different in a window of 60 captures than in a window of 4, and the caller cannot supply + /// that number itself. Lite's twin returns the SAME sentences word for word. + /// + private static async Task EmptyTrend( + /* SINGULAR ("blocking", "deadlock"): it is the subject of "No {subject} was recorded". */ + string subject, + string serverName, + int hoursBack, + List captures, + Func> hasEverCapturedAsync) + { + var captureCount = captures.Sum(c => c.Runs); + var hints = new + { + server = serverName, + hours_back = hoursBack, + capture_count = captureCount, + captures = captures.Select(c => new + { + collector = c.CollectorName, + runs = c.Runs, + first_run_at = c.FirstRunAt?.ToString("o"), + last_run_at = c.LastRunAt?.ToString("o"), + }), + }; + + if (captureCount > 0) + return McpHelpers.Status( + "empty", + $"No {subject} was recorded for {serverName} in the last {hoursBack} hour(s). {captureCount} collector run(s) DID execute over this window, so this is a genuine all-clear rather than missing data — see hints.captures for which collectors ran and when.", + hints); + + var everCaptured = await hasEverCapturedAsync(); + return McpHelpers.Status( + "unavailable", + everCaptured + ? $"No {subject} collector runs are recorded for {serverName} in the last {hoursBack} hour(s), so this is NOT an all-clear — nothing was captured and the window says nothing either way. Collection HAS run for this server outside the window, so this is a gap rather than a dead collector: widen hours_back, or use get_collection_health to find where it stopped." + : $"No {subject} collector runs have EVER been recorded for {serverName}, so this is NOT an all-clear — there is nothing to read. Check that collection is running for this server before concluding it was quiet.", + hints); + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigHistoryTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigHistoryTools.cs index f9a620cdb..f48837065 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigHistoryTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigHistoryTools.cs @@ -44,23 +44,27 @@ public sealed class DarlingMcpConfigHistoryTools public static async Task GetServerConfigChanges( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168) + [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var windowStart = NaiveUtcNow().AddHours(-hours_back); + var windowEndNaive = NaiveUtc(windowEnd); + var windowStart = windowEndNaive.AddHours(-hours_back); var snapshots = await DarlingConfigHistoryReader.GetServerConfigSnapshotsAsync(postgres, resolved.ServerId); - /* The tool reads the full unbounded history and only lower-bounds, so pass DateTime.MaxValue as the - (no-op) upper edge — the shared both-edges diff then reproduces the prior behavior exactly. */ - var changes = ConfigChangeDiff.DiffServerConfigChanges(snapshots, windowStart, DateTime.MaxValue); + /* Unanchored, the tool still reads the full history and only lower-bounds — see UpperEdge, which + keeps DateTime.MaxValue as the (no-op) upper edge so the shared both-edges diff reproduces the + prior behaviour exactly. An as_of anchor is what closes the upper edge. */ + var changes = ConfigChangeDiff.DiffServerConfigChanges(snapshots, windowStart, UpperEdge(as_of, windowEndNaive)); if (changes.Count == 0) - return NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "server_config") + ?? NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); var result = changes.Select(c => new { @@ -92,21 +96,24 @@ public static async Task GetServerConfigChanges( public static async Task GetDatabaseConfigChanges( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168) + [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var windowStart = NaiveUtcNow().AddHours(-hours_back); + var windowEndNaive = NaiveUtc(windowEnd); + var windowStart = windowEndNaive.AddHours(-hours_back); var snapshots = await DarlingConfigHistoryReader.GetDatabaseConfigSnapshotsAsync(postgres, resolved.ServerId); - var changes = ConfigChangeDiff.DiffDatabaseConfigChanges(snapshots, windowStart, DateTime.MaxValue); + var changes = ConfigChangeDiff.DiffDatabaseConfigChanges(snapshots, windowStart, UpperEdge(as_of, windowEndNaive)); if (changes.Count == 0) - return NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "database_config") + ?? NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); var result = changes.Select(c => new { @@ -135,21 +142,24 @@ public static async Task GetDatabaseConfigChanges( public static async Task GetTraceFlagChanges( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168) + [Description("Hours of history to retrieve. Default 168 (7 days).")] int hours_back = 168, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var windowStart = NaiveUtcNow().AddHours(-hours_back); + var windowEndNaive = NaiveUtc(windowEnd); + var windowStart = windowEndNaive.AddHours(-hours_back); var snapshots = await DarlingConfigHistoryReader.GetTraceFlagSnapshotsAsync(postgres, resolved.ServerId); - var changes = ConfigChangeDiff.DiffTraceFlagChanges(snapshots, windowStart, DateTime.MaxValue); + var changes = ConfigChangeDiff.DiffTraceFlagChanges(snapshots, windowStart, UpperEdge(as_of, windowEndNaive)); if (changes.Count == 0) - return NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "trace_flags") + ?? NoChanges(resolved.ServerName, hours_back, DistinctCaptures(snapshots.Select(s => s.CaptureTime))); var result = changes.Select(c => new { @@ -190,9 +200,10 @@ public static async Task GetDatabaseScopedConfig( { var rows = await DarlingConfigHistoryReader.GetLatestDatabaseScopedConfigAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - "No database-scoped configuration data available. The config collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "database_scoped_config") + ?? McpHelpers.Status( + "unavailable", + "No database-scoped configuration data available. The config collector may not have run yet."); IEnumerable filtered = rows; if (!string.IsNullOrEmpty(database_name)) @@ -237,9 +248,10 @@ public static async Task GetQueryStoreHealth( { var rows = await DarlingConfigHistoryReader.GetLatestQueryStoreHealthAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - "No Query Store health data available. The query_store_health collector runs hourly (SQL Server 2016+); a server with no rows either predates Query Store or has not completed a cycle yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_store_health") + ?? McpHelpers.Status( + "unavailable", + "No Query Store health data available. The query_store_health collector runs hourly (SQL Server 2016+); a server with no rows either predates Query Store or has not completed a cycle yet."); IEnumerable filtered = rows; if (!string.IsNullOrEmpty(database_name)) @@ -302,5 +314,19 @@ private static string Scope(bool? isGlobal, bool? isSession) => _ => "", }; - private static DateTime NaiveUtcNow() => DateTime.SpecifyKind(DateTime.UtcNow, DateTimeKind.Unspecified); + private static DateTime NaiveUtc(DateTime utc) => DateTime.SpecifyKind(utc, DateTimeKind.Unspecified); + + /// + /// The diff's upper edge. With no as_of the read stays open-ended () + /// — byte-for-byte the pre-#2495 behaviour, which is the whole compatibility contract of that change; an + /// anchored read bounds at the anchor, because that is the point of sending one. + /// + /// Lite's twin computes now for the same unanchored case rather than an open edge, and the + /// two are NOT reconciled here on purpose. capture_time is stamped DateTime.UtcNow by the + /// collector in the same process that later reads it, so a snapshot can never carry a timestamp after + /// the read's own now — the two edges cannot select different rows, and unifying them would be an + /// unrelated behaviour change to two shipped reads inside a change that promises not to make one. + /// + private static DateTime UpperEdge(string? asOf, DateTime anchorNaiveUtc) => + string.IsNullOrWhiteSpace(asOf) ? DateTime.MaxValue : anchorNaiveUtc; } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigTools.cs index fa5a7c5a5..5c98d5547 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpConfigTools.cs @@ -44,9 +44,10 @@ public static async Task GetServerConfig( { var rows = await DarlingCurrentConfigReader.GetLatestServerConfigAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - "No server configuration data available. The config collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "server_config") + ?? McpHelpers.Status( + "unavailable", + "No server configuration data available. The config collector may not have run yet."); return JsonSerializer.Serialize(new { @@ -82,9 +83,10 @@ public static async Task GetDatabaseConfig( { var rows = await DarlingCurrentConfigReader.GetLatestDatabaseConfigAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - "No database configuration data available. The config collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "database_config") + ?? McpHelpers.Status( + "unavailable", + "No database configuration data available. The config collector may not have run yet."); IEnumerable filtered = rows; if (!string.IsNullOrEmpty(database_name)) @@ -139,7 +141,8 @@ public static async Task GetTraceFlags( { var rows = await DarlingCurrentConfigReader.GetLatestTraceFlagsAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("empty", "No trace flags found (none enabled, or the config collector has not run yet)."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "trace_flags") + ?? McpHelpers.Status("empty", "No trace flags found (none enabled, or the config collector has not run yet)."); return JsonSerializer.Serialize(new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDataTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDataTools.cs index 202840f3d..8d1292cea 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDataTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDataTools.cs @@ -52,19 +52,21 @@ public sealed class DarlingMcpDataTools public static async Task GetCpuUtilization( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 4.")] int hours_back = 4) + [Description("Hours of history. Default 4.")] int hours_back = 4, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var rows = await DarlingDataReader.GetCpuUtilizationAsync(postgres, resolved.ServerId, DateTime.UtcNow.AddHours(-hours_back)); + var rows = await DarlingDataReader.GetCpuUtilizationAsync(postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No CPU utilization data available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "cpu_utilization") + ?? McpHelpers.Status("unavailable", "No CPU utilization data available."); /* Downsample to 1-minute buckets to avoid overwhelming LLM context (Lite's projection). */ var bucketed = rows @@ -100,22 +102,24 @@ public static async Task GetWaitStats( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, - [Description("Maximum rows to return. Default 20.")] int limit = 20) + [Description("Maximum rows to return. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingDataReader.GetWaitStatsAsync(postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No wait stats data available for the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "wait_stats") + ?? McpHelpers.Status("unavailable", "No wait stats data available for the specified time range."); var result = rows.Take(limit).Select(r => { @@ -148,20 +152,44 @@ public static async Task GetWaitStats( public static async Task GetWaitTypes( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var types = await DarlingDataReader.GetDistinctWaitTypesAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); + if (types.Count == 0) + { + /* + An empty list said nothing about which nothing this is. A server that collected and was + quiet in THIS window wants the window widened; a server nothing has been stored for + wants somebody to look at collection, and widening will never fill it. Probed only here, + against the SAME source the read walks. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "wait_stats"); + if (gated != null) + { + return gated; + } + + return await DarlingDataReader.HasAnyWaitStatAsync(postgres, resolved.ServerId) + ? McpHelpers.Status( + "empty", + $"No wait types recorded for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS collected wait stats before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.") + : McpHelpers.Status( + "unavailable", + $"No wait stats have EVER been recorded for {resolved.ServerName}. This is not an empty window — nothing has been stored for this server at all. Delta wait stats need a SECOND collection cycle before the first row exists, so on a newly added server this clears itself; otherwise check that collection is running and that the server is enabled."); + } + return JsonSerializer.Serialize(new { server = resolved.ServerName, @@ -180,21 +208,32 @@ public static async Task GetWaitTrend( NpgsqlDataSource postgres, [Description("The exact wait type name, e.g. CXPACKET, PAGEIOLATCH_SH.")] string wait_type, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var start = now.AddHours(-hours_back); var points = await DarlingDataReader.GetWaitTrendAsync(postgres, resolved.ServerId, wait_type, start, now); if (points.Count == 0) { + /* The engine question comes BEFORE the distinct-values probe, not after it. Both are on + the miss path, so either order keeps the property that matters — but a permanently gated + engine takes this branch on every call, forever, and the probe below could never tell it + anything. Asking first makes that case one query instead of two. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "wait_stats"); + if (gated != null) + { + return gated; + } + /* Distinguish "unknown wait type here" from "nothing collected at all", handing back the ones that do have data — Lite's get_wait_trend miss vocabulary. */ var collected = await DarlingDataReader.GetDistinctWaitTypesAsync(postgres, resolved.ServerId, start, now); @@ -243,7 +282,8 @@ public static async Task GetMemoryStats( { var stats = await DarlingDataReader.GetLatestMemoryStatsAsync(postgres, resolved.ServerId); if (stats == null) - return McpHelpers.Status("unavailable", "No memory stats available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_stats") + ?? McpHelpers.Status("unavailable", "No memory stats available."); var utilization = stats.TotalPhysicalMemoryMb > 0 ? (stats.TotalPhysicalMemoryMb - stats.AvailablePhysicalMemoryMb) / stats.TotalPhysicalMemoryMb * 100 @@ -281,6 +321,21 @@ public static async Task GetMemoryClerks( try { var rows = await DarlingDataReader.GetLatestMemoryClerksAsync(postgres, resolved.ServerId); + + if (rows.Count == 0) + /* + ONE branch here, deliberately, and it is the reason this read gets no existence probe. + The read is "every clerk at MAX(collection_time)", so zero rows back is logically the + same statement as zero rows in the table — any probe against that source would agree + with the read by construction and tell the caller nothing it did not already have. What + the caller does need is to be told that an empty clerk list is NEVER a quiet period, + because on a live SQL Server it cannot be: the DMV always has clerks. + */ + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_clerks") + ?? McpHelpers.Status( + "unavailable", + $"No memory-clerk snapshot is available for {resolved.ServerName}. This read returns the LATEST snapshot rather than a window, so an empty result is never a quiet period — a live SQL Server always has memory clerks. It means nothing the memory_clerks collector stored is still retained, either because it has not run for this server or because its rows have aged out. Check get_collection_health and get_collection_log for the memory_clerks collector."); + var result = rows.Select(r => new { clerk_type = r.ClerkType, @@ -311,7 +366,8 @@ public static async Task GetFileIoStats( { var rows = await DarlingDataReader.GetLatestFileIoStatsAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No file I/O stats available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "file_io_stats") + ?? McpHelpers.Status("unavailable", "No file I/O stats available."); var result = rows.Select(r => new { @@ -346,19 +402,21 @@ public static async Task GetFileIoStats( public static async Task GetTempDbTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var rows = await DarlingDataReader.GetTempDbTrendAsync(postgres, resolved.ServerId, DateTime.UtcNow.AddHours(-hours_back)); + var rows = await DarlingDataReader.GetTempDbTrendAsync(postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No TempDB data available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "tempdb_stats") + ?? McpHelpers.Status("unavailable", "No TempDB data available."); var result = rows.Select(r => new { @@ -400,7 +458,8 @@ public static async Task GetPerfmonStats( { var rows = await DarlingDataReader.GetLatestPerfmonStatsAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No perfmon stats available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "perfmon_stats") + ?? McpHelpers.Status("unavailable", "No perfmon stats available."); IEnumerable filtered = rows; if (!string.IsNullOrEmpty(counter_name)) @@ -439,7 +498,8 @@ public static async Task GetTopQueriesByCpu( [Description("Filter to a specific database.")] string? database_name = null, [Description("If true, only return queries whose cached plan has EVER run at DOP > 1. Note: max_dop comes from sys.dm_exec_query_stats and is a lifetime-max for the plan's time in cache, so a plan compiled before MAXDOP was lowered keeps reporting the old higher value until it is evicted or recompiled. Confirm current parallelism with analyze_query_plan, which reads the actual plan.")] bool parallel_only = false, [Description("Minimum DOP to filter on. Implies parallel filtering. Filters the same lifetime-max value as parallel_only, not current parallelism.")] int min_dop = 0, - [Description("Grouping. 'query_hash' (default) is one row per (database, query_hash, host_object). 'host_object' rolls every statement of a hosting procedure/function into ONE row — use it when dynamic SQL built with per-value literals fragments one logical statement across many query_hash values, which makes top-N-by-hash structurally unable to surface it (measured at 21 fragments for one statement, whose combined CPU was the largest on the instance while no single fragment ranked). Ad-hoc statements have no host object and stay grouped per hash in both modes. distinct_query_hashes reports how many hashes a row rolled up.")] string group_by = "query_hash") + [Description("Grouping. 'query_hash' (default) is one row per (database, query_hash, host_object). 'host_object' rolls every statement of a hosting procedure/function into ONE row — use it when dynamic SQL built with per-value literals fragments one logical statement across many query_hash values, which makes top-N-by-hash structurally unable to surface it (measured at 21 fragments for one statement, whose combined CPU was the largest on the instance while no single fragment ranked). Ad-hoc statements have no host object and stay grouped per hash in both modes. distinct_query_hashes reports how many hashes a row rolled up.")] string group_by = "query_hash", + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; @@ -454,18 +514,19 @@ the exact wrong conclusion this option exists to prevent. */ $"group_by must be 'query_hash' or 'host_object' (got '{group_by}')."); } - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(top, "top"); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingDataReader.GetTopQueriesByCpuAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, top, database_name, rollUpByHostObject: rollUp); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No query stats available for the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_stats") + ?? McpHelpers.Status("unavailable", "No query stats available for the specified time range."); var filtered = rows .Where(r => !(parallel_only || min_dop > 1) || (r.MaxDop > 1 && r.MaxDop >= (min_dop > 1 ? min_dop : 2))) @@ -566,24 +627,26 @@ public static async Task GetTopProceduresByCpu( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, [Description("Number of top procedures. Default 20.")] int top = 20, - [Description("Filter to a specific database.")] string? database_name = null) + [Description("Filter to a specific database.")] string? database_name = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(top, "top"); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingDataReader.GetTopProceduresByCpuAsync(postgres, resolved.ServerId, now.AddHours(-hours_back), now, top, database_name); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - "No procedure stats available. Delta-based collection requires at least two collection cycles (~30 minutes) to produce non-zero values."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "procedure_stats") + ?? McpHelpers.Status( + "unavailable", + "No procedure stats available. Delta-based collection requires at least two collection cycles (~30 minutes) to produce non-zero values."); /* #2320: same attributed-CPU disclosure as the queries tool — one shared computation, same concurrent independent reads. */ @@ -650,19 +713,20 @@ public static async Task GetQueryStoreTop( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, [Description("Number of top queries. Default 20.")] int top = 20, - [Description("Filter to a specific database.")] string? database_name = null) + [Description("Filter to a specific database.")] string? database_name = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(top, "top"); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var requestedStart = now.AddHours(-hours_back); var rows = await DarlingDataReader.GetQueryStoreTopAsync(postgres, resolved.ServerId, requestedStart, now, top, database_name); @@ -677,13 +741,23 @@ that was served rather than echo the one that was asked for. */ var truncated = floor is DateTime f && f > requestedStart.AddMinutes(90); if (rows.Count == 0) - return McpHelpers.Status( - "unavailable", - $"No Query Store rows for this server in the {hours_back}-hour window searched. Query Store " + - "may not be enabled on the target databases -- or the window reaches past what the raw tier " + - "retains (query_store_stats is dropped at 4 days when the rollups are armed), in which case " + - "nothing was read for the older part of it. Try a shorter window before concluding the " + - "queries did not run."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_store") + /* #2546: the sentence below GUESSES ("may not be enabled"), and it has to, because the + read had no way to find out. The store has known all along -- query_store_health + records actual_state per database every hour for exactly this purpose. Asking it turns + a hedge into a fact plus the ALTER DATABASE that fixes it, and it answers for the + database this read was scoped to rather than for the server's most flattering one. */ + ?? await DarlingRuntimePrecondition.QueryStoreStatusAsync(postgres, resolved.ServerId, resolved.ServerName, database_name) + /* And the collector's own last run, for the case Query Store is on and the collector is + the thing that cannot read it. */ + ?? await DarlingRuntimePrecondition.StatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_store") + ?? McpHelpers.Status( + "unavailable", + $"No Query Store rows for this server in the {hours_back}-hour window searched. Query Store " + + "may not be enabled on the target databases -- or the window reaches past what the raw tier " + + "retains (query_store_stats is dropped at 4 days when the rollups are armed), in which case " + + "nothing was read for the older part of it. Try a shorter window before concluding the " + + "queries did not run."); var result = rows.Select(r => new { @@ -808,7 +882,7 @@ internal static string RenderServerList( }, McpHelpers.JsonOptions); } - [McpServerTool(Name = "get_collection_health"), Description("Shows the health status of all data collectors for a server — whether they're running successfully, failing, or stale. Check this before investigating data to ensure collectors are working properly. Each row also carries last_note/note_count: what a NON-failing run reported, e.g. an enumeration that came back with 0 items. note_count equal to total_runs means the collector has been collecting nothing all window — not a fault (the target may be legitimately empty), but the reason a HEALTHY collector can still have no data. target_has_user_databases tells those two apart: true means the target DID have user databases in the same window, so an all-window empty enumeration is worth investigating (a login that cannot enter them, an exclusion filter that matched everything); false means either no user databases or no inventory to go on. The sweep_pressure block is the server-level roll-up: it compares the collectors' combined execution demand (average duration amortized by cadence) against the minute the fastest cadence holds. SATURATED means the collection body cannot fit inside its cadence, so relaunches are skipped and the server collects at a multiple of its configured interval while every collector still reads healthy — heaviest_collectors names where that budget goes.")] + [McpServerTool(Name = "get_collection_health"), Description("Shows the health status of all data collectors for a server — whether they're running successfully, failing, or stale. Check this before investigating data to ensure collectors are working properly. Each row also carries last_note/note_count: what a NON-failing run reported, e.g. an enumeration that came back with 0 items. note_count equal to total_runs means the collector has been collecting nothing all window — not a fault (the target may be legitimately empty), but the reason a HEALTHY collector can still have no data. target_has_user_databases tells those two apart: true means the target DID have user databases in the same window, so an all-window empty enumeration is worth investigating (a login that cannot enter them, an exclusion filter that matched everything); false means either no user databases or no inventory to go on. The sweep_pressure block is the server-level roll-up: it compares the collectors' combined execution demand (average duration amortized by cadence) against the minute the fastest cadence holds. SATURATED means the collection body cannot fit inside its cadence, so relaunches are skipped and the server collects at a multiple of its configured interval while every collector still reads healthy — heaviest_collectors names where that budget goes. That verdict is the SUSTAINED answer only. peak_cycle_risk is the separate single-sweep answer: peak_cycle_ms is what the body costs on the cycle where every scheduled cadence comes due together, and BODY_OVERRUN means that one body cannot fit the budget even when the verdict reads OK — the signature of one infrequent heavy collector, which amortization hides and heaviest_collectors therefore ranks out of sight. peak_collector names it, and peak_cycle_note explains it. Read both fields: a server can be OK/BODY_OVERRUN (a schedule-shape problem, fix by moving or splitting that collector) or SATURATED/BODY_OVERRUN (a capacity problem). Every collector row carries avg_duration_ms, p95_duration_ms and max_duration_ms, because a collector's runs are not always one population: query_store on one dogfood server averaged 13,834 ms over 1,155 runs of which 958 yielded nothing and cost about 36 ms, which puts the other 197 at roughly 80,900 ms EACH - each one, on its own, larger than the whole sweep budget. Read the three together: avg close to p95 close to max is one population, avg far below p95 is two, and p95 far below max is one pathological run. peak_cycle_ms is built from p95 (floored at the mean, so it can never read lower than a mean-based figure) for exactly that reason, and peak_collector carries peak_run_ms beside avg_duration_ms so the gap is visible. Those three still describe RUNS, and a collector that runs once per DATABASE writes one blended row, so no run-level statistic can say which database cost what. Five fan out from an enumeration on any SQL Server target (query_store, plan_correction, query_store_health, index_object_stats, database_scoped_config); separately, eight more fan out over a per-database connection loop when the target is Azure SQL DB, and pg_autovacuum_stats always does on PostgreSQL. The per-collector `fanout` block is that answer, null for a collector that does not fan out: `items` is how wide the fan-out was, `slowest`/`slowest_ms` name the dearest database and its cost on the window's worst run, `run_ms` is that whole run, and `dominance` is slowest_ms * items / run_ms — 1.0 for a perfectly even fan-out, rising with concentration. It matters because the remedies diverge there: near 1.0 the cost is the fan-out's WIDTH and bounded parallelism is the lever, while around 2.0 or above one database dominates and a per-database schedule override or a stagger is what helps. Do not try to infer this from p95 versus avg — on a per-database collector that ratio is usually saturated by empty-versus-productive runs and says nothing about databases.")] public static async Task GetCollectionHealth( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null) @@ -833,6 +907,16 @@ here is a lock-contention signal about the monitored server. */ yields = r.YieldCount, failure_rate_pct = Math.Round(r.FailureRatePercent, 1), avg_duration_ms = Math.Round(r.AvgDurationMs, 0), + /* #2460: the mean above is a blend whenever a collector's runs come in two sizes, and + on this fleet one of them plainly does — query_store averaged 13,834 ms over 1,155 + runs where 958 yielded nothing at ~36 ms, which puts the other 197 at ~80,900 ms + each. p95 is what a HEAVY run of this collector costs and is what the peak-cycle + arithmetic below is built from; max is carried beside it so a routine tail can be + told from a single pathological cycle, which is the one thing a max alone cannot + say about itself. Read the three together: avg ~= p95 ~= max is one population, + avg << p95 is two, and p95 << max is one bad run. */ + p95_duration_ms = Math.Round(r.P95DurationMs, 0), + max_duration_ms = Math.Round(r.MaxDurationMs, 0), last_success = r.LastSuccessTime?.ToString("o"), last_error = r.LastError, /* #1837: what a NON-failing run reported — an enumeration that came back with 0 items, @@ -849,7 +933,23 @@ diagnosing an empty collector gets it as a boolean instead of parsing it out of /* The same string both WPF grids render, composed on this side so the web dashboard and any other consumer cannot re-derive it differently. */ note_summary = CollectorHealthClassifier.FormatCollectionNote( - r.LastNote, r.NoteCount, r.TotalRuns, r.CollectorName, r.TargetHasUserDatabases) + r.LastNote, r.NoteCount, r.TotalRuns, r.CollectorName, r.TargetHasUserDatabases), + /* #2472: the per-database breakdown of a collector that fans out, null for one that does + not. Emitted as a nested object rather than four sibling fields so a consumer cannot + read a slowest item without the width it has to be judged against — the parts only mean + something together, and `dominance` is that meaning, computed here so every consumer + gets the same arithmetic instead of three of them inventing it. + + This is the thing avg/p95/max structurally cannot say: they aggregate over runs, and one + run is one blended row however many databases it covered. */ + fanout = r.FanoutDominance is null ? null : new + { + items = r.FanoutItems, + slowest = r.SlowestItem, + slowest_ms = r.SlowestItemMs, + run_ms = r.SlowestRunDurationMs, + dominance = Math.Round(r.FanoutDominance.Value, 2) + } }); /* #2296: the roll-up that makes half-rate collection visible. Every collector on a saturated @@ -859,7 +959,7 @@ server reads HEALTHY — from each one's own seat nothing is wrong — so the co by cadence) against the minute the fastest cadence holds; heaviest_collectors names where the budget goes, which is the actionable half of the answer. */ var pressure = SweepPressureClassifier.Compute( - rows.Select(r => (r.CollectorName, r.AvgDurationMs, r.FrequencyMinutes))); + rows.Select(r => (r.CollectorName, r.AvgDurationMs, r.P95DurationMs, r.FrequencyMinutes))); var heaviest = rows .Where(r => r.FrequencyMinutes > 0 && r.AvgDurationMs > 0) .OrderByDescending(r => r.AvgDurationMs / r.FrequencyMinutes) @@ -868,9 +968,45 @@ server reads HEALTHY — from each one's own seat nothing is wrong — so the co { collector = r.CollectorName, avg_duration_ms = Math.Round(r.AvgDurationMs, 0), - frequency_minutes = r.FrequencyMinutes + p95_duration_ms = Math.Round(r.P95DurationMs, 0), + max_duration_ms = Math.Round(r.MaxDurationMs, 0), + frequency_minutes = r.FrequencyMinutes, + /* #2446: the ranking key said out loud, beside the single-run cost it is derived from. + The list still ranks by amortized contribution, because that is what explains + busy_percent — but an operator reading it to find the collector that overran a body + was reading the wrong column with nothing on the row to say so. */ + amortized_ms_per_minute = Math.Round(r.AvgDurationMs / r.FrequencyMinutes, 0), + /* #2460: "% of the budget PER RUN" now comes from the run that actually costs + something — PeakRunMs, the p95 floored at the mean — rather than from a mean that + on a bimodal collector describes no run at all. It is the same number the peak + cycle charges this collector, so the column and the cycle reconcile by hand; + taken from the mean, this row said query_store cost 23% of a body when its heavy + run costs 135% of one. Through the shared helper rather than re-derived here, so + the floor rule cannot drift between the two SKUs' tools. */ + pct_of_sweep_budget_per_run = Math.Round( + SweepPressureClassifier.PeakRunMs(r.AvgDurationMs, r.P95DurationMs) / SweepPressureClassifier.SweepBudgetMs * 100.0, 1) }); + /* #2446: the collector that owns the most of ONE sweep, which is a different collector from + the ones above whenever it is infrequent enough for amortization to hide it. Named on every + server, not only on BODY_OVERRUN — knowing where a body's time concentrates is worth having + before it is a problem, and this is exactly the row heaviest_collectors ranks out of sight. */ + var peakCollector = pressure.PeakCollectorName == null ? null : new + { + collector = pressure.PeakCollectorName, + /* #2460: what one aligned body is charged for this collector — its p95, floored at its + mean — with the mean kept beside it, because on a bimodal collector the GAP between + the two is the finding. amortized_ms_per_minute stays derived from the mean: that is + what amortization means, and a rate built from a tail would claim work the server + never sustains. */ + peak_run_ms = Math.Round(pressure.PeakCollectorPeakRunMs, 0), + avg_duration_ms = Math.Round(pressure.PeakCollectorAvgDurationMs, 0), + frequency_minutes = pressure.PeakCollectorFrequencyMinutes, + amortized_ms_per_minute = Math.Round(pressure.PeakCollectorAvgDurationMs / pressure.PeakCollectorFrequencyMinutes, 0), + pct_of_sweep_budget_per_run = Math.Round(pressure.PeakCollectorPeakRunMs / SweepPressureClassifier.SweepBudgetMs * 100.0, 1) + }; + var peakCycleNote = SweepPressureClassifier.FormatPeakCycleNote(pressure); + return JsonSerializer.Serialize(new { server = resolved.ServerName, @@ -879,6 +1015,18 @@ server reads HEALTHY — from each one's own seat nothing is wrong — so the co busy_ms_per_minute = Math.Round(pressure.BusyMsPerMinute, 0), busy_percent = Math.Round(pressure.BusyPercent, 1), verdict = pressure.Verdict, + /* #2446: the second dimension, and deliberately NOT folded into verdict. verdict + answers "does sustained demand fit the cadence on average"; this answers "does one + scheduled body fit at all". They disagree exactly when an infrequent heavy collector + owns most of a single sweep — which an amortized number cannot see by construction, + since dividing by that collector's own long cadence is what makes it small. Its own + vocabulary (FITS / BODY_OVERRUN) so it can never be read as a fourth verdict band, + and its own field so a fleet scan can filter on it. */ + peak_cycle_ms = Math.Round(pressure.PeakCycleMs, 0), + peak_cycle_percent = Math.Round(pressure.PeakCyclePercent, 1), + peak_cycle_risk = pressure.PeakCycleRisk, + peak_collector = peakCollector, + peak_cycle_note = string.IsNullOrEmpty(peakCycleNote) ? null : peakCycleNote, heaviest_collectors = heaviest, note = pressure.Verdict switch { @@ -910,7 +1058,8 @@ public static async Task GetServerProperties( { var row = await DarlingDataReader.GetLatestServerPropertiesAsync(postgres, resolved.ServerId); if (row == null) - return McpHelpers.Status("unavailable", "No server properties available. The properties collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "server_properties") + ?? McpHelpers.Status("unavailable", "No server properties available. The properties collector may not have run yet."); return JsonSerializer.Serialize(new { @@ -940,26 +1089,22 @@ public static async Task GetServerProperties( /* ─────────────────────────── list_servers helpers ─────────────────────────── */ - /// Older than twice the ~1-minute collector cadence = the collection has visibly lagged. - private static readonly TimeSpan StaleThreshold = TimeSpan.FromMinutes(2); - - /// Older than this (or no collection at all) = the server is treated as Offline. - private static readonly TimeSpan OfflineThreshold = TimeSpan.FromMinutes(15); - /// - /// The freshness-derived status the headless viewer's cards use (ServerSummaryItem.ClassifyFreshness): - /// Fresh → Online, Stale → Warning, long-dead → Offline, never-collected → AwaitingFirstCollection - /// (the service hasn't reached the server yet — a bootstrap state, not an outage; additive status - /// value, existing values unchanged). Both instants are UTC. + /// The freshness-derived status this tool reports: Fresh → Online, Stale → Warning, long-dead → Offline, + /// never-collected → AwaitingFirstCollection (the service hasn't reached the server yet — a bootstrap + /// state, not an outage). Both instants are UTC. + /// + /// It used to classify freshness itself, against its OWN copies of the 2-minute and 15-minute + /// thresholds — so ServerHealthThresholds could move and list_servers would silently keep + /// answering with the old numbers. It now shares the ladder with every other status surface (#2473). What + /// it does NOT share is the vocabulary: spells the + /// never-collected state as one word because that value was published to MCP clients, and a status value + /// a client keys on is a consumer API. /// - private static string FreshnessStatus(DateTime? lastCollectionUtc, DateTime nowUtc) - { - if (!lastCollectionUtc.HasValue) return "AwaitingFirstCollection"; - var age = nowUtc - lastCollectionUtc.Value; - if (age > OfflineThreshold) return "Offline"; - if (age > StaleThreshold) return "Warning"; - return "Online"; - } + private static string FreshnessStatus(DateTime? lastCollectionUtc, DateTime nowUtc) => + ServerCollectionStatusRules + .FromFreshness(ServerHealthClassifier.ClassifyFreshness(lastCollectionUtc, nowUtc)) + .McpToken(); /// Product-name label for a sql_major_version (the viewer's SqlVersionLabel); 2016+ is /// what the product supports, older/unknown majors fall back to a bare version tag, null to empty. @@ -975,4 +1120,271 @@ private static string FreshnessStatus(DateTime? lastCollectionUtc, DateTime nowU 17 => "SQL Server 2025", _ => $"SQL Server v{sqlMajorVersion}", }; + + [McpServerTool(Name = "get_collection_log"), Description("Gets the RAW per-run collection log for a server, newest first: one row per collector run with its total duration, the part spent querying the monitored server, the part spent writing to the store, rows collected, status and any error. get_collection_health rolls seven days of these into a per-collector verdict; this is the underlying runs, which is what you need when the rollup says healthy and collection still looks wrong, or when you want to see what a collector was doing during a specific incident window.")] + public static async Task GetCollectionLog( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return, newest first. Default 200.")] int limit = 200, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + /* The shared row-cap contract every sibling read uses: rejects out of range rather than + silently clamping, so a caller asking for 5000 is told no instead of quietly given 1000. */ + var invalidLimit = McpHelpers.ValidateTop(limit); + if (invalidLimit != null) return invalidLimit; + + /* ResolveAsOf here, deliberately NOT ValidateWindow. These three reads have never capped + hours_back -- they Math.Abs() it and window on the result -- so routing them through the + shared validator would impose the 168-hour ceiling every other read carries, and take reach + away from exactly the read whose premise is looking FURTHER back than the default. The anchor + is validated because it is new; the span keeps the behaviour callers already have. */ + var anchorError = McpHelpers.ResolveAsOf(as_of, out var windowEnd); + if (anchorError != null) return anchorError; + + try + { + var end = windowEnd; + var start = end.AddHours(-Math.Abs(hours_back)); + + /* Over-fetch by one so truncation is OBSERVED rather than inferred. Comparing count to the + cap cannot tell a window holding exactly `limit` runs from one holding more, and this + read's whole premise is that the cap announces itself instead of being guessed at. */ + var rows = await DarlingDataReader.GetCollectionLogAsync( + postgres, resolved.ServerId, start, end, limit + 1); + var truncated = rows.Count > limit; + if (truncated) rows = rows.Take(limit).ToList(); + + if (rows.Count == 0) + { + /* + Zero rows is two completely different facts and they need different answers. + A server that has collected before and simply did nothing in THIS window is a + true negative -- the caller narrowed to a quiet period, and widening the window + is the move. A server with no log rows at all has never collected, which is a + fault, and telling that caller "nothing in the last 24 hours" would send them + off widening a window that will never fill. So we ask which one it is rather + than emitting one sentence that is true of both. + */ + var everCollected = await DarlingDataReader.HasAnyCollectionLogAsync(postgres, resolved.ServerId); + return everCollected + ? McpHelpers.Status( + "empty", + $"No collector runs recorded for {resolved.ServerName} in the last {Math.Abs(hours_back)} hour(s). This server HAS collected before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent runs.") + : McpHelpers.Status( + "unavailable", + $"No collector runs have EVER been recorded for {resolved.ServerName}. This is not an empty window — collection has not run at all for this server. Check that the service is running and that the server is enabled for collection; get_collection_health will be equally empty until it does."); + } + + var result = rows.Select(r => new + { + collector = r.CollectorName, + collection_time = r.CollectionTime.ToString("o"), + duration_ms = r.DurationMs is null ? (double?)null : Math.Round(r.DurationMs.Value, 0), + /* + The split matters more than the total. A collector slow because the monitored + server is slow needs work on that server; one slow because the store is slow + needs work here. The total alone cannot tell those apart, and it is the + question people actually ask of this log. + */ + sql_duration_ms = r.SqlDurationMs is null ? (double?)null : Math.Round(r.SqlDurationMs.Value, 0), + store_duration_ms = r.StoreDurationMs is null ? (double?)null : Math.Round(r.StoreDurationMs.Value, 0), + rows_collected = r.RowsCollected, + status = r.Status, + error_message = r.ErrorMessage, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back = Math.Abs(hours_back), + run_count = rows.Count, + /* Observed by the over-fetch above, not inferred from the row count. */ + truncated, + runs = result, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_collection_log", ex); + } + } + + [McpServerTool(Name = "get_current_waits_trend"), Description("Gets the two Current Waits series over time for a server: waiting-task total wait duration per wait type per collection, and blocked-session counts per database per collection. get_waiting_tasks answers 'what is waiting right now' and can never say whether it is worse than an hour ago — this is that question. Use it to tell a server that is always mildly blocked from one that just started, and to see which database owns the blocking over the window rather than in one snapshot.")] + public static async Task GetCurrentWaitsTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 4.")] int hours_back = 4, + [Description("Limit the blocked-session series to one database. Omit for all databases.")] string? database_name = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + /* ResolveAsOf here, deliberately NOT ValidateWindow. These three reads have never capped + hours_back -- they Math.Abs() it and window on the result -- so routing them through the + shared validator would impose the 168-hour ceiling every other read carries, and take reach + away from exactly the read whose premise is looking FURTHER back than the default. The anchor + is validated because it is new; the span keeps the behaviour callers already have. */ + var anchorError = McpHelpers.ResolveAsOf(as_of, out var windowEnd); + if (anchorError != null) return anchorError; + + try + { + var end = windowEnd; + var start = end.AddHours(-Math.Abs(hours_back)); + + var waits = await DarlingDataReader.GetWaitingTaskTrendAsync(postgres, resolved.ServerId, start, end); + var blocked = await DarlingDataReader.GetBlockedSessionTrendAsync( + postgres, resolved.ServerId, start, end, database_name); + + if (waits.Count == 0 && blocked.Count == 0) + { + /* + Both series empty is two facts again, and here the wrong one is actively reassuring: + "nothing was waiting" reads as an all-clear, while the truth may be that the + waiting_tasks collector never ran. A caller told all-clear stops looking. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "waiting_tasks"); + if (gated != null) + { + return gated; + } + + var everCollected = await DarlingDataReader.HasAnyWaitingTaskSampleAsync(postgres, resolved.ServerId); + return everCollected + ? McpHelpers.Status( + "empty", + $"Nothing was waiting on {resolved.ServerName} in the last {Math.Abs(hours_back)} hour(s). The collector HAS sampled this server, so this is a genuine all-clear for the window rather than missing data.") + : McpHelpers.Status( + "unavailable", + $"No waiting-task samples have EVER been recorded for {resolved.ServerName}, so this is NOT an all-clear — there is nothing to read. Check that collection is running for this server before concluding it was quiet."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back = Math.Abs(hours_back), + database_name, + /* + Two series in one payload because they are read together: a wait-type spike with no + blocked sessions is a resource wait, and the same spike WITH them is contention. Split + across two tools a caller can fetch one and draw the wrong conclusion. + */ + waiting_tasks = waits.Select(w => new + { + collection_time = w.CollectionTime.ToString("o"), + wait_type = w.WaitType, + total_wait_ms = w.TotalWaitMs, + }), + blocked_sessions = blocked.Select(b => new + { + collection_time = b.CollectionTime.ToString("o"), + database_name = b.DatabaseName, + blocked_count = b.BlockedCount, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_current_waits_trend", ex); + } + } + + [McpServerTool(Name = "get_blocking_stats"), Description("Gets blocking SEVERITY over time for a server: per-minute blocking duration (event count, total, max and average wait) and per-minute deadlock severity (victim count plus total, max and average wait across every process in the graphs). get_blocking_trend and get_deadlock_trend count incidents; this is how BAD they were. Ten one-second blocks and one ten-minute block are the same count and are not the same problem, which is the distinction this read exists to make.")] + public static async Task GetBlockingStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + /* ResolveAsOf here, deliberately NOT ValidateWindow. These three reads have never capped + hours_back -- they Math.Abs() it and window on the result -- so routing them through the + shared validator would impose the 168-hour ceiling every other read carries, and take reach + away from exactly the read whose premise is looking FURTHER back than the default. The anchor + is validated because it is new; the span keeps the behaviour callers already have. */ + var anchorError = McpHelpers.ResolveAsOf(as_of, out var windowEnd); + if (anchorError != null) return anchorError; + + try + { + var end = windowEnd; + var start = end.AddHours(-Math.Abs(hours_back)); + + var blocking = await DarlingDataReader.GetBlockingDurationStatsAsync(postgres, resolved.ServerId, start, end); + + /* Parsed and bucketed by the shared aggregator rather than re-derived here: a second copy of + "what counts as a victim" is how two surfaces end up disagreeing about one deadlock. */ + var graphs = await DarlingDataReader.GetDeadlockGraphsAsync(postgres, resolved.ServerId, start, end); + var deadlocks = DeadlockSeverityAggregator.Aggregate(graphs); + + if (blocking.Count == 0 && deadlocks.Count == 0) + { + /* + The denominator is whether we LOOKED, not whether we ever FOUND anything. Blocking and + deadlocks are edge tables: a server collected perfectly for months that simply never + blocked has no rows at all, so asking "was an event ever captured" answers no and + reports a healthy server as uncollected -- the reassuring-answer failure inverted, and + a false alarm sends someone to fix collection that is working. + + So this asks collection_log for a SUCCESSFUL run of either capture path. Both are + checked because either can be off alone, and the deadlock collector is separate from + both -- the verdict covers its series too. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "blocked_process_report"); + if (gated != null) + { + return gated; + } + + var everRan = + await DarlingBlockingTrendReader.HasAnyBlockingCollectorRunAsync(postgres, resolved.ServerId) + || await DarlingBlockingTrendReader.HasAnyDeadlockCollectorRunAsync(postgres, resolved.ServerId); + return everRan + ? McpHelpers.Status( + "empty", + $"No blocking or deadlocks recorded for {resolved.ServerName} in the last {Math.Abs(hours_back)} hour(s). The blocking collectors HAVE run successfully for this server, so the window is genuinely clear rather than blind.") + : McpHelpers.Status( + "unavailable", + $"The blocking collectors have NEVER run successfully for {resolved.ServerName}, so this is NOT a clean bill of health — nothing looked. Blocked-process reports need the XE session running, or the DMV blocking snapshot collector enabled; check those before concluding this server does not block."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back = Math.Abs(hours_back), + /* + Severity, not counts. get_blocking_trend already answers how OFTEN; ten one-second + blocks and one ten-minute block share a count and are different problems. + */ + blocking_duration = blocking.Select(b => new + { + time = b.Time.ToString("o"), + event_count = b.EventCount, + total_duration_ms = b.TotalDurationMs, + max_duration_ms = b.MaxDurationMs, + avg_duration_ms = Math.Round(b.AvgDurationMs, 0), + }), + deadlock_severity = deadlocks.Select(d => new + { + time = d.Time.ToString("o"), + victim_count = d.VictimCount, + /* Every process's wait, not just the victims' -- the Dashboard analyzer's semantics. */ + total_wait_ms = d.TotalWaitMs, + max_wait_ms = d.MaxWaitMs, + avg_wait_ms = Math.Round(d.AvgWaitMs, 0), + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_blocking_stats", ex); + } + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDefaultTraceTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDefaultTraceTools.cs index fda50e00c..0644bf952 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDefaultTraceTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpDefaultTraceTools.cs @@ -42,17 +42,18 @@ public static async Task GetDefaultTraceEvents( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of events to return. Default 100.")] int limit = 100) + [Description("Maximum number of events to return. Default 100.")] int limit = 100, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back) ?? McpHelpers.ValidateTop(limit); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd) ?? McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var all = await DarlingDefaultTraceReader.ReadEventsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); @@ -63,7 +64,8 @@ public static async Task GetDefaultTraceEvents( .ToList(); if (significant.Count == 0) - return McpHelpers.Status("empty", "No significant default trace events found in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "default_trace_events") + ?? McpHelpers.Status("empty", "No significant default trace events found in the requested time range."); var events = significant.Take(limit).Select(r => { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthParserTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthParserTools.cs index 1abf5e596..7c288e3cb 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthParserTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthParserTools.cs @@ -22,10 +22,11 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// /// The system_health parse-on-read MCP tools — get_health_parser_cpu_tasks / _io_issues / _memory_broker / -/// _memory_conditions / _memory_node_oom / _scheduler_issues / _severe_errors / _system_health — served over -/// Darling's Postgres store, the SAME eight tools the Dashboard's McpHealthParserTools exposes. Where -/// the Dashboard reads its server-side-parsed collect.HealthParser_* tables, these PARSE ON READ: the -/// raw system_health_events the collector captured are shredded by the shared +/// _memory_conditions / _memory_node_oom / _scheduler_issues / _severe_errors / _significant_waits / +/// _system_health — served over Darling's Postgres store, the SAME nine tools the Dashboard's +/// McpHealthParserTools exposes. Where the Dashboard reads its server-side-parsed +/// collect.HealthParser_* tables, these PARSE ON READ: the raw system_health_events the +/// collector captured are shredded by the shared /// (PerformanceMonitor.Common — reused, NOT re-implemented) and gated by /// , exactly as the viewer's System Events tab does. The /// result is the same SIGNIFICANT warning set the Dashboard surfaces (its collector runs sp_HealthParser at @@ -45,23 +46,34 @@ namespace PerformanceMonitor.Darling.Service.Mcp; [McpServerToolType] public sealed class DarlingMcpHealthParserTools { + /// + /// The collector every one of these nine reads is served by. Named once so the #2511 capability probe + /// asks about the same collector on every read and on both SKUs; a test scans both MCP trees for the + /// names passed to the probe and holds them to CollectorCatalog, because an unknown name would + /// answer "supported" and silently restore the old wrong message. + /// + private const string SystemHealthCollectorName = "system_health_events"; + [McpServerTool(Name = "get_health_parser_system_health"), Description("Gets parsed system_health extended event data: overall health indicators captured by sp_HealthParser.")] public static async Task GetSystemHealth( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { /* The corruption + contention counter series is UNGATED (no warnings-only filter) — the viewer's GetSystemHealthAsync keeps every SYSTEM snapshot that has a timestamp. */ - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.SpServerDiagnosticsEvent, xml => One(SystemHealthParser.ParseSystemHealth(xml)), r => r.EventTime.HasValue); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No system health data found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No system health data found in the requested time range."); return JsonSerializer.Serialize(new { @@ -98,17 +110,18 @@ public static async Task GetSevereErrors( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back) ?? McpHelpers.ValidateTop(limit); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd) ?? McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; /* database_id → name resolution needs the collected size-stats mapping (the shred left it null). */ var mapTask = DarlingSystemHealthReader.GetDatabaseNameMapAsync(postgres, resolved.ServerId); var xmls = await DarlingSystemHealthReader.ReadEventXmlAsync( @@ -120,7 +133,9 @@ public static async Task GetSevereErrors( .Where(r => r != null && SystemHealthSignificance.IsSignificant(r)) .Select(r => r!) .ToList(); - if (rows.Count == 0) return McpHelpers.Status("empty", "No severe errors found in the requested time range."); + if (rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No severe errors found in the requested time range."); return JsonSerializer.Serialize(new { @@ -148,17 +163,20 @@ public static async Task GetIOIssues( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { /* IO_SUBSYSTEM fans one event out to one row per pending-request file — a many-per-event shred. */ - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.SpServerDiagnosticsEvent, SystemHealthParser.ParseIoIssues, SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No I/O issues found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No I/O issues found in the requested time range."); return JsonSerializer.Serialize(new { @@ -186,16 +204,19 @@ public static async Task GetSchedulerIssues( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.SchedulerMonitorEvent, xml => One(SystemHealthParser.ParseSchedulerIssue(xml)), SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No scheduler issues found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No scheduler issues found in the requested time range."); return JsonSerializer.Serialize(new { @@ -225,16 +246,19 @@ public static async Task GetMemoryConditions( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.SpServerDiagnosticsEvent, xml => One(SystemHealthParser.ParseMemoryConditions(xml)), SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No memory condition events found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No memory condition events found in the requested time range."); return JsonSerializer.Serialize(new { @@ -287,16 +311,19 @@ public static async Task GetCPUTasks( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.SpServerDiagnosticsEvent, xml => One(SystemHealthParser.ParseCpuTasks(xml)), SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No CPU task events found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No CPU task events found in the requested time range."); return JsonSerializer.Serialize(new { @@ -328,16 +355,19 @@ public static async Task GetMemoryBroker( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.MemoryBrokerEvent, xml => One(SystemHealthParser.ParseMemoryBroker(xml)), SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No memory broker events found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No memory broker events found in the requested time range."); return JsonSerializer.Serialize(new { @@ -371,18 +401,21 @@ public static async Task GetMemoryNodeOOM( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, - [Description("Maximum number of entries. Default 50.")] int limit = 50) + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { try { /* Memory-node OOM is never gated — every recorded OOM is significant (sp_HealthParser applies no WHERE filter to this category). */ - var c = await CollectAsync(postgres, server_name, hours_back, limit, + var c = await CollectAsync(postgres, server_name, hours_back, limit, as_of, SystemHealthParser.MemoryNodeOomEvent, xml => One(SystemHealthParser.ParseMemoryNodeOom(xml)), SystemHealthSignificance.IsSignificant); if (c.EarlyReturn != null) return c.EarlyReturn; - if (c.Rows.Count == 0) return McpHelpers.Status("empty", "No memory node OOM events found in the requested time range."); + if (c.Rows.Count == 0) + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, c.ServerId, c.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status("empty", "No memory node OOM events found in the requested time range."); return JsonSerializer.Serialize(new { @@ -426,15 +459,112 @@ public static async Task GetMemoryNodeOOM( catch (Exception ex) { return McpHelpers.FormatError("get_health_parser_memory_node_oom", ex); } } + [McpServerTool(Name = "get_health_parser_significant_waits"), Description("Gets significant individual waits from system_health: one row per wait_info event where a real session's non-BACKUP statement waited at least 500 ms on a wait type that is not idle/background — the wait type, total and signal duration, the wait resource, the session id and the waiting statement. get_wait_stats gives the instance-wide totals and can never name the statement that paid them; this is the individual waits, with their SQL text.")] + public static async Task GetSignificantWaits( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to retrieve. Default 24.")] int hours_back = 24, + [Description("Maximum number of entries. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd) ?? McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + /* + Written out rather than routed through CollectAsync because the raw event count is + load-bearing here: zero significant waits out of a hundred captured events is a healthy + server, and zero out of zero is a blind one. CollectAsync returns only the surviving rows, + so the two would arrive indistinguishable. + */ + var now = windowEnd; + var xmls = await DarlingSystemHealthReader.ReadEventXmlAsync( + postgres, resolved.ServerId, now.AddHours(-hours_back), now, SystemHealthParser.WaitInfoEvent); + + var rows = xmls + .Select(SystemHealthParser.ParseSignificantWait) + .Where(r => r != null && SystemHealthSignificance.IsSignificant(r)) + .Select(r => r!) + .ToList(); + + if (rows.Count == 0) + { + /* + Three different nothings, and only one of them is good news. Events captured but none + significant is the healthy state and costs no extra query -- we already counted them. + Nothing captured in the window needs the probe to tell a quiet window from a server + whose wait_info has never been collected, because "no significant waits" is exactly + what an operator wants to hear and a caller who believes it stops looking. + */ + if (xmls.Count > 0) + { + return McpHelpers.Status( + "empty", + $"{xmls.Count} wait_info event(s) were captured for {resolved.ServerName} in the last {hours_back} hour(s) and none was significant (needs a real session, a non-BACKUP statement, at least {SystemHealthSignificance.SignificantWaitMinDurationMs} ms, and a wait type off the idle list). Events ARE being captured, so this is the healthy answer for this read rather than missing data."); + } + + var everCaptured = await DarlingSystemHealthReader.HasAnyEventOfTypeAsync( + postgres, resolved.ServerId, SystemHealthParser.WaitInfoEvent); + if (everCaptured) + { + return McpHelpers.Status( + "empty", + $"No wait_info events were captured for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS captured them before, so the window is genuinely quiet rather than blind — widen hours_back to reach the most recent events."); + } + + /* + #2511 adds a FOURTH nothing, and it is the one that was being mis-explained. On an engine + whose system_health collector is gated off there is no session to start and no collection + to check, so the advice below is advice about something that cannot exist. The engine + answer goes first because it is the stronger claim; the text after it stays exactly right + for every engine that DOES collect this. + */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, SystemHealthCollectorName) + ?? McpHelpers.Status( + "unavailable", + $"No wait_info events have EVER been captured for {resolved.ServerName}, so this is NOT an all-clear — there is nothing here to be clear about. This read is served from the collected system_health ring buffer: check that collection is running for this server and that its system_health session is started before concluding nothing was waiting."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + wait_count = rows.Count, + shown = Math.Min(rows.Count, limit), + waits = rows.Take(limit).Select(r => new + { + event_time = r.EventTime?.ToString("o"), + wait_type = r.WaitType, + duration_ms = r.DurationMs, + /* Signal duration is the part spent runnable AFTER the resource was granted, so a + signal close to the total is CPU pressure wearing a wait type's name. */ + signal_duration_ms = r.SignalDurationMs, + wait_resource = r.WaitResource, + session_id = r.SessionId, + query_text = r.QueryText + }) + }, McpHelpers.JsonOptions); + } + catch (Exception ex) { return McpHelpers.FormatError("get_health_parser_significant_waits", ex); } + } + /* ─────────────────────────── shared read + parse + filter ─────────────────────────── */ /// The outcome of resolve + validate + read + shred + significance-filter: either an /// string (a resolution error or #1224 validation) the tool returns verbatim, - /// or the resolved + the SIGNIFICANT parsed . - private readonly record struct Collected(string? EarlyReturn, string ServerName, List Rows); + /// or the resolved / + the SIGNIFICANT parsed + /// . The id rides along for the #2511 engine-capability probe on the zero-row path — + /// re-resolving the name there would be a second chance to match a DIFFERENT server, since resolution is + /// first-wins over a partial. + private readonly record struct Collected(string? EarlyReturn, int ServerId, string ServerName, List Rows); /// - /// Resolves the server, validates hours_back + limit, reads the raw event_xml for + /// Resolves the server, validates hours_back + as_of + limit, reads the raw event_xml for /// over the window, shreds each blob with (the /// reused — 0..n records per event), and keeps only the rows /// accepts. The seven gated categories pass their @@ -442,16 +572,16 @@ public static async Task GetMemoryNodeOOM( /// predicate (ungated, matching the viewer's chart read). /// private static async Task> CollectAsync( - NpgsqlDataSource postgres, string? serverName, int hoursBack, int limit, string eventType, + NpgsqlDataSource postgres, string? serverName, int hoursBack, int limit, string? asOf, string eventType, Func> shred, Func significant) where T : class { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, serverName); - if (error != null) return new Collected(error, "", new List()); + if (error != null) return new Collected(error, 0, "", new List()); - var validation = McpHelpers.ValidateHoursBack(hoursBack) ?? McpHelpers.ValidateTop(limit); - if (validation != null) return new Collected(validation, "", new List()); + var validation = McpHelpers.ValidateWindow(hoursBack, asOf, out var windowEnd) ?? McpHelpers.ValidateTop(limit); + if (validation != null) return new Collected(validation, 0, "", new List()); - var now = DateTime.UtcNow; + var now = windowEnd; var xmls = await DarlingSystemHealthReader.ReadEventXmlAsync( postgres, resolved.ServerId, now.AddHours(-hoursBack), now, eventType); @@ -465,7 +595,7 @@ private static async Task> CollectAsync( } } - return new Collected(null, resolved.ServerName, rows); + return new Collected(null, resolved.ServerId, resolved.ServerName, rows); } /// Wraps a single-record shred (0-or-1) as the 0..n sequence expects. diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthTools.cs index 3c4156d86..39564f525 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHealthTools.cs @@ -8,6 +8,7 @@ using System; using System.ComponentModel; +using System.Linq; using System.Text.Json; using System.Threading.Tasks; using ModelContextProtocol.Server; @@ -19,13 +20,27 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// -/// The health-overview MCP tools — get_server_summary (Lite's one-shot per-server health) and -/// get_daily_summary (the Dashboard's daily rollup) — served over Darling's Postgres store, the same names -/// those SKUs expose. Both read through (STORED reads, no live -/// monitored-server hit). get_server_summary is the fast "is this server OK" check — current SQL CPU, memory, -/// recent blocking, recent deadlocks — before drilling in; get_daily_summary folds a day's signals into the -/// SHARED composite health band (DailyHealthBandCalculator), so an agent gets the same Healthy / -/// Warning / Critical verdict the Performance Calendar shows. +/// The health-overview MCP tools — get_server_summary (Lite's one-shot per-server health), get_daily_summary +/// (the Dashboard's daily rollup for ONE day) and get_daily_summary_range (#2484: the same rollup across a +/// span of days, which is what the Performance Calendar's month grid draws) — served over Darling's Postgres +/// store, the same names those SKUs expose. All read through (STORED reads, +/// no live monitored-server hit). get_server_summary is the fast "is this server OK" check — current SQL CPU, +/// memory, recent blocking, recent deadlocks — before drilling in; the two daily reads fold a day's signals +/// into the SHARED composite health band (DailyHealthBandCalculator), so an agent gets the same +/// Healthy / Warning / Critical verdict the Performance Calendar shows. +/// +/// Why the range is a SIBLING rather than a wider get_daily_summary. Four reasons, and the first +/// two are mechanical. (1) The response SHAPE is the contract: get_daily_summary returns a flat object of +/// scalars, which is what a stat tile consumes and what its own description promises; a range returns rows. A +/// single tool that returned either depending on whether a span argument arrived would make every consumer +/// branch on a parameter it may not have sent, and the web stat panel would simply stop rendering. (2) The web +/// server page must not fetch one read twice — there is a pin for it — and the Overview tab already reads +/// get_daily_summary for today's band, so the month grid beside it CANNOT be the same read. (3) They are +/// different questions with different defaults: "how was Tuesday" versus "which of the last thirty days were +/// bad", the second of which is a screening read whose answer is the shape of the month rather than one day's +/// numbers. (4) get_daily_summary is a shipped name on both SKUs; changing its payload shape would break +/// callers for no gain. The two share the ONE aggregate underneath (DailySummarySql.RangeSql), which is +/// what stops them ever disagreeing about a day. /// [McpServerToolType] public sealed class DarlingMcpHealthTools @@ -112,4 +127,103 @@ public static async Task GetDailySummary( return McpHelpers.FormatError("get_daily_summary", ex); } } + + [McpServerTool(Name = "get_daily_summary_range"), Description("Gets the daily health summary for a SPAN of days rather than one: one row per collected day, each with its composite health band (Healthy/Warning/Critical), total wait time, top wait type, unique query count, deadlocks, blocking events with the peak block wait, high-CPU samples, memory pressure, collection errors and actionable alert count. This is what the desktop viewer's Performance Calendar month grid draws, and it is the read to use when the question is WHICH day rather than how one day went — scan the bands, then call get_daily_summary for the day that stands out. A day on which anything at all was collected appears here even if every signal was quiet (that day is Healthy, not missing), so a gap in the returned days is a gap in COLLECTION.")] + public static async Task GetDailySummaryRange( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Days of history, ending on the anchor day (inclusive). Default 30; max 366 (a year).")] int days_back = 30, + [Description(McpHelpers.AsOfDaysDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + /* A year, not the calendar's month, because "how did last quarter look" is a real question — but + bounded, because the aggregate underneath scans the RAW per-collection series for every signal + except the query count (which routes to the rollup). The ceiling is SHARED with Lite so the two + SKUs cannot accept different spans. */ + if (days_back <= 0 || days_back > McpHelpers.MaxDailySummaryDaysBack) + return $"Invalid days_back value '{days_back}'. Must be a positive integer (1-{McpHelpers.MaxDailySummaryDaysBack})."; + + /* The anchor is the ONLY source of "now" in this body — see AsOfWindowAnchorTests, which fails a tool + that advertises as_of and then reads the process clock anyway. (That check is a source scan and + the rule is absolute, so this comment cannot name the property either.) The resolver returns the + present when the caller sent nothing, which is exactly the pre-anchor behaviour. */ + var anchorError = McpHelpers.ResolveAsOf(as_of, out var windowEnd); + if (anchorError != null) return anchorError; + + try + { + /* Days, not hours: the anchor names a DAY here and only its UTC date is used, because the + aggregate buckets on date_trunc('day', ...) and a half-day is not a row it can return. The + range is half-open [from, to) over whole days, so `days_back` days ending ON the anchor day + means the anchor day is the last one included rather than the first one excluded. */ + var lastDay = windowEnd.Date; + var fromDate = lastDay.AddDays(-(days_back - 1)); + var toDate = lastDay.AddDays(1); + + var rows = await DarlingHealthReader.GetDailySummaryRangeAsync( + postgres, resolved.ServerId, fromDate, toDate); + + if (rows.Count == 0) + { + /* + Zero DAYS is two facts. The aggregate's day spine is a UNION over nine sources and one of + them is the collection log, where ANY run marks the day collected — that is why a quiet + but monitored day comes back Healthy rather than absent. So a range with no rows at all + cannot be "the server was quiet"; it is either a range that predates this server's + history, or a server nothing has ever been collected for. + + The denominator is therefore the DATA, probed on v_collection_log — the spine member that + guarantees a collected day appears at all. collection_log is PERIODIC: every collector run + writes a row whatever it found, so its presence is proof somebody looked, and unlike an + edge table it cannot report a healthy server as uncollected. + */ + var everCollected = await DarlingDataReader.HasAnyCollectionLogAsync(postgres, resolved.ServerId); + return everCollected + ? McpHelpers.Status( + "empty", + $"No collected days for {resolved.ServerName} between {fromDate:yyyy-MM-dd} and {lastDay:yyyy-MM-dd}. A day with ANY collection appears here even when every signal was quiet, so this range is outside what the store holds for this server rather than a stretch of quiet days — widen days_back, or move as_of.", + new { from_date = fromDate.ToString("yyyy-MM-dd"), to_date = lastDay.ToString("yyyy-MM-dd") }) + : McpHelpers.Status( + "unavailable", + $"No collector runs have EVER been recorded for {resolved.ServerName}, so the calendar is empty because nothing has been collected — not because those days were quiet. Check that the service is running and that the server is enabled for collection.", + new { from_date = fromDate.ToString("yyyy-MM-dd"), to_date = lastDay.ToString("yyyy-MM-dd") }); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + days_back, + /* The bounds the read actually used, echoed back: with an anchor in play, a caller cannot + otherwise tell which days they were given from the days they got. */ + from_date = fromDate.ToString("yyyy-MM-dd"), + to_date = lastDay.ToString("yyyy-MM-dd"), + /* Days WITH data, not days in the span. The two differ exactly where collection has a hole, + and that difference is the most useful thing on this payload. */ + day_count = rows.Count, + days = rows.Select(row => new + { + summary_date = row.SummaryDate.ToString("yyyy-MM-dd"), + overall_health = row.OverallHealth, + health_band = row.HealthBand.ToString(), + total_wait_time_sec = row.TotalWaitTimeSec, + top_wait_type = row.TopWaitType, + unique_queries = row.UniqueQueries, + deadlock_count = row.DeadlockCount, + blocking_events = row.BlockingEvents, + high_cpu_events = row.HighCpuEvents, + memory_pressure_events = row.MemoryPressureEvents, + memory_critical_events = row.MemoryCriticalEvents, + collection_errors = row.CollectionErrors, + alert_count = row.AlertCount, + max_block_duration_ms = row.MaxBlockDurationMs, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_daily_summary_range", ex); + } + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs index 00c24d73b..90cc70e8c 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpHostService.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -137,6 +137,9 @@ protected override async Task ExecuteAsync(CancellationToken stoppingToken) network exposure block is restart-only by design). */ DarlingConfig? config = null; var lastFailedStartUtc = DateTime.MinValue; + /* #2389: the last control-plane-override report emitted, so a steady disagreement is stated once per + distinct state instead of on every 5s poll tick. */ + string? lastOverrideReport = null; while (!stoppingToken.IsCancellationRequested) { if (config is null && DateTime.UtcNow - lastFailedStartUtc >= FailedStartBackoff) @@ -169,13 +172,29 @@ protected override async Task ExecuteAsync(CancellationToken stoppingToken) } var published = _state.Read(); - var enabled = published?.Enabled ?? config.Mcp.Enabled; - var desiredPort = published?.Port ?? config.Mcp.Port; - switch (DecideMcpAction(_app is not null, _runningPort, enabled, desiredPort)) + /* #2389: the store still wins whenever the worker has published (unchanged), but the resolution + now carries WHICH plane supplied each value, so neither the start line nor a disagreement has to + be inferred from two INFO lines five seconds apart. */ + var toggle = DarlingHostBinding.ResolveEndpointToggle( + published is null ? null : (published.Enabled, published.Port), config.Mcp.Enabled, config.Mcp.Port); + + /* Report the DISAGREEMENT at the point of override, not the outcome. Once per distinct state (the + last-reported string, the same shape as the firewall check's ShouldReport) so a steady mismatch + says its piece once per service start rather than every poll tick, while a LATER re-divergence — + someone toggling the store after boot — is still reported. */ + var overrideReport = DarlingHostBinding.DescribeToggleOverride(toggle, "mcp", "MCP", config.Mcp.Enabled, config.Mcp.Port); + if (overrideReport is not null && !string.Equals(overrideReport, lastOverrideReport, StringComparison.Ordinal)) + { + _logger.LogWarning("{Report}", overrideReport); + } + + lastOverrideReport = overrideReport; + + switch (DecideMcpAction(_app is not null, _runningPort, toggle.Enabled, toggle.Port)) { case McpSupervisorAction.Start when DateTime.UtcNow - lastFailedStartUtc >= FailedStartBackoff: - if (!await TryStartServerAsync(config, desiredPort, stoppingToken)) + if (!await TryStartServerAsync(config, toggle, stoppingToken)) { lastFailedStartUtc = DateTime.UtcNow; } @@ -188,9 +207,9 @@ protected override async Task ExecuteAsync(CancellationToken stoppingToken) case McpSupervisorAction.Restart: _logger.LogInformation( - "MCP port changed via the control plane ({Old} -> {New}) — rebinding", _runningPort, desiredPort); + "MCP port changed via the control plane ({Old} -> {New}) — rebinding", _runningPort, toggle.Port); await StopServerAsync(stoppingToken); - if (!await TryStartServerAsync(config, desiredPort, stoppingToken)) + if (!await TryStartServerAsync(config, toggle, stoppingToken)) { lastFailedStartUtc = DateTime.UtcNow; } @@ -258,15 +277,21 @@ private async Task DisposeFailedStartAsync() } /// - /// One start ATTEMPT of the inner MCP web app at (#1560): the whole + /// One start ATTEMPT of the inner MCP web app at 's port (#1560): the whole /// pre-supervisor startup body, with two changes — the port comes from the live control-plane value /// rather than the file, and every bail path returns false so the supervisor can retry with backoff /// instead of standing down for the process lifetime. The bind/network/token decisions still come from /// the FILE-loaded config (network exposure is deliberately restart-only); returns true when the app /// is started and listening. + /// #2389: the toggle carries the enable/port PROVENANCE, not just the port, so the start line names + /// the plane each half of the bind came from — the operator greps that line and stops reading, so it has + /// to admit when it is starting on file values the control plane may be about to contradict. /// - private async Task TryStartServerAsync(DarlingConfig config, int effectivePort, CancellationToken stoppingToken) + private async Task TryStartServerAsync( + DarlingConfig config, DarlingHostBinding.EndpointToggle toggle, CancellationToken stoppingToken) { + var effectivePort = toggle.Port; + /* Decide the effective bind PURELY, then map the reason -> severity here (Round-4 #7: the caller, not the pure fn, chooses LogCritical vs LogWarning; tests assert (Mode, Reason) without a logger). */ var bind = ResolveMcpBind(config.Mcp, config.Postgres.Managed); @@ -510,6 +535,16 @@ the subset Gemini/Antigravity accepts — collapsing nullable type unions and and the Dashboard expose, over Darling's Postgres store (STORED reads, no live hit). These are the tools the analysis findings' next_tools recommendations point at. */ .WithGeminiCompatibleTools() + /* get_query_store_regressions (#2484) — the viewer's Query Store Regressions tab. Every + other Query Store read answers what is EXPENSIVE; this answers what got WORSE, which is + not derivable from the first (the costliest query is usually the one that always was). + A STORED read over the same query_store_stats the tools above read. */ + .WithGeminiCompatibleTools() + /* get_query_heatmap (#2484) — the viewer's Query Heatmap tab. The interactive plot is + desktop-only by design; the READ behind it is not, and a bucketed table is the same + answer. It is the only query read with a TIME axis: the rankings above cannot show that + a window had a quiet half and a bad half. A STORED read over the same query_stats. */ + .WithGeminiCompatibleTools() /* The diagnostic-depth data-read tools (blocking/deadlocks, sessions, config-history, index/object) — get_blocking / get_deadlocks / get_deadlock_detail / get_blocked_process_xml, get_session_stats / get_active_queries / get_waiting_tasks, @@ -556,6 +591,10 @@ without lying about a unit or emitting mostly-null columns. */ pg_statement_stats collector. Carries Aurora's I/O source split and per-statement peak memory, neither of which the SQL Server tools have an equivalent for. */ .WithGeminiCompatibleTools() + /* get_pg_plans — the plan itself, not a pointer to one (#2567). Registered beside the + statement tools because that is the join: a plan is read alongside the statement it + belongs to, on query_id. */ + .WithGeminiCompatibleTools() /* get_pg_wraparound_risk — XID/MultiXact freeze headroom, the highest-consequence PostgreSQL signal and one with no SQL Server counterpart. Not Aurora-gated. */ .WithGeminiCompatibleTools() @@ -581,6 +620,47 @@ with the ROOT attributed. The one PostgreSQL read whose caveat has to travel WIT sampled, so "no blocking" here means "none was sampled" and the tool reports its own capture count so that distinction cannot be lost. */ .WithGeminiCompatibleTools() + /* get_pg_database_stats — four questions off one cluster-wide view: temp-file spills (the + PostgreSQL answer to "why is this query slow" that no other read here can give on a stock + target), the buffer-cache hit ratio, a server-recorded deadlock count, and the + commit/rollback split. The one read whose reset handling is part of its contract: a + statistics reset is reported as a reset rather than surfacing as a negative rate. */ + .WithGeminiCompatibleTools() + /* get_pg_index_usage — per-index scan counts with the catalog facts that decide whether an + index can actually go. The half that is not in pg_stat_user_indexes is the point: a + unique index backing a constraint enforces it without ever registering a scan, so advice + derived from the counter alone tells somebody to drop their primary key. */ + .WithGeminiCompatibleTools() + /* get_pg_table_bloat — the damage the vacuum reads above measure the cause of. The only + read here whose headline number is an ESTIMATE, and the one whose contract is that it + suppresses that number rather than captioning it when its inputs cannot be trusted. */ + .WithGeminiCompatibleTools() + /* get_pg_session_states — the session side of the xmin horizon, and the one read here whose + job includes REFUSING a causal claim. get_pg_xmin_horizon says a session is holding the + horizon; this says which one, and — measured on a live instance — says when an + idle-in-transaction session that looks identical is holding nothing at all, because a + READ COMMITTED transaction that only read has already released its snapshot. */ + .WithGeminiCompatibleTools() + /* #2659: these six shipped REGISTERED NOWHERE. They were implemented, documented, dispatched + by the web API and counted in the instructions census, and an agent could not call one of + them — the web dashboard could, which is why it went unnoticed. Registration here is + per-class and explicit, with no assembly scan, so a tools class is reachable only if + someone remembers this line and nothing failed when they did not. + McpToolTypeRegistrationTests now derives the check by reflection instead of trusting it. */ + .WithGeminiCompatibleTools() + .WithGeminiCompatibleTools() + .WithGeminiCompatibleTools() + .WithGeminiCompatibleTools() + .WithGeminiCompatibleTools() + .WithGeminiCompatibleTools() + /* get_pg_deadlocks / get_pg_deadlock_detail (#2661) - the reports themselves, out of the + server log, rather than pg_stat_database's count. */ + .WithGeminiCompatibleTools() + /* get_pg_wait_trend / get_pg_query_duration_trend / get_pg_io_trend / + get_pg_database_trend (#2663) - the PostgreSQL time series. Fourteen trend reads shipped + and none worked on this engine. All four live on one tools class, so this line covers + the later two as well - which is the only reason adding them needed no edit here. */ + .WithGeminiCompatibleTools() .WithGeminiCompatibleTools() .WithGeminiCompatibleTools() .WithGeminiCompatibleTools() @@ -595,10 +675,12 @@ capture count so that distinction cannot be lost. */ get_mute_rules via the service-side PgMuteRuleStore), the CURRENT-config snapshot trio (get_server_config / get_database_config / get_trace_flags — latest capture, the companion to the *_changes diff tools), and the health overview (get_server_summary + the daily rollup - get_daily_summary, folded through the shared DailyHealthBandCalculator). Same names Lite and - the Dashboard expose, all STORED reads over Darling's Postgres store (no live hit). The - blocking-trend / deadlock-trend, memory-pressure-event, and wait-type siblings ride along on - the existing blocking / memory-grant / core data-read classes above. */ + get_daily_summary and its #2484 range sibling get_daily_summary_range — the Performance + Calendar's month grid — both folded through the shared DailyHealthBandCalculator). Same + names Lite and the Dashboard expose, all STORED reads over Darling's Postgres store (no live + hit). The blocking-trend / deadlock-trend / lock-wait-trend, memory-pressure-event, and + wait-type siblings ride along on the existing blocking / memory-grant / core data-read + classes above. */ .WithGeminiCompatibleTools() .WithGeminiCompatibleTools() .WithGeminiCompatibleTools() @@ -613,8 +695,9 @@ get_fleet_overview this is a cross-server read the central store makes possible. .WithGeminiCompatibleTools() /* The system_health parse-on-read family — get_health_parser_cpu_tasks / _io_issues / _memory_broker / _memory_conditions / _memory_node_oom / _scheduler_issues / - _severe_errors / _system_health — the same names the Dashboard exposes. Where the Dashboard - reads its server-side-parsed collect.HealthParser_* tables, these shred the raw + _severe_errors / _significant_waits / _system_health — the same names the Dashboard + exposes. Where the Dashboard reads its server-side-parsed collect.HealthParser_* + tables, these shred the raw system_health_events on read via the shared SystemHealthParser (Common) and gate with the service-side twin of the viewer's SystemEventSignificance, exactly as the viewer's System Events tab does — the same SIGNIFICANT warning set, no live hit. */ @@ -643,10 +726,23 @@ or tear down FLEET monitoring conversationally. The service-side twin of the Vie resolver the read tools use. The mcp role carries the narrow INSERT/UPDATE/DELETE grant on ONLY config.config_monitored_servers (the encrypted_password column stays SELECT-carved) — never the config pivot or a schema-wide write. */ - .WithGeminiCompatibleTools(); + .WithGeminiCompatibleTools() + /* Optional GCF (Graph Compact Format) output: a single call-tool filter that, + when DARLING_OUTPUT_FORMAT=gcf, re-encodes each tool's JSON result as a GCF + generic wire. Registered once; covers every tool. Opt-in, lossless, and + never larger than the JSON (see GcfCallToolFilter / GcfOutput). */ + .WithRequestFilters(filters => filters.AddCallToolFilter(GcfCallToolFilter.Instance)); _app = builder.Build(); + /* #2479 item 5: every gate below used to refuse silently, so "is my token wrong or my CIDR + wrong" was answerable only from the client, which sees one opaque status code. One log per + refusal is rate-limited per (gate, source) because this port is LAN-exposed on purpose and + an exposed port meets a scanner eventually - see DarlingHttpRefusalLog for the shape and + what it deliberately never writes. Created here, per started server, so a rebind starts + with a clean budget rather than inheriting the previous listener's scan. */ + var refusals = new DarlingHttpRefusalLog(); + /* DNS-rebinding guard (#1648) — the FIRST middleware, in BOTH modes, mirroring the web host's #1576 fix. The loopback bind is tokenless by design (the network gates below install only in network mode), so a browser ON this host that loads attacker content could be rebound to @@ -661,6 +757,12 @@ preflight applies. Require the Host header to name an address we actually bind { if (!DarlingHostBinding.IsAllowedHost(context.Request.Host.Host, networkListenIp)) { + refusals.Report( + _logger, "MCP", DarlingRefusalGate.HostAllowlist, StatusCodes.Status400BadRequest, + context.Connection.RemoteIpAddress, + $"the Host header '{DarlingHttpRefusalLog.Sanitize(context.Request.Host.Host)}' is not an address this endpoint binds" + + " (a loopback name/IP, or mcp.network.listen when LAN-exposed)", + DateTime.UtcNow); context.Response.StatusCode = StatusCodes.Status400BadRequest; return; } @@ -682,8 +784,33 @@ keep working. Both run BEFORE MapMcp (D3-b: "first ... before any handler/handsh _app.Use(async (context, next) => { - if (!IsBearerTokenAuthorized(context.Request.Headers.Authorization.ToString(), token)) + /* Materialized once: StringValues.ToString() allocates, and the refusal path below needs + the same header again to tell "no credential" from "wrong credential" (review catch on + #2479). IsBearerTokenAuthorized keeps taking the raw header rather than returning what + it parsed - its signature is pinned by DarlingMcpHostTests and DarlingHostBindingTests, + and threading a result type through it to save one parse on an ALREADY-REFUSED request + is not a trade worth making. */ + var authorization = context.Request.Headers.Authorization.ToString(); + + if (!IsBearerTokenAuthorized(authorization, token)) { + /* THREE client states, never the token's value. Each one is a different next step + for the operator, which is the whole point of logging this at all: + + no header -> a client that was never configured with a token + header, not a Bearer -> a client configured wrong (Basic, a bare token, an + empty "Bearer ") - it IS sending something + a Bearer that misses -> a token that does not match this endpoint's + + Review catch on #2479: ExtractBearerToken returns null for the first TWO, so + testing only it reported "nothing was presented" about a client that presented + a malformed header - collapsing precisely the ambiguity this exists to resolve. + None of the three says anything about what the token IS. */ + refusals.Report( + _logger, "MCP", DarlingRefusalGate.Token, StatusCodes.Status401Unauthorized, + context.Connection.RemoteIpAddress, + DescribeBearerRefusal(authorization), + DateTime.UtcNow); context.Response.StatusCode = StatusCodes.Status401Unauthorized; context.Response.Headers.WWWAuthenticate = "Bearer"; return; @@ -696,6 +823,11 @@ keep working. Both run BEFORE MapMcp (D3-b: "first ... before any handler/handsh { if (!IsRemoteAddressAllowed(context.Connection.RemoteIpAddress, cidr)) { + refusals.Report( + _logger, "MCP", DarlingRefusalGate.SourceCidr, StatusCodes.Status403Forbidden, + context.Connection.RemoteIpAddress, + $"its address is outside mcp.network.allowFrom ({cidr})", + DateTime.UtcNow); context.Response.StatusCode = StatusCodes.Status403Forbidden; return; } @@ -706,17 +838,41 @@ keep working. Both run BEFORE MapMcp (D3-b: "first ... before any handler/handsh _app.MapMcp(); + /* #2389: name the authority for each half of what is being started. enabled/port come from + whichever plane the supervisor resolved; listen/allowFrom/token are always darling.json. */ + var origin = DarlingHostBinding.DescribeToggleOrigin(toggle); if (networkMode) { _logger.LogInformation( - "Starting MCP server on http://{Listen}:{Port} (LAN-exposed to {Cidr} behind a bearer token + in-app CIDR; loopback also bound)", - primaryBind, effectivePort, allowedCidr); + "Starting MCP server on http://{Listen}:{Port} (LAN-exposed to {Cidr} behind a bearer token + in-app CIDR; loopback also bound) — " + + "enabled/port from {Origin}; listen/allowFrom/token from darling.json mcp.network (file-only, restart-only)", + primaryBind, effectivePort, allowedCidr, origin); } else { - _logger.LogInformation("Starting MCP server on http://localhost:{Port} (loopback only)", effectivePort); + _logger.LogInformation( + "Starting MCP server on http://localhost:{Port} (loopback only) — enabled/port from {Origin}", + effectivePort, origin); } + /* #2479 item 6: the network block is read ONCE and held for the process lifetime by design. + Say so at every start, in BOTH modes - the loopback line above never mentioned the block at + all, and loopback-when-you-expected-LAN is exactly the state being diagnosed. + + The null-conditional is load-bearing, not defensive. config.Mcp.Network is McpNetworkConfig? + and is NULL on the default secure config - no mcp.network block at all - which is exactly the + state this line exists to describe. The dereferences at 313/317 are safe because they sit + inside if (networkMode), where the bind resolution has already proven a block exists; this + one runs unconditionally, so a bare .IsConfigured throws on every start of an un-exposed + server, gets swallowed by the catch below, and retry-fails forever because the config never + changes. Review catch on #2479. */ + _logger.LogInformation( + "{Report}", + DarlingHostBinding.DescribeNetworkBlockLifetime( + "mcp", "MCP", config.Mcp.Network?.IsConfigured ?? false, networkMode, + networkMode ? primaryBind.ToString() : null, + networkMode ? allowedCidr.ToString() : null)); + /* StartAsync, not RunAsync (#1560): the supervisor loop owns the wait — the app keeps serving until StopServerAsync (toggle-off, port change, or shutdown). */ await _app.StartAsync(stoppingToken); @@ -867,6 +1023,33 @@ internal static bool IsBearerTokenAuthorized(string? authorizationHeaderValue, s return string.IsNullOrEmpty(token) ? null : token; } + /// + /// PURE: why a bearer check refused, in the operator's terms — three states, not two (#2479). + /// + /// answers null for BOTH "no header" and "a header that is not a + /// well-formed Bearer", so a refusal line built on it alone tells an operator nothing was presented + /// while their client is sending Authorization: Basic … every second. Those are different + /// faults with different fixes — one client has no token configured, the other has it configured + /// wrong — and telling them apart is the reason this line exists. + /// + /// Says nothing about the token's value, and cannot: it reads only the header's SHAPE. + /// + internal static string DescribeBearerRefusal(string? authorizationHeaderValue) + { + if (string.IsNullOrWhiteSpace(authorizationHeaderValue)) + { + return "no 'Authorization: Bearer ' header was presented"; + } + + if (ExtractBearerToken(authorizationHeaderValue) is null) + { + return "an Authorization header WAS presented but is not a 'Bearer ' " + + "(wrong scheme, or an empty token after 'Bearer')"; + } + + return "the presented bearer token does not match mcp.network.encryptedToken"; + } + /// /// PURE severity map for a reason (Round-4 #7): the fail-closed degrades /// (// diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpInstructions.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpInstructions.cs index e11a2082b..85922b7c0 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpInstructions.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpInstructions.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -62,17 +62,35 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) ## Tool Reference - This server exposes 101 tools. 76 are the same names Performance Monitor Lite exposes, spanning diagnostic analysis, plan analysis, data reads at core and diagnostic depth, resource contention + jobs, trends, system-health parse-on-read, alerts + health overview, and the Default Trace. The remaining 25 are unique to Darling: eight are the PostgreSQL reads (Aurora/PostgreSQL targets only Darling's central store can hold), eight are the Custom Views tools (seven manage the saved views — the one view-authoring write surface — and `describe_custom_view_catalog` returns the read-only compose vocabulary those authoring tools draw from), three are alert-tuning write tools (`update_alert_settings` tunes the alert engine's thresholds; `create_mute_rule` / `delete_mute_rule` manage the mute rules) that write only the shared alert configuration in the monitoring store, two are server-onboarding write tools (`add_servers` bulk-adds monitored servers; `remove_server` removes one) that add or remove rows in the monitoring store's monitored-server registry, `get_fleet_overview` and `get_ag_health` are the two cross-server reads only a central store can answer, `get_store_metrics` reads the monitoring store's OWN hourly size/compression/growth series for capacity forecasting, and `get_blocking` is Darling's name for the blocked-process-report read that Lite exposes as `get_blocked_process_reports` — a naming difference, not a capability gap. Every data-read tool reads the data the collectors already captured into the store — a stored read, never a live query against the monitored server. + This server exposes 134 tools. 86 are the same names Performance Monitor Lite exposes, spanning diagnostic analysis, plan analysis, data reads at core and diagnostic depth, resource contention + jobs, trends, system-health parse-on-read, alerts + health overview, and the Default Trace. The remaining 48 are unique to Darling: thirty-one are the PostgreSQL reads (Aurora/PostgreSQL targets only Darling's central store can hold), eight are the Custom Views tools (seven manage the saved views — the one view-authoring write surface — and `describe_custom_view_catalog` returns the read-only compose vocabulary those authoring tools draw from), three are alert-tuning write tools (`update_alert_settings` tunes the alert engine's thresholds; `create_mute_rule` / `delete_mute_rule` manage the mute rules) that write only the shared alert configuration in the monitoring store, two are server-onboarding write tools (`add_servers` bulk-adds monitored servers; `remove_server` removes one) that add or remove rows in the monitoring store's monitored-server registry, `get_fleet_overview` and `get_ag_health` are the two cross-server reads only a central store can answer, `get_store_metrics` reads the monitoring store's OWN hourly size/compression/growth series for capacity forecasting, and `get_blocking` is Darling's name for the blocked-process-report read that Lite exposes as `get_blocked_process_reports` — a naming difference, not a capability gap. Every data-read tool reads the data the collectors already captured into the store — a stored read, never a live query against the monitored server. + + ### Reading an empty result + + When a read comes back with no data, the `status` word says WHICH kind of nothing it is, and the four are not interchangeable. `empty` is a true negative: we looked and there was nothing to find. `unavailable` means this server could have that data and does not have it right now, so collection health is worth a look. `not_collected` means this server does not collect that at all — and when the reason is the ENGINE, the gap is PERMANENT: the collector serving that read does not run on this server's engine (an Azure SQL Database has no system_health session, no default trace and no SQL Agent; a PostgreSQL target collects none of the SQL Server signals at all, and the `get_pg_*` reads are the ones that answer there), so there is no session to start, no collector to enable, and nothing to check. The message names the engine and the collector. Do not send anyone to go and fix it. `precondition` is the one that IS worth acting on: this server could have that data, the collector is running, and a setup step on the monitored server is in the way — a Query Store that is off or has gone READ_ONLY, an Extended Events capture session that is not running, an extension that was never created, a grant the monitoring login was refused. The message names the precondition, quotes what the monitored server itself said, and gives the statement or grant that satisfies it. It is re-derived on EVERY read rather than decided when the connection was made, so once somebody does the thing it asked for the next call answers with data — usually with nothing to restart on the monitoring side. A few preconditions are the exception and SAY SO IN THEIR OWN MESSAGE: the fact that gates them is read once when the service connects to that server and cached for the connection's life, so satisfying them also needs the service to reconnect before collection resumes. Read the message rather than assuming the general case — it tells you which kind you have, and telling somebody to retry a connect-scoped one without reconnecting sends them round a loop that never terminates. + + ### Asking about a PAST window + + Every tool below that takes `hours_back` also takes `as_of`: an optional ISO-8601 UTC instant that moves the END of the window off "now". `hours_back` stays the window's LENGTH. So the four hours around last Tuesday 03:00 is `as_of=2026-08-19T05:00:00Z, hours_back=4` — not `hours_back=170`. + + Reach for it whenever the question is about a time rather than about the present, and do NOT substitute a wider `hours_back`: for an aggregate read a wider window is a DIFFERENT answer, not the same answer with more rows. It changes what a top-N returns, what an average is taken over, and how much a capped read truncates. + + - Accepted forms: `2026-08-19T05:00:00Z`, `2026-08-19T05:00:00` (read as UTC), `2026-08-19T07:00:00+02:00`, `2026-08-19` (midnight UTC). + - An unparseable `as_of`, or one in the future, is REFUSED with a message rather than quietly answered as "now" — a read that silently reverts to now is indistinguishable from a correct one. + - An `as_of` older than anything the store still holds is NOT refused. It returns the read's normal `empty` / `unavailable` status, which means exactly what it says: we looked in the window you named and there was nothing in it. + - Tools that take no window at all (latest-snapshot reads like `get_memory_stats`, `get_file_io_stats`, `get_index_usage`, and the configuration reads) do not take `as_of` — they read the newest row, and there is no window to move. + - The analysis family DOES take it (#2506), and the anchor reaches the ENGINE rather than stopping at the tool: `get_analysis_facts` and `analyze_server` re-run fact collection and scoring over the anchored window, and `analyze_server`'s anomaly detection moves with it, so the window is compared against the hour-of-day x day-of-week baseline for the hours it actually covers instead of for the hours you happen to be asking in. `compare_analysis` hangs BOTH windows off the anchor, since `baseline_hours_back` has always been measured from the comparison window's end. `get_analysis_findings` is the odd one and worth reading twice: its window is on ANALYSIS TIME, so anchoring it asks what a scheduled analysis pass was SAYING then, which is a different question from re-analyzing that window now (that is `analyze_server` with the same anchor). + - `analyze_server` with an `as_of` is EXPLORATORY and does NOT persist its findings; the result says so in `persisted` / `persistence_note`. A finding row is stamped with the time the analysis RAN, and `get_analysis_findings` and the viewer's Recommendations tab treat the newest `analysis_time` as the server's CURRENT state — so writing a backdated run would make last week's findings today's headline and would inflate the occurrence stats of any live incident sharing a story path. Run it without `as_of` when you want the present analyzed and recorded. + - `get_pvs_stats` and `get_fleet_overview` do not take it. Each mixes a latest-snapshot measurement with a windowed one, so anchoring only the windowed half would return a result whose two halves describe different instants. `get_store_metrics` does not take it either: it windows in DAYS over the store's own growth series. ### Diagnostic-analysis tools | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `analyze_server` | Runs the inference engine: scores facts, traverses relationship graph, returns evidence-backed findings with severity and recommended next tools. A remediable finding also carries `remediation_command` — the full copy-paste T-SQL remediation (identical to the viewer card), with a two-sided risk-disclosure header on destructive changes; advisory only, never executed. Force-plan findings additionally carry `structured_remediation`: the verdict as machine-readable fields (eligible + named blockers), evidence, and split force/unforce/verify SQL | `server_name`, `hours_back` (default 4) | - | `get_analysis_facts` | Exposes raw scored facts from the collect+score pipeline — every observation the engine sees with base severity, amplifiers, and metadata | `server_name`, `hours_back` (default 4), `source` (filter), `min_severity` | - | `compare_analysis` | Compares two time periods (e.g., peak vs off-peak, before vs after a change) showing severity deltas for each fact | `server_name`, `hours_back` (default 4), `baseline_hours_back` (default 28) | + | `analyze_server` | Runs the inference engine: scores facts, traverses relationship graph, returns evidence-backed findings with severity and recommended next tools. A remediable finding also carries `remediation_command` — the full copy-paste T-SQL remediation (identical to the viewer card), with a two-sided risk-disclosure header on destructive changes; advisory only, never executed. Force-plan findings additionally carry `structured_remediation`: the verdict as machine-readable fields (eligible + named blockers), evidence, and split force/unforce/verify SQL. With `as_of` it analyzes a PAST window — anomaly baseline included — and is EXPLORATORY: the findings come back in full but are NOT persisted, which `persisted` / `persistence_note` state on every result | `server_name`, `hours_back` (default 4), `as_of` | + | `get_analysis_facts` | Exposes raw scored facts from the collect+score pipeline — every observation the engine sees with base severity, amplifiers, and metadata | `server_name`, `hours_back` (default 4), `source` (filter), `min_severity`, `as_of` | + | `compare_analysis` | Compares two time periods (e.g., peak vs off-peak, before vs after a change) showing severity deltas for each fact. When NEITHER window produced facts the result is `unavailable` rather than an all-zero comparison, because "nothing to compare" is not "nothing changed"; when only ONE window is empty the payload carries a `caveat` saying so, since every fact then counts as new or resolved by default. `baseline_hours_back` is measured from the comparison window's END, so `as_of` moves BOTH windows together | `server_name`, `hours_back` (default 4), `baseline_hours_back` (default 28), `as_of` | | `audit_config` | Edition-aware configuration audit: evaluates CTFP, MAXDOP, max memory, and max worker threads against best practices | `server_name` | - | `get_analysis_findings` | Retrieves persisted findings from previous analysis runs (the service also analyzes on its own schedule, every 30 minutes per server), deduplicated to one entry per diagnostic chain (`story_path_hash` + `incident_id`): the latest occurrence plus `occurrences`/`first_seen`/`last_seen`/`peak_severity` spanning the window; each remediable finding carries `remediation_command` — the full copy-paste T-SQL remediation (identical to the viewer card), rendered from the persisted action, advisory only and never executed; force-plan findings additionally carry `structured_remediation` (verdict + evidence + split artifacts, machine-readable) | `server_name`, `hours_back` (default 24) | + | `get_analysis_findings` | Retrieves persisted findings from previous analysis runs (the service also analyzes on its own schedule, every 30 minutes per server), deduplicated to one entry per diagnostic chain (`story_path_hash` + `incident_id`): the latest occurrence plus `occurrences`/`first_seen`/`last_seen`/`peak_severity` spanning the window; each remediable finding carries `remediation_command` — the full copy-paste T-SQL remediation (identical to the viewer card), rendered from the persisted action, advisory only and never executed; force-plan findings additionally carry `structured_remediation` (verdict + evidence + split artifacts, machine-readable). Its window is on ANALYSIS TIME, so `as_of` asks what analysis was saying then rather than re-analyzing that window now | `server_name`, `hours_back` (default 24), `as_of` | | `mute_analysis_finding` | Mutes a finding pattern by story_path_hash so it won't appear in future runs | `story_path_hash` (required), `server_name`, `reason` | ### Plan-analysis tools @@ -93,20 +111,25 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_cpu_utilization` | CPU % over time (SQL / other-process / total / idle), 1-minute averages | `server_name`, `hours_back` (default 4) | - | `get_wait_stats` | Top wait types aggregated over the window (wait/signal/resource ms, signal %) | `server_name`, `hours_back` (default 24), `limit` (default 20) | - | `get_wait_trend` | A single wait type's per-second trend over time | `wait_type` (required), `server_name`, `hours_back` (default 24) | - | `get_wait_types` | The distinct wait types observed on the server (heaviest first) — pick a `wait_type` for get_wait_trend | `server_name`, `hours_back` (default 24) | + | `get_cpu_utilization` | CPU % over time (SQL / other-process / total / idle), 1-minute averages | `server_name`, `hours_back` (default 4), `as_of` | + | `get_wait_stats` | Top wait types aggregated over the window (wait/signal/resource ms, signal %) | `server_name`, `hours_back` (default 24), `limit` (default 20), `as_of` | + | `get_wait_trend` | A single wait type's per-second trend over time | `wait_type` (required), `server_name`, `hours_back` (default 24), `as_of` | + | `get_wait_types` | The distinct wait types observed on the server (heaviest first) — pick a `wait_type` for get_wait_trend. An empty result distinguishes a quiet window (`empty`, widen `hours_back`) from a server no wait stats have ever been stored for (`unavailable`) | `server_name`, `hours_back` (default 24), `as_of` | | `get_memory_stats` | Latest memory snapshot: physical / buffer pool / plan cache / utilization %, memory model | `server_name` | - | `get_memory_clerks` | Latest top memory consumers by clerk type | `server_name` | + | `get_memory_clerks` | Latest top memory consumers by clerk type. An empty result is `unavailable`, never a quiet period — a live SQL Server always has clerks, so nothing retained means the collector has not run or its rows aged out | `server_name` | | `get_file_io_stats` | Latest per-file I/O: reads/writes/bytes/stall and computed read/write latency | `server_name` | - | `get_tempdb_trend` | TempDB space over time (user / internal / version store / unallocated) + top consumer | `server_name`, `hours_back` (default 24) | + | `get_tempdb_trend` | TempDB space over time (user / internal / version store / unallocated) + top consumer | `server_name`, `hours_back` (default 24), `as_of` | | `get_perfmon_stats` | Latest perfmon counters (value + delta); filter by counter / instance | `server_name`, `counter_name`, `instance_name` | - | `get_top_queries_by_cpu` | Expensive queries from query stats (plan cache) with query_hash / sql_handle; `cpu_attribution.attributed_cpu_ratio` says how much of the box's measured CPU the returned rows explain | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name`, `parallel_only`, `min_dop` | - | `get_top_procedures_by_cpu` | Most expensive stored procedures by total CPU, with the same `cpu_attribution` disclosure | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name` | - | `get_query_store_top` | Expensive queries from Query Store with query_id / plan_id (survives restarts) | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name` | + | `get_top_queries_by_cpu` | Expensive queries from query stats (plan cache) with query_hash / sql_handle; `cpu_attribution.attributed_cpu_ratio` says how much of the box's measured CPU the returned rows explain | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name`, `parallel_only`, `min_dop`, `as_of` | + | `get_top_procedures_by_cpu` | Most expensive stored procedures by total CPU, with the same `cpu_attribution` disclosure | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name`, `as_of` | + | `get_query_store_top` | Expensive queries from Query Store with query_id / plan_id (survives restarts) | `server_name`, `hours_back` (default 24), `top` (default 20), `database_name`, `as_of` | + | `get_query_heatmap` | The desktop viewer's Query Heatmap as a TABLE: how many DISTINCT queries fell into each (time bin x log-magnitude bucket) cell, with the most-executed query in each cell. The only query read with a TIME axis — `get_top_queries_by_cpu` ranks a whole window and cannot show that the window had a quiet half and a bad half, which is the first question about an incident that has already ended. Bins are **5 minutes** wide by default because that is exactly what the desktop viewer uses, so both surfaces draw the same picture; raise `bucket_minutes` to cover a longer window in fewer cells (it is the lever to reach for before the cap). The seven magnitude buckets are the viewer's, in the metric's own unit, and the labels come back with the result. `limit` caps CELLS, and truncation drops the OLDEST bins rather than the least interesting cells — `first_time_bin` / `last_time_bin` say which slice came back. Zero cells is THREE states and the read says which: never collected (`unavailable`, nobody looked), nothing collected in the window (`empty`, widen it), or collected and genuinely idle — captures exist and every one recorded zero executions (`empty`) | `server_name`, `hours_back` (default 24), `metric` (default `duration`), `database_name`, `bucket_minutes` (default 5), `limit` (default 500), `as_of` | + | `get_query_store_regressions` | Queries whose Query Store performance got WORSE: each (database, query_id) group's averages inside the recent window vs its BASELINE — every capture collected BEFORE that window. Baseline vs recent duration / CPU / logical reads with a regression percent each, the execution-count-weighted `additional_duration_ms` (the ranking key: a 5 ms regression run a million times outranks a 5-second one run twice), the plan counts on both sides, and a severity band. `get_query_store_top` answers what is EXPENSIVE and the costliest query is usually the one that always was; this answers what CHANGED. Kept only where average CPU regressed > 25%. Zero rows is FOUR states and the read says which: never collected (`unavailable`), no BASELINE because all history falls inside the window (`unavailable`, and NOT a clean bill of health — shorten hours_back), nothing collected in the window (`empty`, widen it), or a genuine all-clear (`empty`) | `server_name`, `hours_back` (default 24), `database_name`, `limit` (default 50), `as_of` | | `list_servers` | All monitored servers with collection-freshness status and last collection time, plus `peer_fleets` — the declared SIBLING Darling stores and what each covers (disclosure only; this server cannot read them) and `peer_note`, which says what an EMPTY `peer_fleets` does and does not prove | none | - | `get_collection_health` | Per-collector health (running / failing / stale) over the last 7 days, plus the server's sweep_pressure verdict (a SATURATED body collects at a multiple of its configured cadence with every collector healthy) | `server_name` | + | `get_collection_health` | Per-collector health (running / failing / stale) over the last 7 days, plus the server's sweep_pressure block: a `verdict` for SUSTAINED demand (a SATURATED body collects at a multiple of its configured cadence with every collector healthy) and a separate `peak_cycle_risk` for a SINGLE sweep (BODY_OVERRUN means one scheduled body cannot fit the budget even when the verdict reads OK, the signature of one infrequent heavy collector; `peak_collector` names it). Per-collector rows carry `avg_duration_ms`, `p95_duration_ms` and `max_duration_ms`: a mean far below the p95 means the collector's runs come in two sizes and the mean describes neither | `server_name` | + | `get_collection_log` | The RAW per-run collection log behind that rollup: one row per collector run with total duration split into time on the monitored server and time on the store, rows collected, status and any error. Reach for it when the rollup reads HEALTHY and collection still looks wrong, or to see what a collector was doing during a specific window. An empty result distinguishes a quiet window (`empty`, widen it) from a server that has never collected (`unavailable`, collection is not running) | `server_name`, `hours_back`, `limit`, `as_of` | + | `get_current_waits_trend` | The two Current Waits series over time: waiting-task total wait per wait type per collection, and blocked-session counts per database per collection. `get_waiting_tasks` gives the snapshot and can never say whether now is worse than an hour ago; this is that question. Read the two series together — a wait-type spike with no blocked sessions is a resource wait, the same spike with them is contention. An empty result distinguishes a genuine all-clear (`empty`) from a server the collector has never sampled (`unavailable`), which is NOT an all-clear | `server_name`, `hours_back`, `database_name`, `as_of` | + | `get_blocking_stats` | Blocking SEVERITY per minute: blocking duration (event count, total, max, avg wait) and deadlock severity (victim count plus total/max/avg wait across EVERY process in the graphs, not just victims). `get_blocking_trend` and `get_deadlock_trend` say how OFTEN; this says how BAD — ten one-second blocks and one ten-minute block are the same count and a different problem. An empty result distinguishes a genuinely clear window (`empty`) from a server where neither capture path has ever produced a row (`unavailable`), which is NOT a clean bill of health | `server_name`, `hours_back`, `as_of` | | `get_server_properties` | Instance properties: edition, version, CPU count, memory, socket/core topology, HADR | `server_name` | ### Diagnostic-depth data-read tools @@ -115,19 +138,20 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_blocking` | Recent blocked/blocking pairs from the blocked-process-report XE + the always-on DMV fallback | `server_name`, `hours_back` (default 24), `limit` (default 30), `dedup_key` (optional) | - | `get_deadlocks` | Recent deadlocks: victim process/SQL + a process summary | `server_name`, `hours_back` (default 24), `limit` (default 20), `dedup_key` (optional) | - | `get_deadlock_detail` | The raw deadlock graph XML for the recent deadlocks | `server_name`, `hours_back` (default 24), `limit` (default 5), `dedup_key` (optional) | - | `get_blocked_process_xml` | The raw blocked-process-report XML | `server_name`, `hours_back` (default 24), `limit` (default 5) | - | `get_long_query_completions` | Longest completed queries (rpc/batch over the trace threshold) + attentions/cancels from the opt-in long-query trace, duration DESC (empty until the collector is enabled) | `server_name`, `hours_back` (default 24), `limit` (default 30) | - | `get_blocking_trend` | Per-minute blocking-incident counts over time (XE, DMV-snapshot fallback) | `server_name`, `hours_back` (default 24) | - | `get_deadlock_trend` | Per-minute deadlock counts over time | `server_name`, `hours_back` (default 24) | + | `get_blocking` | Recent blocked/blocking pairs from the blocked-process-report XE + the always-on DMV fallback | `server_name`, `hours_back` (default 24), `limit` (default 30), `dedup_key` (optional), `as_of` | + | `get_deadlocks` | Recent deadlocks: victim process/SQL + a process summary | `server_name`, `hours_back` (default 24), `limit` (default 20), `dedup_key` (optional), `as_of` | + | `get_deadlock_detail` | The raw deadlock graph XML for the recent deadlocks | `server_name`, `hours_back` (default 24), `limit` (default 5), `dedup_key` (optional), `as_of` | + | `get_blocked_process_xml` | The raw blocked-process-report XML | `server_name`, `hours_back` (default 24), `limit` (default 5), `as_of` | + | `get_long_query_completions` | Longest completed queries (rpc/batch over the trace threshold) + attentions/cancels from the opt-in long-query trace, duration DESC (empty until the collector is enabled) | `server_name`, `hours_back` (default 24), `limit` (default 30), `as_of` | + | `get_blocking_trend` | Per-minute blocking-incident counts over time (XE, DMV-snapshot fallback). An empty result distinguishes a genuine all-clear (`empty`, with the collector run counts in `hints` so you can see how many captures the window actually holds) from a window no collector covered (`unavailable`), which is NOT an all-clear | `server_name`, `hours_back` (default 24), `as_of` | + | `get_deadlock_trend` | Per-minute deadlock counts over time. An empty result distinguishes a genuine all-clear (`empty`, with the collector run counts in `hints` so you can see how many captures the window actually holds) from a window no collector covered (`unavailable`), which is NOT an all-clear | `server_name`, `hours_back` (default 24), `as_of` | + | `get_lock_wait_trend` | Every LCK% wait type's wait milliseconds per SECOND at each collection — the aggregate lock-wait lane. The two trends above count incidents and `get_wait_trend` charts ONE named wait type; this is the whole lock family as a rate, which is what shows lock pressure rising when no single type dominates. Rate rather than raw delta, so it compares across servers on different cadences. An empty result distinguishes a genuinely quiet window (`empty`, widen `hours_back`) from a server no wait stats have ever been stored for (`unavailable`), which is NOT a report of a server without lock contention | `server_name`, `hours_back` (default 24), `as_of` | | `get_session_stats` | Latest per-application connection/session counts (running/sleeping/dormant) + resource totals | `server_name` | - | `get_active_queries` | Captured running-query snapshots over the window (waits, CPU, blocking, grants) | `server_name`, `hours_back` (default 1), `database_name`, `blocking_only`, `limit` (default 50) | - | `get_waiting_tasks` | Individual waiting tasks captured at collection time | `server_name`, `hours_back` (default 1), `limit` (default 30) | - | `get_server_config_changes` | sp_configure changes, diffed from config snapshots | `server_name`, `hours_back` (default 168) | - | `get_database_config_changes` | sys.databases setting changes, diffed from config snapshots | `server_name`, `hours_back` (default 168) | - | `get_trace_flag_changes` | Trace flags enabled/disabled/modified, diffed from config snapshots | `server_name`, `hours_back` (default 168) | + | `get_active_queries` | Captured running-query snapshots over the window (waits, CPU, blocking, grants) | `server_name`, `hours_back` (default 1), `database_name`, `blocking_only`, `limit` (default 50), `as_of` | + | `get_waiting_tasks` | Individual waiting tasks captured at collection time | `server_name`, `hours_back` (default 1), `limit` (default 30), `as_of` | + | `get_server_config_changes` | sp_configure changes, diffed from config snapshots | `server_name`, `hours_back` (default 168), `as_of` | + | `get_database_config_changes` | sys.databases setting changes, diffed from config snapshots | `server_name`, `hours_back` (default 168), `as_of` | + | `get_trace_flag_changes` | Trace flags enabled/disabled/modified, diffed from config snapshots | `server_name`, `hours_back` (default 168), `as_of` | | `get_database_scoped_config` | Latest database-scoped configuration (MAXDOP, legacy CE, ...) | `server_name`, `database_name` | | `get_query_store_health` | Per-database Query Store health (latest hourly snapshot) — actual vs desired state, readonly_reason decoded, storage vs cap, cleanup thresholds | `server_name`, `database_name` | | `get_server_config` | CURRENT sys.configurations (latest snapshot) — what CTFP / MAXDOP / max memory are set to now | `server_name` | @@ -144,12 +168,12 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_latch_stats` | Top latch classes by wait time, with per-second rates + severity / description / recommendation | `server_name`, `hours_back` (default 24), `top` (default 10) | - | `get_spinlock_stats` | Top spinlocks by collisions, with per-second rates + description | `server_name`, `hours_back` (default 24), `top` (default 10) | - | `get_resource_semaphore` | Latest workspace-memory semaphores: target / max-target ceiling vs granted / used, waiter / timeout / forced | `server_name`, `hours_back` (default 24) | - | `get_memory_grants` | Latest per-pool grant detail: available / granted / used + waiter / timeout / forced deltas | `server_name`, `hours_back` (default 1) | - | `get_memory_pressure_events` | RING_BUFFER_RESOURCE_MONITOR memory-pressure notifications (process/system indicator scale 0-3+); not on Azure SQL DB | `server_name`, `hours_back` (default 24) | - | `get_plan_cache_bloat` | Plan cache single-use vs multi-use composition + bloat_level classification | `server_name`, `hours_back` (default 24) | + | `get_latch_stats` | Top latch classes by wait time, with per-second rates + severity / description / recommendation | `server_name`, `hours_back` (default 24), `top` (default 10), `as_of` | + | `get_spinlock_stats` | Top spinlocks by collisions, with per-second rates + description | `server_name`, `hours_back` (default 24), `top` (default 10), `as_of` | + | `get_resource_semaphore` | Latest workspace-memory semaphores: target / max-target ceiling vs granted / used, waiter / timeout / forced | `server_name`, `hours_back` (default 24), `as_of` | + | `get_memory_grants` | Latest per-pool grant detail: available / granted / used + waiter / timeout / forced deltas | `server_name`, `hours_back` (default 1), `as_of` | + | `get_memory_pressure_events` | RING_BUFFER_RESOURCE_MONITOR memory-pressure notifications (process/system indicator scale 0-3+); not on Azure SQL DB | `server_name`, `hours_back` (default 24), `as_of` | + | `get_plan_cache_bloat` | Plan cache single-use vs multi-use composition + bloat_level classification | `server_name`, `hours_back` (default 24), `as_of` | | `get_cpu_scheduler_pressure` | Latest scheduler snapshot: runnable queue, worker utilization, pressure_level + warnings | `server_name` | | `get_running_jobs` | Currently running SQL Agent jobs with duration vs historical average / p95 | `server_name` | @@ -159,11 +183,13 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_memory_trend` | Memory usage over time: total / target server memory, buffer pool, plan cache | `server_name`, `hours_back` (default 24) | - | `get_perfmon_trend` | A single performance counter's value + delta over time (summed across instances) | `counter_name` (required), `server_name`, `hours_back` (default 24) | - | `get_file_io_trend` | Per-database file I/O read/write latency over time (top-10 busiest files) | `server_name`, `hours_back` (default 24) | - | `get_query_trend` | One query's per-collection history (deltas, avg cpu/elapsed, DOP) by query_hash | `query_hash` (required), `database_name` (required), `server_name`, `hours_back` (default 24) | - | `get_query_duration_trend` | Overall query elapsed-ms/sec + executions/sec across all queries over time | `server_name`, `hours_back` (default 24) | + | `get_memory_trend` | Memory usage over time: total / target server memory, buffer pool, plan cache. An empty result distinguishes a quiet window (`empty`, widen `hours_back`) from a server nothing has ever been collected for (`unavailable`, collection is not running) | `server_name`, `hours_back` (default 24), `as_of` | + | `get_perfmon_trend` | A single performance counter's value + delta over time (summed across instances) | `counter_name` (required), `server_name`, `hours_back` (default 24), `as_of` | + | `get_file_io_trend` | Per-database file I/O read/write latency over time (top-10 busiest files). An empty result distinguishes a quiet window (`empty`, widen `hours_back`) from a server nothing has ever been collected for (`unavailable`, collection is not running) | `server_name`, `hours_back` (default 24), `as_of` | + | `get_query_trend` | One query's per-collection history (deltas, avg cpu/elapsed, DOP) by query_hash | `query_hash` (required), `database_name` (required), `server_name`, `hours_back` (default 24), `as_of` | + | `get_query_duration_trend` | Overall query elapsed-ms/sec + executions/sec across all queries over time, from the PLAN CACHE. Each point carries `value` (ms/sec), `execution_count` and `executions_per_second` — the last two are the same quantity, and `execution_count` is truncated to an integer, so read `executions_per_second` on a quiet server where the rate is below 1. An empty result distinguishes a quiet window (`empty`, widen `hours_back`) from a server nothing has ever been collected for (`unavailable`, collection is not running) | `server_name`, `hours_back` (default 24), `as_of` | + | `get_procedure_duration_trend` | The same series over `procedure_stats`. NOT a duplicate of the above: query_stats attributes a procedure's work to the individual statements inside it, so a procedure that got slower is smeared across however many statements it runs — this charges the whole call to the procedure. Read the two together to tell an ad-hoc regression from a procedure regression | `server_name`, `hours_back` (default 24), `as_of` | + | `get_query_store_duration_trend` | The same series over Query Store. The plan-cache trends lose everything an eviction or a restart takes with them; Query Store persists per interval, so this is the series that survives a failover and the one to reach for when a regression is older than the cache. Each interval is counted once, at the hour the work RAN. Its `unavailable` names the cause the other two do not have: Query Store may simply be OFF on every database | `server_name`, `hours_back` (default 24), `as_of` | ### System-health parse-on-read tools @@ -171,14 +197,15 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_health_parser_system_health` | SYSTEM-component snapshots: corruption (bad pages / dumps / access violations) + contention (non-yielding / latch / sick spinlock / CPU) counters | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_severe_errors` | error_reported severity >= 19 (excl. 17830 / 18056), with database_id resolved to a name | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_scheduler_issues` | Non-yielding / offline scheduler WARNING rows | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_memory_conditions` | RESOURCE_MEMPHYSICAL_LOW memory-pressure snapshots | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_memory_broker` | RESOURCE_MEMPHYSICAL_LOW memory-broker ratio changes | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_memory_node_oom` | Every recorded per-NUMA-node out-of-memory event (never gated) | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_cpu_tasks` | QUERY_PROCESSING WARNING rows with pendingTasks >= 10 (worker-thread exhaustion) | `server_name`, `hours_back` (default 24), `limit` (default 50) | - | `get_health_parser_io_issues` | IO_SUBSYSTEM WARNING rows (15-second I/O warnings), one per pending-request file | `server_name`, `hours_back` (default 24), `limit` (default 50) | + | `get_health_parser_system_health` | SYSTEM-component snapshots: corruption (bad pages / dumps / access violations) + contention (non-yielding / latch / sick spinlock / CPU) counters | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_severe_errors` | error_reported severity >= 19 (excl. 17830 / 18056), with database_id resolved to a name | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_scheduler_issues` | Non-yielding / offline scheduler WARNING rows | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_memory_conditions` | RESOURCE_MEMPHYSICAL_LOW memory-pressure snapshots | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_memory_broker` | RESOURCE_MEMPHYSICAL_LOW memory-broker ratio changes | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_memory_node_oom` | Every recorded per-NUMA-node out-of-memory event (never gated) | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_cpu_tasks` | QUERY_PROCESSING WARNING rows with pendingTasks >= 10 (worker-thread exhaustion) | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_io_issues` | IO_SUBSYSTEM WARNING rows (15-second I/O warnings), one per pending-request file | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | + | `get_health_parser_significant_waits` | Individual wait_info events: a real session's non-BACKUP statement waited 500 ms+ on a non-idle wait type. Returns wait type, duration and signal duration, wait resource, session id and the waiting SQL text — `get_wait_stats` gives the instance-wide totals and can never name the statement that paid them. An empty result says WHICH nothing it is: events captured but none significant (the healthy answer), a quiet window (`empty`, widen it), or a server whose wait_info has never been captured (`unavailable`, NOT an all-clear) | `server_name`, `hours_back` (default 24), `limit` (default 50), `as_of` | ### Alert + health-overview tools @@ -186,11 +213,12 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_alert_history` | Alerts that fired (metric, value vs threshold, delivery success/failure, muted); omit server_name for the whole fleet, each row names its server | `server_name` (optional — all servers if omitted), `hours_back` (default 24), `limit` (default 50) | + | `get_alert_history` | Alerts that fired (metric, value vs threshold, delivery success/failure, muted); omit server_name for the whole fleet, each row names its server | `server_name` (optional — all servers if omitted), `hours_back` (default 24), `limit` (default 50), `as_of` | | `get_alert_settings` | The current alert config the service uses: per-alert enable + thresholds, cooldown, excluded databases, delivery mode, and the scheduled-analysis cadence | none | - | `get_mute_rules` | The alert mute rules in force, so a suppressed server is distinguishable from a healthy-quiet one | `enabled_only` (default true) | + | `get_mute_rules` | The alert mute rules in force, so a suppressed server is distinguishable from a healthy-quiet one. An empty result distinguishes no rule ever written from rules that exist but have all lapsed, with the configured count in `hints` | `enabled_only` (default true) | | `get_server_summary` | One-shot per-server health: current CPU %, memory, recent blocking count, recent deadlock count | `server_name` | | `get_daily_summary` | A day's composite health band (Healthy / Warning / Critical) plus the signals behind it (waits, deadlocks, blocking, high CPU, memory pressure, alerts) | `server_name`, `summary_date` (yyyy-MM-dd, default today) | + | `get_daily_summary_range` | The SAME rollup across a span of days — one row per collected day, which is the desktop viewer's Performance Calendar month grid. Use it when the question is WHICH day rather than how one day went: scan the bands, then call `get_daily_summary` for the day that stands out. A day with ANY collection appears even when every signal was quiet (Healthy, not missing), so a day absent from the result is a gap in COLLECTION. `as_of` anchors the LAST day of the range, so a past month is `as_of` its last day with `days_back` its length. An empty result distinguishes a range outside this server's history (`empty`) from a server nothing has ever been collected for (`unavailable`) | `server_name`, `days_back` (default 30, max 366), `as_of` | **Tuning the alerting (write).** Three Darling-only tools change the shared alert configuration the service delivers on — the SAME config `get_alert_settings` / `get_mute_rules` read, and the same the Viewer's Settings window writes. They are the only alert writes here; none touches a monitored server or the collected data, and a change hot-reloads into the running service within one collection sweep. @@ -229,7 +257,7 @@ public static string Build(DarlingPeerDirectory.Snapshot peers) | Tool | Purpose | Key Parameters | |------|---------|----------------| - | `get_default_trace_events` | Significant Default Trace events — auto-grow/shrink stalls, severe ErrorLog, schema DDL, security audits, memory change — each categorized (config-change events excluded) | `server_name`, `hours_back` (default 24), `limit` (default 100) | + | `get_default_trace_events` | Significant Default Trace events — auto-grow/shrink stalls, severe ErrorLog, schema DDL, security audits, memory change — each categorized (config-change events excluded) | `server_name`, `hours_back` (default 24), `limit` (default 100), `as_of` | ### Custom Views (create & manage) diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpJobTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpJobTools.cs index 5390cca2a..603fdb3b9 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpJobTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpJobTools.cs @@ -40,7 +40,33 @@ public static async Task GetRunningJobs( { var rows = await DarlingJobReader.GetRunningJobsAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("empty", "No running SQL Agent jobs found (or the running_jobs collector has not run yet)."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "running_jobs") + /* #2546: the msdb case. "No running SQL Agent jobs found" is an affirmative claim about + the server's Agent, and it is the wrong one when the monitoring login was refused the + job tables — the collector runs, is denied, and records that denial with the GRANT to + issue. Reporting it here is the difference between "nothing is running" and "we cannot + see what is running". */ + ?? await DarlingRuntimePrecondition.StatusAsync(postgres, resolved.ServerId, resolved.ServerName, "running_jobs") + /* #2559: the case the line above cannot see. StatusAsync reports what the collector's last + run RECORDED, and a collector whose AppliesTo gate is off never runs — the runner returns + before writing any collection_log row, deliberately, because a per-cycle fake row was + thousands of rows a day of noise. So the gated-off server produced no evidence, this fell + through to the "empty" line below, and we went back to asserting the Agent is idle on a + server we were never permitted to look at. That is the exact claim #2546 set out to + remove, surviving in the one case that records nothing to read. + + The gate is !IsAzureSqlDb && !IsAwsRds. The engine half is already answered above, + so AWS RDS is the only remaining candidate and the message can name it outright rather + than hedging — #2559 removed HasMsdbAccess from this gate, which is what turned two + unpersisted candidates into one. */ + ?? await DarlingRuntimePrecondition.GatedOffStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "running_jobs", + "For this collector the gate is: this is an AWS RDS instance, where the Agent job " + + "tables are not reachable to a monitoring login at all and no grant changes that. " + + "Since #2559 msdb access is NOT a gate — a login without it now attempts and is " + + "reported as a permission denial, so the grant takes effect on the next cycle " + + "rather than the next reconnect.") + ?? McpHelpers.Status("empty", "No running SQL Agent jobs found (or the running_jobs collector has not run yet)."); var jobs = rows.Select(r => new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLatchSpinlockTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLatchSpinlockTools.cs index 74384728a..843650aa1 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLatchSpinlockTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLatchSpinlockTools.cs @@ -42,23 +42,25 @@ public static async Task GetLatchStats( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of data to analyze. Default 24.")] int hours_back = 24, - [Description("Number of top latch classes to return. Default 10.")] int top = 10) + [Description("Number of top latch classes to return. Default 10.")] int top = 10, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(top); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingLatchSpinlockReader.GetLatchStatsTopNAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, top); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No latch statistics available in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "latch_stats") + ?? McpHelpers.Status("unavailable", "No latch statistics available in the requested time range."); var latches = rows.Select(r => new { @@ -95,23 +97,25 @@ public static async Task GetSpinlockStats( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of data to analyze. Default 24.")] int hours_back = 24, - [Description("Number of top spinlocks to return. Default 10.")] int top = 10) + [Description("Number of top spinlocks to return. Default 10.")] int top = 10, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(top); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingLatchSpinlockReader.GetSpinlockStatsTopNAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, top); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No spinlock statistics available in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "spinlock_stats") + ?? McpHelpers.Status("unavailable", "No spinlock statistics available in the requested time range."); var spinlocks = rows.Select(r => new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLongQueryTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLongQueryTools.cs index 735405449..709456b31 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLongQueryTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpLongQueryTools.cs @@ -34,23 +34,30 @@ public static async Task GetLongQueryCompletions( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, - [Description("Maximum rows. Default 30.")] int limit = 30) + [Description("Maximum rows. Default 30.")] int limit = 30, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingLongQueryReader.GetRecentLongQueryCompletionsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No long-running query completions found in the specified time range. The long_query_completions collector is opt-in (default OFF) — enable it in the collector schedule to capture data."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "long_query_completions") + /* #2546: this collector is opt-in, so the fall-through below already sends the reader to + the schedule — which is the wrong place when the collector IS enabled and its session + is missing. The precondition answer names that state instead of quietly blaming a knob + that is already switched on. */ + ?? await DarlingRuntimePrecondition.StatusAsync(postgres, resolved.ServerId, resolved.ServerName, "long_query_completions") + ?? McpHelpers.Status("empty", "No long-running query completions found in the specified time range. The long_query_completions collector is opt-in (default OFF) — enable it in the collector schedule to capture data."); var result = rows.Take(limit).Select(r => new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpMemoryGrantTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpMemoryGrantTools.cs index fcadf76f0..f40694756 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpMemoryGrantTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpMemoryGrantTools.cs @@ -34,21 +34,23 @@ public sealed class DarlingMcpMemoryGrantTools public static async Task GetResourceSemaphore( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingMemoryGrantReader.GetResourceSemaphoreLatestAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No memory grant data available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_grant_stats") + ?? McpHelpers.Status("unavailable", "No memory grant data available."); var grants = rows.Select(r => new { @@ -85,21 +87,23 @@ public static async Task GetResourceSemaphore( public static async Task GetMemoryGrants( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 1.")] int hours_back = 1) + [Description("Hours of history. Default 1.")] int hours_back = 1, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingMemoryGrantReader.GetMemoryGrantsLatestAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No memory grant data available."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_grant_stats") + ?? McpHelpers.Status("unavailable", "No memory grant data available."); var grants = rows.Select(r => new { @@ -140,21 +144,23 @@ Not available on Azure SQL DB (ring buffer not exposed).")] public static async Task GetMemoryPressureEvents( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingMemoryGrantReader.GetMemoryPressureEventsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No memory pressure events found in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_pressure_events") + ?? McpHelpers.Status("empty", "No memory pressure events found in the requested time range."); return JsonSerializer.Serialize(new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpObjectStatsTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpObjectStatsTools.cs index a07a55058..0cef25cd8 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpObjectStatsTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpObjectStatsTools.cs @@ -49,7 +49,8 @@ public static async Task GetTableIndexSizes( var rows = await DarlingObjectStatsReader.GetObjectSizeGrowthAsync( postgres, resolved.ServerId, now.AddDays(-7), now.AddDays(-30), TableSizesTop); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No object size data available. Index/object stats are collected daily."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "index_object_stats") + ?? McpHelpers.Status("unavailable", "No object size data available. Index/object stats are collected daily."); var result = rows.Select(r => new { @@ -78,19 +79,54 @@ public static async Task GetTableIndexSizes( } } - [McpServerTool(Name = "get_index_usage"), Description("Gets per-index usage (seeks, scans, lookups, updates) from the latest daily snapshot, classifying each index as Unused, Write-only, or Active. Unused and write-only indexes are listed first - these are drop candidates. Counters are cumulative since the last instance restart.")] + [McpServerTool(Name = "get_index_usage"), Description("Gets per-index usage (seeks, scans, lookups, updates) from the latest daily snapshot, classifying each index as Unused, Write-only, or Active. Unused and write-only indexes are listed FIRST because they are drop candidates - which means that on a server with many unused indexes the row limit can be filled entirely by one database's unused indexes, hiding every Active index elsewhere. Pass database_name to ask about one database, which is almost always what you want; the response carries matching_index_count and truncated so a short answer is never mistaken for an absent one. Counters are cumulative since the last instance restart.")] public static async Task GetIndexUsage( NpgsqlDataSource postgres, - [Description("Server name or display name.")] string? server_name = null) + [Description("Server name or display name.")] string? server_name = null, + [Description("Limit to one database. Strongly recommended: without it, unused-first ordering can fill the whole result from one database.")] string? database_name = null, + [Description("Maximum rows to return. Default 200.")] int limit = IndexUsageTop) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; + var validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + try { - var rows = await DarlingObjectStatsReader.GetIndexUsageAsync(postgres, resolved.ServerId, IndexUsageTop); + var database = string.IsNullOrWhiteSpace(database_name) ? null : database_name; + + var rows = await DarlingObjectStatsReader.GetIndexUsageAsync(postgres, resolved.ServerId, limit, database); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No index usage data available. Index/object stats are collected daily."); + { + /* #2636: a database filter that matches nothing is a DIFFERENT answer from a server that + collects no index stats, and the reporter hit the first while being told the second. The + capability check still runs first — a wrong-engine target has no index_object_stats at all + — and only then does the filter get blamed for its own empty result. */ + if (database is not null) + { + var anyOnServer = await DarlingObjectStatsReader.GetIndexUsageMatchCountAsync(postgres, resolved.ServerId); + + if (anyOnServer > 0) + { + return McpHelpers.Status( + "empty", + $"No index usage rows for database '{database}' on {resolved.ServerName} at the " + + $"latest snapshot, though the server has {anyOnServer:N0} across its other " + + "databases. Check the database name against get_database_sizes — the filter " + + "matches exactly, and an excluded or renamed database looks identical to one " + + "with no indexes."); + } + } + + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "index_object_stats") + ?? McpHelpers.Status("unavailable", "No index usage data available. Index/object stats are collected daily."); + } + + /* Counted BEFORE the cap, by a second query. A count taken over the returned rows is a count of + the page, which is the whole defect this answers. */ + var matching = await DarlingObjectStatsReader.GetIndexUsageMatchCountAsync(postgres, resolved.ServerId, database); + var truncated = matching > rows.Count; var result = rows.Select(r => new { @@ -113,6 +149,17 @@ public static async Task GetIndexUsage( return JsonSerializer.Serialize(new { server = resolved.ServerName, + database_name = database, + returned_index_count = rows.Count, + matching_index_count = matching, + truncated, + note = truncated + ? $"TRUNCATED: {matching:N0} indexes match and {rows.Count:N0} were returned. Rows are " + + "ordered UNUSED FIRST across the whole server, so the ones omitted are the ACTIVE " + + "indexes and they may be concentrated in databases with no rows here at all. This is " + + "not evidence that a database was not collected — pass database_name to ask about " + + "one, or raise limit." + : "Complete: every index matching this filter at the latest snapshot is included.", indexes = result }, McpHelpers.JsonOptions); } @@ -134,7 +181,8 @@ public static async Task GetObjectLocking( { var rows = await DarlingObjectStatsReader.GetIndexLockingAsync(postgres, resolved.ServerId, ObjectLockingTop); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No locking/contention data recorded. Index/object stats are collected daily."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "index_object_stats") + ?? McpHelpers.Status("unavailable", "No locking/contention data recorded. Index/object stats are collected daily."); var result = rows.Select(r => new { @@ -178,7 +226,8 @@ public static async Task GetDatabaseSizes( { var rows = await DarlingObjectStatsReader.GetLatestDatabaseSizesAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No database size data available. The size collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "database_size_stats") + ?? McpHelpers.Status("unavailable", "No database size data available. The size collector may not have run yet."); return JsonSerializer.Serialize(new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgAutovacuumTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgAutovacuumTools.cs index ea6d85e86..5a092a3c6 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgAutovacuumTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgAutovacuumTools.cs @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -61,24 +62,35 @@ public static async Task GetPgAutovacuumHealth( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze, used for the dead-tuple growth comparison. Default 24.")] int hours_back = 24, - [Description("Maximum tables to return, worst first. Default 20.")] int limit = 20) + [Description("Maximum tables to return, worst first. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgAutovacuumReader.GetPgAutovacuumAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, limit); if (rows.Count == 0) { + /* "No table on this server has dead tuples" is the healthy answer for a PostgreSQL target + and a fabricated one for a SQL Server target — the collector has never run there. Ask the + engine before making the claim (#2532). */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_autovacuum_stats"); + if (gated != null) + { + return gated; + } + return JsonSerializer.Serialize(new { server = resolved.ServerName, diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgBlockingTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgBlockingTools.cs index fc22c91e9..9253c7aba 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgBlockingTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgBlockingTools.cs @@ -7,13 +7,16 @@ */ using System; +using System.Collections.Generic; using System.ComponentModel; +using System.Globalization; using System.Linq; using System.Text.Json; using System.Threading.Tasks; using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -73,17 +76,18 @@ internal static string RemedyFor(string? rootState, bool idleInTransaction, long + "client that stopped reading."; } - [McpServerTool(Name = "get_pg_blocking"), Description("Gets PostgreSQL blocking chains that were captured for a server, assembled from the stored edge list into one entry per chain with its ROOT blocker identified and attributed. Use this when sessions are timing out, waiting, or piling up on a PostgreSQL target, or to check whether a past slowdown involved lock contention. Reports for each captured chain: the root blocker's pid, state, application, username and query text, how many sessions were behind it in total and directly, how deep the chain went, the longest-waiting victim, and how many separate captures that same backend has been the root of - which distinguishes one stuck session from a recurring pattern. Also returns a specific remedy per root state, because an 'idle in transaction' root is an application defect while an 'active' root is a query-tuning problem and the two need opposite responses. IMPORTANT: this is a periodic SAMPLE, not an event log. Unlike SQL Server's blocked-process report, PostgreSQL records nothing on its own, so blocking shorter than the collection interval is never seen and an empty result means 'none was sampled', not 'none happened' - the capture counts in the response say how many samples the window actually contains. Works on any PostgreSQL target including standbys.")] + [McpServerTool(Name = "get_pg_blocking"), Description("Gets PostgreSQL blocking chains that were captured for a server, assembled from the stored edge list into one entry per chain with its ROOT blocker identified and attributed. Use this when sessions are timing out, waiting, or piling up on a PostgreSQL target, or to check whether a past slowdown involved lock contention. Reports for each captured chain: the root blocker's pid, state, application, username and query text, how many sessions were behind it in total and directly, how deep the chain went, the longest-waiting victim, and how many separate captures that same backend has been the root of - which distinguishes one stuck session from a recurring pattern. Also returns a specific remedy per root state, because an 'idle in transaction' root is an application defect while an 'active' root is a query-tuning problem and the two need opposite responses. IMPORTANT: this is a periodic SAMPLE, not an event log. Unlike SQL Server's blocked-process report, PostgreSQL records nothing on its own, so blocking shorter than the collection interval is never seen and an empty result means 'none was sampled', not 'none happened' - the capture counts in the response say how many samples the window actually contains. root_backend_id is returned as a STRING, not a number, and it is the value to compare a root blocker across captures with: it is a 64-bit composite of the backend's start time and its pid, always well past what a JSON number survives, so a numeric wire form would round DIFFERENT backends onto the same id rather than merely lose one. Works on any PostgreSQL target including standbys.")] public static async Task GetPgBlocking( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, - [Description("Maximum chains to return, worst-first by victim count. Default 50.")] int limit = 50) + [Description("Maximum chains to return, worst-first by victim count. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; var limitValidation = McpHelpers.ValidateTop(limit); @@ -91,7 +95,7 @@ public static async Task GetPgBlocking( try { - var now = DateTime.UtcNow; + var now = windowEnd; var startUtc = now.AddHours(-hours_back); var chains = await DarlingPgBlockingReader.GetPgBlockingChainsAsync( @@ -107,27 +111,7 @@ tool would report "no blocking" from a capture that recorded a deadlock. */ var cycles = await DarlingPgBlockingReader.GetPgBlockingCyclesAsync( postgres, resolved.ServerId, startUtc, now, limit); - var cycleEntries = cycles.Select(c => new - { - captured_at = c.CapturedAt, - participant_count = c.ParticipantCount, - pids = c.Pids, - database = c.DatabaseName, - application = c.ApplicationName, - /* Sessions queued behind the deadlock without being part of it. Previously invisible to - BOTH reads — chains cannot see them (no cycle member qualifies as a root) and the cycle - walk cannot either (their walks never close). This count is usually what decides - urgency: a two-way deadlock is a bug, a two-way deadlock with forty sessions behind it - is an outage. */ - blocked_behind_count = c.BlockedBehindCount, - blocked_behind_pids = c.BlockedBehindPids, - finding = - "These backends were each waiting on a lock held by another member of the same set — a " - + "genuine cycle, which is a deadlock. PostgreSQL's deadlock detector resolves it after " - + "deadlock_timeout by killing one participant, so this capture landed inside that " - + "window and is likely the only record that will ever exist. The fix is ordering: make " - + "every code path acquire these objects in the same sequence.", - }).ToList(); + var cycleEntries = BuildCycleEntries(cycles); if (chains.Count == 0 && cycleEntries.Count > 0) { @@ -149,6 +133,20 @@ is an outage. */ if (chains.Count == 0) { + /* THREE, once the store knows the engine (#2532). "No captures exist, check the collector" + is the right advice on a PostgreSQL target and the wrong cause on a SQL Server one, where + there is no collector to check. Asked only when there are no captures at all: a window + with captures is a window this collector ran in, so the engine cannot be in question. */ + if (captures.CapturesTotal == 0) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_blocking"); + if (gated != null) + { + return gated; + } + } + /* Two very different empty answers, and conflating them would be the whole failure mode of a sampled signal. No captures at all means the collector never ran — nothing is known about this window either way. Captures with no blocking is a real all-clear, bounded by @@ -173,76 +171,131 @@ the sampling interval. */ }, McpHelpers.JsonOptions); } - var entries = chains.Select(c => new - { - captured_at = c.CapturedAt, - root_pid = c.RootPid, - /* Surfaced because it is what samples_as_root counts, and a reader comparing pids across - captures without it can be fooled by pid reuse. */ - root_backend_id = c.RootBackendId, - databases = c.Databases, - root_username = c.RootUsername, - root_application = c.RootApplicationName, - root_state = c.RootState, - root_is_idle_in_transaction = c.RootIsIdleInTransaction, - root_query = c.RootQuery, - root_xact_duration_ms = c.RootXactDurationMs, - root_query_duration_ms = c.RootQueryDurationMs, - total_victims = c.TotalVictims, - direct_victims = c.DirectVictims, - max_chain_depth = c.MaxDepth, - worst_victim_wait_ms = c.WorstVictimWaitMs, - worst_victim_query = c.WorstVictimQuery, - /* The one-off vs. pattern discriminator, keyed on the stable backend id. NULL when the - root's own identity did not resolve — reported as unknown rather than as 1, because a - fabricated "seen once" reads as a real finding. */ - samples_as_root = c.SamplesAsRoot, - samples_as_root_note = c.SamplesAsRoot is null - ? "Unknown: this root had already left pg_stat_activity when the edge was captured, so " - + "it has no stable backend identity to count appearances of. Not a sign it is new." - : null, - /* Chain-wide, not root-only: the stored flag is an OR across both sides of an edge, so it - answers "some text in this chain may be clipped". */ - query_text_may_be_truncated = c.QueryTextMayBeTruncated, - chain_may_be_truncated = c.ChainMayBeTruncated, - chain_truncation_note = c.ChainMayBeTruncated - ? "This chain hit the read's 32-level walk cap, so total_victims, max_depth and the " - + "worst victim are computed over a truncated walk and are FLOORS, not totals." - : null, - recommended_action = RemedyFor(c.RootState, c.RootIsIdleInTransaction, c.RootXactDurationMs), - }).ToList(); - - var worst = entries[0]; - - return JsonSerializer.Serialize(new - { - server = resolved.ServerName, - hours_back, - status = "blocking_sampled", - captures_total = captures.CapturesTotal, - captures_with_blocking = captures.CapturesWithBlocking, - pct_of_captures_with_blocking = captures.CapturesTotal > 0 - ? Math.Round((double)captures.CapturesWithBlocking / captures.CapturesTotal * 100, 1) - : 0, - /* Lead with the worst chain and what to do about it, the same shape as the xmin tool: - the cause, then the action for that cause. */ - worst_chain_victims = worst.total_victims, - worst_chain_root_state = worst.root_state, - worst_chain_root_application = worst.root_application, - recommended_action = worst.recommended_action, - sampling_caveat = - "These are periodic samples of pg_stat_activity, not an event log. PostgreSQL records " - + "no blocking on its own, so any episode shorter than the collection interval is " - + "invisible here and the counts below are a floor, not a total.", - chains = entries, - /* Always present, even when empty, so its absence is never mistaken for "not checked". */ - cycles_sampled = cycleEntries.Count, - cycles = cycleEntries, - }, McpHelpers.JsonOptions); + return BuildBlockingChainsJson( + resolved.ServerName, hours_back, chains, cycleEntries, captures); } catch (Exception ex) { return McpHelpers.Status("error", $"Reading PostgreSQL blocking chains failed: {ex.Message}"); } } + + /// + /// The cycle rows, projected. Split out with so the WIRE SHAPE can + /// be asserted directly (#2548) without a live store behind it. + /// + internal static List BuildCycleEntries( + IReadOnlyList cycles) + { + return cycles.Select(object (c) => new + { + captured_at = c.CapturedAt, + participant_count = c.ParticipantCount, + pids = c.Pids, + database = c.DatabaseName, + application = c.ApplicationName, + /* Sessions queued behind the deadlock without being part of it. Previously invisible to + BOTH reads — chains cannot see them (no cycle member qualifies as a root) and the cycle + walk cannot either (their walks never close). This count is usually what decides + urgency: a two-way deadlock is a bug, a two-way deadlock with forty sessions behind it + is an outage. */ + blocked_behind_count = c.BlockedBehindCount, + blocked_behind_pids = c.BlockedBehindPids, + finding = + "These backends were each waiting on a lock held by another member of the same set — a " + + "genuine cycle, which is a deadlock. PostgreSQL's deadlock detector resolves it after " + + "deadlock_timeout by killing one participant, so this capture landed inside that " + + "window and is likely the only record that will ever exist. The fix is ordering: make " + + "every code path acquire these objects in the same sequence.", + }).ToList(); + } + + /// + /// The blocking_sampled response body, split out from the tool so the WIRE SHAPE can be asserted + /// directly (#2548) — the tool itself needs a live store and a resolved server, which a serialization + /// guard should not have to stand up, and a guard that re-implemented the projection would keep passing + /// while the shipped one drifted underneath it. + /// + internal static string BuildBlockingChainsJson( + string serverName, + int hoursBack, + IReadOnlyList chains, + List cycleEntries, + DarlingPgBlockingReader.PgBlockingCaptureCounts captures) + { + var entries = chains.Select(c => new + { + captured_at = c.CapturedAt, + root_pid = c.RootPid, + /* Surfaced because it is what samples_as_root counts, and a reader comparing pids across + captures without it can be fooled by pid reuse. + #2548: a STRING, not a number, and this field is the worst case of that rule rather than a + precaution. backend_id is built by CONCATENATING the backend's start epoch with its + zero-padded pid, so it is structurally a 17-digit integer — about 2x past 2^53, in the + range where consecutive doubles are 2 apart. Half of all backend ids are therefore odd and + cannot be represented at all: a JSON-number wire form rounds them onto the NEIGHBOURING + even id, which belongs to a DIFFERENT backend. That is a worse failure than queryid's, + where a rounded key merely joins to nothing — here it joins to the wrong thing, and this + is the field the comment above tells a reader to prefer for exactly that comparison. */ + root_backend_id = c.RootBackendId.ToString(CultureInfo.InvariantCulture), + databases = c.Databases, + root_username = c.RootUsername, + root_application = c.RootApplicationName, + root_state = c.RootState, + root_is_idle_in_transaction = c.RootIsIdleInTransaction, + root_query = c.RootQuery, + root_xact_duration_ms = c.RootXactDurationMs, + root_query_duration_ms = c.RootQueryDurationMs, + total_victims = c.TotalVictims, + direct_victims = c.DirectVictims, + max_chain_depth = c.MaxDepth, + worst_victim_wait_ms = c.WorstVictimWaitMs, + worst_victim_query = c.WorstVictimQuery, + /* The one-off vs. pattern discriminator, keyed on the stable backend id. NULL when the + root's own identity did not resolve — reported as unknown rather than as 1, because a + fabricated "seen once" reads as a real finding. */ + samples_as_root = c.SamplesAsRoot, + samples_as_root_note = c.SamplesAsRoot is null + ? "Unknown: this root had already left pg_stat_activity when the edge was captured, so " + + "it has no stable backend identity to count appearances of. Not a sign it is new." + : null, + /* Chain-wide, not root-only: the stored flag is an OR across both sides of an edge, so it + answers "some text in this chain may be clipped". */ + query_text_may_be_truncated = c.QueryTextMayBeTruncated, + chain_may_be_truncated = c.ChainMayBeTruncated, + chain_truncation_note = c.ChainMayBeTruncated + ? "This chain hit the read's 32-level walk cap, so total_victims, max_depth and the " + + "worst victim are computed over a truncated walk and are FLOORS, not totals." + : null, + recommended_action = RemedyFor(c.RootState, c.RootIsIdleInTransaction, c.RootXactDurationMs), + }).ToList(); + + var worst = entries[0]; + + return JsonSerializer.Serialize(new + { + server = serverName, + hours_back = hoursBack, + status = "blocking_sampled", + captures_total = captures.CapturesTotal, + captures_with_blocking = captures.CapturesWithBlocking, + pct_of_captures_with_blocking = captures.CapturesTotal > 0 + ? Math.Round((double)captures.CapturesWithBlocking / captures.CapturesTotal * 100, 1) + : 0, + /* Lead with the worst chain and what to do about it, the same shape as the xmin tool: + the cause, then the action for that cause. */ + worst_chain_victims = worst.total_victims, + worst_chain_root_state = worst.root_state, + worst_chain_root_application = worst.root_application, + recommended_action = worst.recommended_action, + sampling_caveat = + "These are periodic samples of pg_stat_activity, not an event log. PostgreSQL records " + + "no blocking on its own, so any episode shorter than the collection interval is " + + "invisible here and the counts below are a floor, not a total.", + chains = entries, + /* Always present, even when empty, so its absence is never mistaken for "not checked". */ + cycles_sampled = cycleEntries.Count, + cycles = cycleEntries, + }, McpHelpers.JsonOptions); + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDatabaseTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDatabaseTools.cs new file mode 100644 index 000000000..bae79d39b --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDatabaseTools.cs @@ -0,0 +1,336 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for the per-database pg_stat_database counters, paired with the +/// pg_database_stats collector (#2539). Four questions, one read: temp-file spills, buffer-cache hit +/// ratio, deadlocks, and the commit/rollback split. +/// +[McpServerToolType] +public sealed class DarlingMcpPgDatabaseTools +{ + /// The name PostgreSQL's own NULL-datname row means: activity against SHARED relations + /// (pg_database, pg_authid and the rest of the cluster-wide catalog), which belongs to no + /// single database. Labelled rather than filtered, because its block counters are real and dropping it + /// would overstate every other database's hit ratio. + internal const string SharedRelationsLabel = "(shared relations)"; + + /// + /// What a temp-file figure means and what to do about it — the reason this tool exists. + /// The split that matters is FILE COUNT against BYTES, because they point at different fixes: many + /// small files is a work_mem that is slightly too low for a plan that runs constantly, while a + /// handful of enormous ones is usually a plan or an index problem that no amount of work_mem + /// makes acceptable. Reporting only "it spilled" would leave a reader choosing between those at + /// random. + /// + internal static string SpillFinding(long tempFiles, long tempBytes) + { + if (tempFiles <= 0) + { + return "No temp files written in this window - every sort, hash and materialization fit inside " + + "work_mem. This is the all-clear for the spill question."; + } + + var averageBytes = tempBytes / tempFiles; + var scale = averageBytes >= 256L * 1024 * 1024 + ? "A handful of very large spill files. work_mem is unlikely to be the whole answer at this " + + "size: something is sorting or hashing far more data than it should, so look for a missing " + + "index, a bad row estimate, or a plan that materializes an intermediate it does not need." + : averageBytes >= 8L * 1024 * 1024 + ? "Mid-sized spill files. This is the shape a work_mem increase usually does fix - but raise " + + "it per session or per role rather than globally, because work_mem is charged PER SORT OR " + + "HASH NODE per query, and a global bump multiplies across every concurrent connection." + : "Many small spill files. work_mem is a little under what the plan needs, so a modest " + + "increase often removes them entirely - and small files are also the cheapest to leave " + + "alone if the queries are meeting their deadline."; + + return "This database wrote temp files, which means work_mem could not hold a sort, hash or " + + "materialization and PostgreSQL pushed it to disk. That is the single most common reason a " + + "query is slow for a reason invisible in its plan shape. " + scale + + " pg_stat_database is database-scoped and cannot name the QUERY: on an Amazon Aurora target, " + + "get_pg_top_queries carries per-statement temp_blks_written (in 8 kB blocks) and is where the " + + "attribution comes from; on stock PostgreSQL this counter is the only temp-file evidence " + + "available at all."; + } + + /// + /// The cache hit ratio, with the caveat that makes it honest. The conventional target is above 99%, and + /// the number is genuinely useful — but blks_read is "not found in shared_buffers", not "read + /// from disk": the OS page cache, and on Aurora the storage layer's own cache, sit underneath. So a low + /// ratio bounds the problem rather than measuring it, and saying otherwise would send someone buying + /// memory for latency that is not there. + /// + internal static string CacheHitFinding(double? hitPct) + { + if (hitPct is not { } pct) + { + return "No block accesses in this window, so there is no hit ratio to report - not a ratio of " + + "zero."; + } + + var verdict = pct >= 99 + ? "At or above the conventional 99% target: shared_buffers is absorbing effectively all of this " + + "database's block accesses." + : pct >= 95 + ? "Below the conventional 99% target but not alarming on its own. Worth pairing with " + + "get_pg_io_stats, which splits the misses by CONTEXT - a ratio dragged down by bulkread " + + "(sequential scans deliberately using a small ring buffer) will not improve with more " + + "memory, and one dragged down by the normal context might." + : "Well below the conventional 99% target. Before adding memory, check get_pg_io_stats for " + + "which context the misses are in and get_pg_autovacuum_health for table bloat, because " + + "scanning bloated heaps produces exactly this signature and more shared_buffers only " + + "caches the bloat."; + + return $"{pct.ToString("0.##", CultureInfo.InvariantCulture)}% of block accesses were served from " + + "shared_buffers. " + verdict + + " Read this as an upper bound on the problem rather than a disk-I/O measurement: blks_read " + + "means 'not in shared_buffers', and the OS page cache - or, on Aurora, the storage layer's " + + "own cache - may still have served it without touching a disk."; + } + + /// + /// The commit/rollback split. The RATIO is the finding rather than the rollback count: a busy database + /// legitimately rolls back more transactions than a quiet one, so a raw count says nothing without its + /// denominator. + /// + internal static string RollbackFinding(long commits, long rollbacks) + { + var total = commits + rollbacks; + if (total <= 0) + { + return "No transactions completed in this window."; + } + + var pct = (double)rollbacks / total * 100; + + var verdict = pct >= 10 + ? "A rollback storm. At this share something is failing systematically rather than " + + "occasionally - a deploy that broke a constraint, a deadlock victim loop, a client timing out " + + "mid-transaction and abandoning it." + : pct >= 2 + ? "Higher than a healthy application usually runs. Worth finding out whether it is one code " + + "path or a general rise." + : "A normal share for a healthy application."; + + return $"{pct.ToString("0.##", CultureInfo.InvariantCulture)}% of completed transactions rolled " + + $"back ({rollbacks:N0} of {total:N0}). " + verdict + + " PostgreSQL counts an ERROR-aborted transaction as a rollback, so this includes constraint " + + "violations and statement timeouts, not only explicit ROLLBACK statements."; + } + + [McpServerTool(Name = "get_pg_database_stats"), Description("Gets the per-database PostgreSQL counters from pg_stat_database, differenced across the requested window - four separate questions one cheap cluster-wide view answers. (1) TEMP FILE SPILLS: temp_files / temp_bytes are work that did not fit in work_mem and went to disk, which is the most common reason a PostgreSQL query is slow for a reason its plan shape does not show; on stock PostgreSQL this is the only temp-file evidence available anywhere, and on Aurora get_pg_top_queries carries the per-statement attribution. (2) CACHE HIT RATIO: blks_hit versus blks_read, conventionally targeted above 99% - reported as an upper bound rather than a disk measurement, because a miss here may still be served by the OS page cache or by Aurora's storage layer. (3) DEADLOCKS: a server-recorded count, so a zero really is an all-clear for the window rather than 'none was sampled'. (4) COMMIT vs ROLLBACK: the ratio a rollback storm shows up in. Every figure is a windowed difference clamped per interval, and a statistics RESET is reported explicitly (stats_reset_count, counter_rewind_count) rather than being allowed to surface as a negative rate or a spike. Per-database rows, plus PostgreSQL's own shared-relations row. Works on every PostgreSQL major and on a standby, where sorts spill exactly the way they do on a writer.")] + public static async Task GetPgDatabaseStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum databases to return, most temp bytes first. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + var rows = await DarlingPgDatabaseReader.GetPgDatabaseStatsAsync( + postgres, resolved.ServerId, start, end, limit); + + if (rows.Count == 0) + { + return await EmptyAsync(postgres, resolved.ServerId, resolved.ServerName, hours_back, start, end); + } + + var totalTempFiles = rows.Sum(r => r.TempFiles); + var totalTempBytes = rows.Sum(r => r.TempBytes); + var totalDeadlocks = rows.Sum(r => r.Deadlocks); + var totalHits = rows.Sum(r => r.BlksHit); + var totalReads = rows.Sum(r => r.BlksRead); + var totalAccesses = totalHits + totalReads; + var resetsSeen = rows.Sum(r => r.StatsResetCount) + rows.Sum(r => r.CounterRewindCount); + + var databases = rows.Select(r => + { + var accesses = r.BlksHit + r.BlksRead; + double? hitPct = accesses > 0 ? Math.Round((double)r.BlksHit / accesses * 100, 2) : null; + var transactions = r.XactCommit + r.XactRollback; + + return new + { + /* PostgreSQL's NULL name is a real value, not missing data, so it is LABELLED rather + than passed through as null - a null here reads as "the read could not tell", which + is the one thing it does not mean. */ + database = r.DatabaseName ?? SharedRelationsLabel, + is_shared_relations = r.DatabaseName is null, + temp_files = r.TempFiles, + temp_bytes = r.TempBytes, + avg_temp_file_bytes = r.TempFiles > 0 ? r.TempBytes / r.TempFiles : (long?)null, + spill_finding = SpillFinding(r.TempFiles, r.TempBytes), + blks_hit = r.BlksHit, + blks_read = r.BlksRead, + cache_hit_pct = hitPct, + cache_finding = CacheHitFinding(hitPct), + deadlocks = r.Deadlocks, + xact_commit = r.XactCommit, + xact_rollback = r.XactRollback, + transactions, + rollback_pct = transactions > 0 + ? Math.Round((double)r.XactRollback / transactions * 100, 2) + : (double?)null, + rollback_finding = RollbackFinding(r.XactCommit, r.XactRollback), + /* The reset evidence, per database, because stats_reset is per database. Both counts + travel even when zero: their absence from the payload would be indistinguishable from + a reader that never looked. */ + stats_reset = r.StatsReset, + stats_reset_count = r.StatsResetCount, + counter_rewind_count = r.CounterRewindCount, + counters_were_reset = r.StatsResetCount > 0 || r.CounterRewindCount > 0, + reset_note = r.StatsResetCount > 0 || r.CounterRewindCount > 0 + ? "This database's statistics were RESET during the window (pg_stat_reset, or a " + + "crash restart discarding them). The interval spanning the reset contributes zero " + + "rather than a negative figure, so every total above is a LOWER BOUND on the real " + + "activity - work done between the reset and the next collection is not counted." + : null, + sample_count = r.SampleCount, + first_sample_at = r.FirstSampleAt?.ToString("o"), + last_sample_at = r.LastSampleAt?.ToString("o"), + }; + }) + .ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "database_activity", + database_count = databases.Count, + total_temp_files = totalTempFiles, + total_temp_bytes = totalTempBytes, + total_deadlocks = totalDeadlocks, + /* OF_RETURNED, not cluster_. Every total here is summed over the rows the read's LIMIT let + through, so on a cluster with more active databases than `limit` this is a ratio over the + top-N and NOT the instance's. The name says which, because a `cluster_` prefix would be a + number that silently changes when a caller raises the limit - a field claiming more + comprehensiveness than it structurally has. Computing the true instance ratio would cost a + second unfiltered aggregate on every call for a figure the per-database rows already + support, so the honest name is the fix rather than the extra query. */ + cache_hit_pct_of_returned = totalAccesses > 0 + ? Math.Round((double)totalHits / totalAccesses * 100, 2) + : (double?)null, + /* And the discriminator that makes the caveat actionable rather than fine print: when the + limit bit, the caller knows the totals are a top-N and can raise it. */ + limit_reached = databases.Count >= limit, + /* Named at the top so a reader sees it before drawing a conclusion from any total below. */ + statistics_were_reset_in_window = resetsSeen > 0, + top_spiller = totalTempBytes > 0 ? databases[0].database : null, + note = (resetsSeen > 0 + ? "All figures are windowed differences, clamped per interval so a statistics reset " + + "cannot produce a negative rate or a spike. At least one database's statistics WERE " + + "reset in this window - see stats_reset_count / counter_rewind_count per database - " + + "so its totals are lower bounds rather than exact counts." + : "All figures are windowed differences, clamped per interval so a statistics reset " + + "cannot produce a negative rate or a spike. No reset was detected in this window, so " + + "the totals are complete for the samples collected.") + + (databases.Count >= limit + ? $" The row limit of {limit} was REACHED, so every total above covers only the " + + "databases returned - more may have had activity. Raise limit for the full picture." + : string.Empty), + databases, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_database_stats", ex); + } + } + + /// + /// Which KIND of nothing an empty result is. The engine question is asked first (#2532) — "this database + /// was quiet" is a statement about a PostgreSQL instance, and said about a SQL Server target it is not a + /// weak answer but a false one. + /// + /// The denominator is the DATA, on the same relation the read walks: pg_stat_database is a + /// periodic surface, so any stored sample proves somebody looked. Three misses, not two, because a + /// cumulative counter needs a SECOND sample before it can be differenced at all — reporting a + /// single-sample window as a quiet one would be a confident wrong answer for exactly as long as it takes + /// the next cycle to land. + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, int serverId, string serverName, int hoursBack, DateTime start, DateTime end) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, serverId, serverName, "pg_database_stats"); + if (gated != null) + { + return gated; + } + + var (samplesInWindow, everCollected) = await DarlingPgDatabaseReader.GetCoverageAsync( + postgres, serverId, start, end); + + var hints = new + { + server = serverName, + hours_back = hoursBack, + samples_in_window = samplesInWindow >= 2 ? "2+" : samplesInWindow.ToString(CultureInfo.InvariantCulture), + ever_collected = everCollected, + }; + + if (samplesInWindow >= 2) + { + return McpHelpers.Status( + "empty", + $"No database recorded transactions, block accesses, temp files or deadlocks for {serverName} " + + $"in the last {hoursBack} hour(s), and no statistics reset either. Collection DID run over " + + "this window - at least two snapshots exist to difference - so this is a genuine all-clear " + + "rather than missing data.", + hints); + } + + if (samplesInWindow == 1) + { + return McpHelpers.Status( + "unavailable", + $"Only ONE pg_stat_database snapshot exists for {serverName} in the last {hoursBack} hour(s), " + + "so this is NOT a report of a quiet server. These are cumulative counters and a windowed " + + "difference needs two snapshots before it produces anything at all. On a newly added " + + "server this clears itself on the next collection cycle; otherwise widen hours_back.", + hints); + } + + return McpHelpers.Status( + "unavailable", + everCollected + ? $"No pg_stat_database snapshots were collected for {serverName} in the last {hoursBack} " + + "hour(s), so this is NOT an all-clear - the window says nothing either way. Collection HAS " + + "run for this server outside the window, so this is a gap rather than a dead collector: " + + "widen hours_back, or use get_collection_health to find where it stopped." + : $"No pg_stat_database snapshots have EVER been collected for {serverName}, so there is " + + "nothing to read and this is NOT a report of a quiet server. Check that collection is " + + "running for this server and that it is enabled.", + hints); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDeadlockTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDeadlockTools.cs new file mode 100644 index 000000000..caab765f8 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgDeadlockTools.cs @@ -0,0 +1,154 @@ +// Copyright (c) Erik Darling Data. All rights reserved. +// Licensed under the terms in the LICENSE file in the repository root. + +using System; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// PostgreSQL deadlocks (#2661) — the reports themselves, read out of the server log, rather than the count +/// pg_stat_database keeps. +/// +[McpServerToolType] +public sealed class DarlingMcpPgDeadlockTools +{ + [McpServerTool(Name = "get_pg_deadlocks"), Description("Gets PostgreSQL deadlocks that were reported in the window, newest first, with the victim process, how many sessions were in the cycle, the lock modes and resources involved, and the victim's full statement text. PostgreSQL writes a complete deadlock report to its server log unconditionally - there is no setting that suppresses it - so this needs nothing configured on the target, unlike plan capture. Each row is one DISTINCT deadlock: the collector re-reads an overlapping tail of the log every cycle on purpose, so the same report is seen several times, and times_seen reports that rather than hiding it. A deadlock that genuinely recurred appears as a separate row, because the participating process IDs differ. Use get_pg_deadlock_detail with a deadlock_hash for the full wait graph and every participant's SQL.")] + public static async Task GetPgDeadlocks( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum deadlocks to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + var limitError = McpHelpers.ValidateTop(limit); + if (limitError != null) return McpHelpers.Status("error", limitError); + + try + { + var rows = await DarlingPgDeadlockReader.GetDeadlocksAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_deadlocks") + ?? McpHelpers.Status( + "no_deadlocks", + $"No deadlock was reported on {resolved.ServerName} in the last {hours_back} " + + "hour(s). Two different things produce that and they are worth telling apart: the " + + "server had no deadlocks, which is the healthy answer; or the log could not be " + + "read, which get_pg_plan_capture_readiness reports on because plan capture reads " + + "the same file the same way. pg_stat_database's cumulative deadlock counter, in " + + "get_pg_database_stats, is the independent check - if it moved and nothing is " + + "here, the log is the problem rather than the server."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "deadlocks", + deadlock_count = rows.Count, + truncated = rows.Count >= limit, + note = "occurred_at is when PostgreSQL wrote the report, not when the collector found it. " + + "times_seen counts how often the collector saw the SAME report while it stayed inside " + + "the log tail it re-reads - it is a property of the read window and not a repeat " + + "deadlock, which would appear as its own row with different process IDs.", + deadlocks = rows.Select(r => new + { + occurred_at = r.OccurredAtUtc, + deadlock_hash = r.DeadlockHash, + /* The session PostgreSQL cancelled. It is the end whose application saw an error, which + is usually the only end anybody noticed. */ + victim_pid = r.VictimPid, + participant_count = r.ParticipantCount, + lock_modes = r.LockModes, + resources = r.Resources, + victim_statement = r.VictimStatement, + times_seen = r.TimesSeen, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL deadlocks failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_deadlock_detail"), Description("Gets PostgreSQL deadlock graphs in full: the complete wait graph as the server wrote it, naming every participant, the lock each was waiting for, who blocked whom, and each participant's entire statement text. Pass a deadlock_hash from get_pg_deadlocks for one specific report, or omit it to get the most recent graphs. This carries MORE than a SQL Server deadlock graph does - PostgreSQL names the SQL of every session in the cycle, where the SQL Server graph often leaves the non-victim side as a handle. The graph is stored verbatim rather than reassembled, so a lock type the parser does not break out separately is still readable here.")] + public static async Task GetPgDeadlockDetail( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("A deadlock_hash from get_pg_deadlocks. Omit for the most recent graphs.")] string? deadlock_hash = null, + [Description("Maximum graphs to return when no hash is given. Default 5.")] int limit = 5) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var limitError = McpHelpers.ValidateTop(limit); + if (limitError != null) return McpHelpers.Status("error", limitError); + + try + { + var rows = await DarlingPgDeadlockReader.GetDeadlockDetailAsync( + postgres, resolved.ServerId, deadlock_hash, limit); + + if (rows.Count == 0) + { + return string.IsNullOrWhiteSpace(deadlock_hash) + ? await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_deadlocks") + ?? McpHelpers.Status( + "empty", + $"No deadlock graph is stored for {resolved.ServerName}. Either the server had no " + + "deadlocks, which is the healthy answer, or its log could not be read - " + + "get_pg_database_stats carries pg_stat_database's cumulative deadlock counter, " + + "which tells those apart.") + : McpHelpers.Status( + "empty", + $"No deadlock with hash '{deadlock_hash}' is stored for {resolved.ServerName}. A " + + "hash identifies one report on ONE server, so one from a different server will " + + "not resolve here - and a report can age out of retention while a hash you are " + + "holding does not."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + status = "deadlock_detail", + graph_count = rows.Count, + note = "graph is PostgreSQL's own DETAIL block, verbatim apart from stripped tab indenting. " + + "It reads as: one line per wait edge naming who waits for what and who blocks them, " + + "then each participant's process ID followed by its full statement.", + deadlocks = rows.Select(r => new + { + deadlock_hash = r.DeadlockHash, + occurred_at = r.OccurredAtUtc, + victim_pid = r.VictimPid, + participant_count = r.ParticipantCount, + lock_modes = r.LockModes, + resources = r.Resources, + graph = r.GraphText, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading the PostgreSQL deadlock failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexTools.cs new file mode 100644 index 000000000..d3313347c --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexTools.cs @@ -0,0 +1,198 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for MEASURED index bloat and for column distribution statistics — the two PostgreSQL +/// reads that answer "why is this plan shaped like that", paired with the pg_index_bloat and +/// pg_column_stats collectors (#2629). +/// +/// +/// get_pg_index_bloat is the measured counterpart of get_pg_table_bloat, and the difference +/// is the whole point: table bloat is ESTIMATED from statistics and is suppressed when those statistics +/// cannot be trusted, while this reads pgstatindex, which walks the index. That costs real I/O, so +/// the collector measures only the largest indexes per cycle and LABELS the rest rather than dropping +/// them — a row carrying skipped_reason is a real index that was not measured, not a healthy one. +/// +/// +[McpServerToolType] +public sealed class DarlingMcpPgIndexTools +{ + [McpServerTool(Name = "get_pg_index_bloat"), Description("Gets MEASURED PostgreSQL index bloat from the pgstattuple extension: average leaf density, leaf fragmentation, empty and deleted pages, and how many bytes a REINDEX could plausibly reclaim. This is measured by walking the index, not estimated - contrast get_pg_table_bloat, which estimates from statistics. Low avg_leaf_density is the bloat signal: a freshly built btree is around 90%, and an index that has churned heavily falls well below that. Because measuring costs real I/O the collector measures only the largest indexes each cycle and LABELS the others, so a row with a skipped_reason is an index that was NOT measured rather than one that is healthy - never read a missing measurement as a clean bill of health.")] + public static async Task GetPgIndexBloat( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (7 days) - this collector runs daily.")] int hours_back = 168, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgIndexBloatReader.GetPgIndexBloatAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_index_bloat") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_index_bloat") + ?? McpHelpers.Status( + "empty", + $"No index bloat measurements for {resolved.ServerName} in the last {hours_back} " + + "hour(s). This collector runs DAILY, so a window shorter than a day can be empty " + + "on a perfectly healthy server — widen it before concluding anything."); + } + + var truncated = rows.Count >= limit; + var measured = rows.Count(r => r.SkippedReason is null); + + var indexes = rows.Select(r => new + { + database_name = r.DatabaseName, + schema_name = r.SchemaName, + table_name = r.TableName, + index_name = r.IndexName, + index_bytes = r.IndexBytes, + index_mb = Math.Round(r.IndexBytes / 1024.0 / 1024.0, 1), + /* Null on a skipped row, and deliberately not zero: zero density would read as a + catastrophically bloated index, which is the opposite of "we did not look". */ + avg_leaf_density = r.AvgLeafDensity, + leaf_fragmentation = r.LeafFragmentation, + tree_level = r.TreeLevel, + empty_pages = r.EmptyPages, + deleted_pages = r.DeletedPages, + estimated_reclaimable_bytes = r.EstimatedReclaimableBytes, + estimated_reclaimable_mb = r.EstimatedReclaimableBytes is { } bytes + ? Math.Round(bytes / 1024.0 / 1024.0, 1) + : (double?)null, + /* Present means NOT MEASURED. Named rather than boolean so the row says WHY. */ + skipped_reason = r.SkippedReason, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + index_count = rows.Count, + truncated, + /* Over the returned rows, and withheld when they are only a page of them: "12 of 25 + measured" reads as a statement about the server's indexes and would not be one. */ + measured_count = truncated ? (int?)null : measured, + note = "avg_leaf_density is the bloat signal — a freshly built btree sits near 90%. Rows " + + "carrying a skipped_reason were NOT measured (measuring walks the index, so the " + + "collector bounds how many it does per cycle); a null measurement on those rows is " + + "absence of data, never a clean result." + + (measured < rows.Count + ? $" {rows.Count - measured} of the {rows.Count} row(s) RETURNED are labelled rather than measured." + : string.Empty) + + (truncated + ? " TRUNCATED at the row limit: there are more indexes than this. Raise the limit " + + "before concluding anything about the server as a whole." + : string.Empty), + indexes, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL index bloat failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_column_stats"), Description("Gets PostgreSQL per-column distribution statistics from pg_stats: n_distinct, null fraction, average width, physical correlation, and the frequency of the single most common value. These are the numbers the PLANNER uses, so they explain plan shapes that otherwise look arbitrary. n_distinct is negative when PostgreSQL expresses it as a RATIO of table rows (-1 means every value is unique) and positive when it is an absolute count - do not compare the two without checking the sign. correlation near 1 or -1 means the column's physical order matches its logical order, which is what makes an index range scan cheap; near 0 makes the same scan expensive. A high top_value_frequency is the classic cause of a plan that is right for the common value and wrong for every other one. Only columns on tables above a size floor are collected, and only where the monitoring login can see the statistics.")] + public static async Task GetPgColumnStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (7 days) - this collector runs daily.")] int hours_back = 168, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgColumnStatsReader.GetPgColumnStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_column_stats") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_column_stats") + ?? McpHelpers.Status( + "empty", + $"No column statistics for {resolved.ServerName} in the last {hours_back} hour(s). " + + "This collector runs DAILY and only reads tables above a size floor, so a small " + + "database can be legitimately empty here. pg_stats is also filtered by " + + "privilege: a monitoring login without SELECT on a table sees no rows for it, " + + "and that looks identical to a table with no statistics."); + } + + var columns = rows.Select(r => new + { + database_name = r.DatabaseName, + schema_name = r.SchemaName, + table_name = r.TableName, + column_name = r.ColumnName, + /* Passed through with its sign intact. Normalising it to an absolute count would need the + row count at the time the sample was taken, which is not stored and would be a guess. */ + n_distinct = r.NDistinct, + null_frac = r.NullFrac, + avg_width = r.AvgWidth, + correlation = r.Correlation, + top_value_frequency = r.TopValueFrequency, + /* Null means NO most-common-value list at all, which is itself informative: a perfectly + uniform column has none. Zero would claim the list exists and is empty. */ + common_value_count = r.CommonValueCount, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + column_count = rows.Count, + note = "n_distinct is a RATIO of table rows when negative and an absolute count when " + + "positive — check the sign before comparing two columns. correlation near ±1 is " + + "what makes an index range scan cheap. A null common_value_count means the column " + + "has no most-common-value list at all, not that the list is empty.", + columns, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL column stats failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexUsageTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexUsageTools.cs new file mode 100644 index 000000000..58e18570a --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIndexUsageTools.cs @@ -0,0 +1,420 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for per-index usage, paired with the pg_index_usage_stats collector (#2541). +/// The tool's job is the half the collector cannot do: turning "nothing scanned this" into advice that +/// will not break a schema. Every remedy it emits is gated on the catalog facts stored beside the counters, +/// and the default answer is "investigate", not "drop". +/// +[McpServerToolType] +public sealed class DarlingMcpPgIndexUsageTools +{ + /// + /// Below this many samples the read has not watched the index long enough to say anything about + /// disuse. Two is the floor rather than a comfortable number: it is the minimum at which + /// scans_in_window is a difference rather than a single reading, and the tool says how many + /// samples it actually had so a caller can judge whether two was enough for their question. + /// + internal const int MinimumSamplesForADisuseClaim = 2; + + /// + /// Why an index that nothing scanned may still not be droppable — and this is the whole reason the tool + /// exists rather than the collector being read directly. + /// + /// The blockers are checked before the finding, not appended after it. The single most + /// common zero-scan index on any schema is a unique index backing a PRIMARY KEY or UNIQUE constraint: + /// enforcing uniqueness is not a scan of the kind idx_scan counts, so a constraint index can + /// enforce correctness on every INSERT for years and still report zero. Advice derived from the counter + /// alone would tell somebody to drop their primary key. + /// + internal static string DroppabilityFinding(DarlingPgIndexUsageReader.PgIndexUsageRow r, long scansInWindow, int sampleCount) + { + /* Invalid first: it is the one state where "unused" and "safe to remove" genuinely coincide, and it + is a finding regardless of the scan counters or how long we have been watching. */ + if (!r.IsValid) + { + return "INVALID index - this is what a failed CREATE INDEX CONCURRENTLY leaves behind. The " + + "planner will never use it, so it reports zero scans forever, while INSERTs and UPDATEs " + + "still maintain it and every VACUUM still cleans it. This is the one case where the scan " + + "count needs no corroboration: it is dead weight by construction. Rebuild it with " + + "REINDEX INDEX CONCURRENTLY, or drop it - but confirm first that no application is " + + "waiting on the index it was meant to become."; + } + + if (r.IsPrimaryKey) + { + return "Backs the PRIMARY KEY. Zero scans does NOT mean unused: enforcing the key on every " + + "INSERT and UPDATE is not a scan and is not counted here. This cannot be dropped without " + + "dropping the constraint."; + } + + if (r.SupportsConstraint || r.IsUnique) + { + return "Backs a UNIQUE or EXCLUSION constraint. Zero scans does NOT mean unused - a uniqueness " + + "check is not counted as a scan - and dropping the index drops the constraint with it, " + + "which changes what data the table will accept. Treat the scan count as irrelevant here " + + "unless you have already decided the constraint itself is not wanted."; + } + + if (r.IsReplicaIdentity) + { + return "This index is the table's REPLICA IDENTITY. Dropping it breaks logical replication of " + + "UPDATE and DELETE for this table - downstream subscribers would start erroring or " + + "silently missing rows. Not droppable while logical replication is in use."; + } + + if (sampleCount < MinimumSamplesForADisuseClaim) + { + return "Not enough history to say anything about disuse: this read has " + + sampleCount.ToString(CultureInfo.InvariantCulture) + " sample(s) of this index, and a " + + "windowed scan count is a DIFFERENCE that needs at least two. PostgreSQL records no " + + "index creation time anywhere, so how long we have been watching is the only evidence " + + "available that an index is old enough to judge."; + } + + if (scansInWindow > 0) + { + return "In use - scanned during this window. Nothing to do."; + } + + /* Zero scans in the window, no structural blocker, enough history to say so. Even here the remedy + is qualified, because two things the data cannot see remain. */ + var qualifier = r.IsPartial + ? " This is a PARTIAL index, which reads as unused in two very different situations: nothing " + + "needs it, or the planner cannot MATCH its predicate to how the queries are now written. " + + "Those want opposite fixes, and only the index definition and the query text together can " + + "tell them apart - check the definition below before deciding." + : r.IsExpression + ? " This is an EXPRESSION index, which the planner only uses when a query repeats the " + + "expression exactly as indexed. Zero scans may mean the expression drifted rather than " + + "that the index is unwanted - compare the definition below against the queries." + : string.Empty; + + return "No scans recorded in this window and no constraint, key or replica-identity role that would " + + "stop it being dropped." + qualifier + + " Two things this data still cannot see: a query that runs less often than the window is " + + "long - a monthly or quarterly report will look identical to a dead index over 90 days - and " + + "a rare query whose plan is only acceptable because this index exists. Widen hours_back to " + + "the longest interval any scheduled job runs on before acting, and prefer making the index " + + "invisible to the planner first where you can, so the change is reversible without a rebuild."; + } + + /// + /// What the index costs, in the units the cost is actually paid in. Size is the headline because + /// storage, write amplification and vacuum cleanup all scale with it — but blocks_hit is what + /// makes the cost visible on an index that is never read: those accesses are the write path maintaining + /// it, counted by the server rather than inferred from the size. + /// + internal static string CostFinding(long indexBytes, long tableBytes, long blocksHit, long blocksRead) + { + if (indexBytes < 0) + { + return "Index size was not measured for this sample."; + } + + var share = tableBytes > 0 + ? $" That is {Math.Round((double)indexBytes / tableBytes * 100, 1).ToString(CultureInfo.InvariantCulture)}% of its table's heap." + : string.Empty; + + var maintenance = blocksHit + blocksRead > 0 + ? $" The server recorded {(blocksHit + blocksRead).ToString(CultureInfo.InvariantCulture)} block " + + "accesses against it. On an index with no scans those are the WRITE path maintaining it and " + + "VACUUM cleaning it, which is the cost being paid stated in the server's own units." + : string.Empty; + + return $"{DescribeBytes(indexBytes)}." + share + maintenance + + " On PostgreSQL an unused index costs more than the space: it is maintained by every write " + + "AND cleaned by every VACUUM, and index cleanup is a large share of vacuum work - so a dead " + + "index on a hot table slows the maintenance that keeps the table healthy."; + } + + /// + /// The row's severity band, computed HERE rather than in the browser — the web renderer never + /// re-derives a band, so a grid can only colour a row if the read hands it one. This is what gives the + /// web dashboard the cue the WPF grid gets from its row-style triggers. + /// + /// The ranking mirrors the desktop's exactly: INVALID is Critical (the planner will never + /// use it while writes still maintain it — the one state where "unused" and "safe to remove" coincide), + /// an unscanned index with no structural blocker and enough history is Warning (a candidate, + /// never a conclusion), and everything else is Healthy. An index we have not watched long enough + /// is Healthy rather than Warning: too-early-to-say must not look like a finding. + /// + internal static string IndexSeverity(DarlingPgIndexUsageReader.PgIndexUsageRow r) + { + if (!r.IsValid) + { + return "Critical"; + } + + var blocked = r.IsPrimaryKey || r.SupportsConstraint || r.IsUnique || r.IsReplicaIdentity; + return !blocked && r.ScansInWindow == 0 && r.SampleCount >= MinimumSamplesForADisuseClaim + ? "Warning" + : "Healthy"; + } + + /// + /// Bytes rendered for PROSE only. The payload carries the raw _bytes plus a numeric _mb, + /// per the house pattern (get_pg_autovacuum_health's total_bytes/total_gb); this + /// exists so a finding sentence can say "6.6 MB" rather than a nine-digit number nobody reads. + /// + internal static string DescribeBytes(long bytes) + { + if (bytes < 0) + { + return "an unmeasured size"; + } + + if (bytes >= 1024L * 1024 * 1024) + { + return Math.Round(bytes / 1024.0 / 1024 / 1024, 2).ToString("0.##", CultureInfo.InvariantCulture) + " GB"; + } + + if (bytes >= 1024L * 1024) + { + return Math.Round(bytes / 1024.0 / 1024, 1).ToString("0.#", CultureInfo.InvariantCulture) + " MB"; + } + + return Math.Round(bytes / 1024.0, 0).ToString("0", CultureInfo.InvariantCulture) + " KB"; + } + + [McpServerTool(Name = "get_pg_index_usage")] + [Description( + "PostgreSQL per-index usage: how often each index was scanned, what it costs, and - the part a raw " + + "pg_stat_user_indexes query cannot give you - whether it is actually safe to drop. Reports BOTH the " + + "server's lifetime scan count since its statistics were last reset AND the scans observed across " + + "the stored window, which is the more useful figure: an index with millions of lifetime scans and " + + "none in ninety days is dead weight today, and only collected history can show that. Every index " + + "carries the catalog facts that decide droppability - primary key, unique or exclusion constraint, " + + "replica identity, partial, expression, invalid - and the tool refuses to recommend dropping " + + "anything those facts do not support, including an index it has not watched for at least two " + + "samples. Ranked by the bytes of an index nothing scanned, largest first, with invalid indexes on " + + "top. Collected hourly per database on writers only - a replica reports its own scan counts, not " + + "the writer's, so calling an index unused from a replica would be confidently wrong.")] + public static async Task GetPgIndexUsage( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (seven days). Widen this to the longest interval any scheduled job runs on before calling an index unused.")] int hours_back = 168, + [Description("Maximum indexes to return, biggest unscanned first. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + var rows = await DarlingPgIndexUsageReader.GetPgIndexUsageAsync( + postgres, resolved.ServerId, start, end, limit); + + if (rows.Count == 0) + { + return await EmptyAsync(postgres, resolved.ServerId, resolved.ServerName, hours_back, start, end); + } + + var unscanned = rows.Where(r => r.ScansInWindow == 0 && r.IsValid && !r.IsPrimaryKey + && !r.SupportsConstraint && !r.IsUnique && !r.IsReplicaIdentity + && r.SampleCount >= MinimumSamplesForADisuseClaim).ToList(); + var reclaimable = unscanned.Sum(r => Math.Max(r.IndexBytes, 0)); + var invalidCount = rows.Count(r => !r.IsValid); + var resetSeen = rows.Any(r => r.StatsWereResetInWindow); + + var indexes = rows.Select(r => new + { + database = r.DatabaseName, + schema = r.SchemaName, + table = r.TableName, + index = r.IndexName, + /* Server-computed, because the browser never re-derives a band. Drives the web grid's cell + colour and matches what the WPF row-style triggers paint. */ + severity = IndexSeverity(r), + index_bytes = r.IndexBytes, + index_mb = r.IndexBytes >= 0 ? Math.Round(r.IndexBytes / 1024.0 / 1024, 2) : (double?)null, + table_bytes = r.TableBytes, + /* Both scan figures travel, always. The lifetime count is what anyone querying the server + directly would see, so omitting it would make this read look like it disagreed with + psql; the windowed one is the answer to the question actually being asked. */ + total_scans_since_stats_reset = r.TotalScans, + scans_in_window = r.ScansInWindow, + tuples_read = r.TuplesRead, + tuples_fetched = r.TuplesFetched, + blocks_read = r.BlocksRead, + blocks_hit = r.BlocksHit, + /* NULL on PostgreSQL 15 and below, where the server does not record it at all, and on 16+ + for an index never scanned since the reset. The note says which rather than letting a + null be read as either. */ + last_scan = r.LastScan?.ToString("o"), + last_scan_note = r.LastScan is null + ? "No last-scan timestamp. On PostgreSQL 16+ this means the index has not been scanned " + + "since the statistics were last reset; on 15 and below the server does not record " + + "this at all, so its absence says nothing." + : null, + /* The denominator for the lifetime counter. NULL means never reset, which is the ordinary + state and means the counters run back to the start of the statistics system - not that + the value is unknown. */ + stats_reset = r.StatsReset?.ToString("o"), + stats_reset_note = r.StatsReset is null + ? "This database's statistics have never been reset, so the lifetime scan count covers " + + "the whole life of the statistics system." + : "Lifetime scan counts cover only the period since this reset.", + stats_were_reset_in_window = r.StatsWereResetInWindow, + is_unique = r.IsUnique, + is_primary_key = r.IsPrimaryKey, + is_valid = r.IsValid, + is_ready = r.IsReady, + is_replica_identity = r.IsReplicaIdentity, + is_partial = r.IsPartial, + is_expression = r.IsExpression, + supports_constraint = r.SupportsConstraint, + index_method = r.IndexMethod, + column_count = r.ColumnCount, + /* The definition travels because it is what turns a name into a decision - a human can see + a partial index's predicate, or an expression index's expression, without opening psql + against production. */ + index_definition = r.IndexDefinition, + droppability_finding = DroppabilityFinding(r, r.ScansInWindow, r.SampleCount), + cost_finding = CostFinding(r.IndexBytes, r.TableBytes, r.BlocksHit, r.BlocksRead), + sample_count = r.SampleCount, + first_seen_at = r.FirstSeenAt.ToString("o"), + measured_at = r.MeasuredAt.ToString("o"), + }) + .ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "index_usage", + index_count = indexes.Count, + /* Deliberately NOT called "droppable". These are the indexes with no scans in the window + and no structural blocker - which is a shortlist to investigate, not a work queue. The + field name has to survive being read by something that will act on it. */ + unscanned_without_a_structural_blocker = unscanned.Count, + bytes_held_by_those_indexes = reclaimable, + mb_held_by_those_indexes = Math.Round(reclaimable / 1024.0 / 1024, 2), + invalid_index_count = invalidCount, + statistics_were_reset_in_window = resetSeen, + /* Same discipline as get_pg_database_stats: every total above covers the rows the LIMIT let + through, and the caller has to be able to tell when the cut bit. */ + limit_reached = indexes.Count >= limit, + note = "An index with no scans is a CANDIDATE, never a conclusion. Three things this data " + + "cannot see, in the order they bite: an index backing a constraint enforces it " + + "without ever registering a scan; a query that runs less often than " + + $"{hours_back} hour(s) looks identical to a dead index; and a rare query may only be " + + "acceptable because this index exists. Counts are cumulative since each database's " + + "statistics were last reset, which is reported per index." + + (resetSeen + ? " At least one index's database had its statistics RESET inside this window, so " + + "its windowed scan count is clamped at zero rather than negative and is a lower " + + "bound." + : string.Empty) + + (invalidCount > 0 + ? $" {invalidCount} INVALID index(es) are present and sorted to the top - those are " + + "failed CREATE INDEX CONCURRENTLY leftovers, maintained by writes and used by " + + "nobody." + : string.Empty) + + (indexes.Count >= limit + ? $" The row limit of {limit} was REACHED, so the totals cover only the indexes " + + "returned. Raise limit for the full picture." + : string.Empty), + indexes, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_index_usage", ex); + } + } + + /// + /// Which KIND of nothing an empty result is. The engine question is asked FIRST (#2532): "no unused + /// indexes" is a statement about a PostgreSQL instance, and said about a SQL Server target it is not a + /// weak answer but a false one — get_index_usage is the tool there. + /// The denominator is the DATA, on the same relation the read walks, because + /// pg_index_usage_stats is a PERIODIC surface: a row exists for every index every cycle, so any + /// stored sample proves somebody looked. One snapshot is called out separately from none, because a + /// windowed scan count is a difference and needs two. + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, int serverId, string serverName, int hoursBack, DateTime start, DateTime end) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, serverId, serverName, "pg_index_usage_stats"); + if (gated != null) + { + return gated; + } + + var probe = await DarlingPgIndexUsageReader.ProbePgIndexUsageAsync(postgres, serverId, start, end); + + var hints = new + { + server = serverName, + hours_back = hoursBack, + rows_in_window = probe.RowsInWindow, + snapshots_in_window = probe.SnapshotsInWindow, + ever_collected = probe.RowsEver > 0, + }; + + if (probe.SnapshotsInWindow >= MinimumSamplesForADisuseClaim) + { + return McpHelpers.Status( + "empty", + $"No indexes were recorded for {serverName} in the last {hoursBack} hour(s) even though " + + "collection ran - at least two snapshots exist. Every index on every collected database is " + + "under the collector's 64 KB size floor, which is the state of a schema whose indexes are " + + "all trivially small. That is a genuine all-clear rather than missing data.", + hints); + } + + if (probe.SnapshotsInWindow == 1) + { + return McpHelpers.Status( + "unavailable", + $"Only ONE index snapshot exists for {serverName} in the last {hoursBack} hour(s), so this " + + "is NOT a report about index usage. A windowed scan count is a DIFFERENCE and needs two " + + "snapshots before it produces anything. This collector runs DAILY, so on a newly added " + + "server it clears itself within a day or two; widen hours_back in the meantime.", + hints); + } + + return McpHelpers.Status( + "unavailable", + probe.RowsEver > 0 + ? $"No index snapshots were collected for {serverName} in the last {hoursBack} hour(s), so " + + "the window says nothing either way. Collection HAS run for this server outside the " + + "window - this collector runs DAILY, so a window shorter than about 48 hours can " + + "legitimately contain fewer than two samples. Widen hours_back, or use " + + "get_collection_health to find where it stopped." + : $"No index snapshots have EVER been collected for {serverName}, so there is nothing to " + + "read. Check that collection is running and enabled for this server - and note this " + + "collector is gated OFF on read replicas, because a replica reports its own scan counts " + + "rather than the writer's, which would make every index on it look unused.", + hints); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIoTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIoTools.cs index ff9e13f60..66e8dd3d2 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIoTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgIoTools.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -26,48 +27,46 @@ public sealed class DarlingMcpPgIoTools /// /// Explains what a context value means, because it is the dimension with no SQL Server /// counterpart and the one that changes what you do about a number. + /// #2530 moved the prose to , beside the query that + /// produces the value it explains, so the MCP surface and the WPF viewer's I/O tab print one copy of it + /// rather than two that drift. Kept as a delegating member because this file's tests name it. /// - internal static string ContextMeaning(string? context) => context switch - { - "normal" => "Ordinary buffer-pool traffic. Reads here are cache misses that shared_buffers could " - + "have absorbed, so a high read share with a low hit share is the classic case for more " - + "memory or a better index.", - "bulkread" => "A sequential scan deliberately using a small ring buffer so it cannot evict the " - + "buffer pool. High volume here is a scan-heavy workload, NOT memory pressure — adding " - + "shared_buffers will not reduce it, because these reads bypass the pool by design.", - "bulkwrite" => "A bulk write (COPY, CREATE TABLE AS, some ALTER TABLE) using its own ring buffer.", - "vacuum" => "Vacuum's ring buffer. Volume here is autovacuum doing its job; pair it with " - + "get_pg_autovacuum_health to see whether it is keeping up.", - "index" => "Index-specific I/O, reported separately from the relation's own.", - "walreplay" => "A standby applying WAL. This is replica catch-up work, not query I/O, and it is the " - + "first thing to check when a reader lags.", - _ => "Unrecognized context — treat the raw counters as authoritative and check the PostgreSQL " - + "documentation for this server's major version.", - }; + internal static string ContextMeaning(string? context) => DarlingPgIoReader.ContextMeaning(context); [McpServerTool(Name = "get_pg_io_stats"), Description("Gets PostgreSQL I/O attributed to WHO did it, to WHAT, and WHY - the (backend_type, object, context) breakdown from pg_stat_io, differenced across the requested window. Richer than SQL Server's file-level dm_io_virtual_file_stats: instead of 'this file is busy' you get 'autovacuum workers are reading relations in the vacuum context', which names the cause. The context dimension is the one with no SQL Server equivalent and the one that changes the remedy - it separates ordinary buffer-pool misses (where more shared_buffers or a better index helps) from sequential scans that deliberately bypass the pool via a ring buffer (where it will not help at all), from vacuum's ring buffer, from a standby applying WAL. Reports whether write counters are TRACKED at all, because on Amazon Aurora they are always null - backends there do not write data files, the storage layer does - and a zero would otherwise read as 'no writes happened'. Requires PostgreSQL 16 or later; valid on a standby.")] public static async Task GetPgIoStats( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, - [Description("Maximum (backend_type, object, context) combinations to return, busiest first. Default 20.")] int limit = 20) + [Description("Maximum (backend_type, object, context) combinations to return, busiest first. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgIoReader.GetPgIoAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, limit); if (rows.Count == 0) { + /* Ask the engine BEFORE offering the idle-server reading (#2532). "No combination recorded + activity" is a statement about a PostgreSQL instance; said about a SQL Server target it + is not a weak answer but a false one, and it is the one an agent asking by name gets. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_io_stats"); + if (gated != null) + { + return gated; + } + return JsonSerializer.Serialize(new { server = resolved.ServerName, @@ -113,14 +112,34 @@ they have completely different remedies. */ writes = r.WriteCountersTracked ? r.Writes : (long?)null, write_time_ms = r.WriteCountersTracked ? Math.Round(r.WriteTimeMs, 1) : (double?)null, write_counters_tracked = r.WriteCountersTracked, + /* The block size an operation moves. Gone from 18, where a read is no longer one + block, so it is null there and read_bytes below is measured instead of derived. */ block_bytes = r.OpBytes > 0 ? r.OpBytes : (long?)null, - read_bytes = r.OpBytes > 0 ? r.Reads * r.OpBytes : (long?)null, + /* One name for the volume answer, and bytes_source says how it was arrived at. From 18 + these are measured totals; below 18 they are reads x block size. Never both, and + never silently swapped: the two are different quantities, and on 18 the old estimate + would UNDERCOUNT because a vectored read covers several blocks. */ + read_bytes = r.ByteCountersTracked + ? r.ReadBytes + : (r.OpBytes > 0 ? r.Reads * r.OpBytes : (decimal?)null), + write_bytes = r.ByteCountersTracked + ? r.WriteBytes + : (r.OpBytes > 0 && r.WriteCountersTracked ? r.Writes * r.OpBytes : (decimal?)null), + extend_bytes = r.ByteCountersTracked ? r.ExtendBytes : (decimal?)null, + bytes_source = r.ByteCountersTracked + ? "measured" + : (r.OpBytes > 0 ? "estimated_from_block_size" : "unavailable"), stats_reset = r.StatsReset, }; }) .ToList(); var anyWritesTracked = rows.Any(r => r.WriteCountersTracked); + /* #2655: PostgreSQL 18 replaced op_bytes with measured byte totals. Said once at the top for + the same reason the write flag is: a caller has to know which quantity it is reading before + it compares two servers, and the two are not comparable. */ + var bytesMeasured = rows.Any(r => r.ByteCountersTracked); + var bytesEstimated = !bytesMeasured && rows.Any(r => r.OpBytes > 0); return JsonSerializer.Serialize(new { @@ -135,12 +154,27 @@ they have completely different remedies. */ and a caller needs to know the write side is unmeasured before it concludes anything from the absence of writes. */ write_counters_tracked_anywhere = anyWritesTracked, + bytes_source = bytesMeasured + ? "measured" + : (bytesEstimated ? "estimated_from_block_size" : "unavailable"), note = anyWritesTracked ? "All counters are windowed differences, clamped per interval so a stats reset cannot " + "produce a negative figure." : "All counters are windowed differences. This server tracks NO write counters — the " + "signature of Amazon Aurora, where backends do not write data files and the storage " + "layer does. Absent writes here mean unmeasured, not zero.", + bytes_note = bytesMeasured + ? "Byte totals are MEASURED, from PostgreSQL 18's read_bytes/write_bytes/extend_bytes. " + + "They are not comparable with the figures a pre-18 server reports, which are " + + "reads x block size - 18 reads several blocks per operation, so the older estimate " + + "undercounts." + : (bytesEstimated + ? "Byte totals are ESTIMATED as count x block_bytes, which is exact below " + + "PostgreSQL 18 because one operation moves one block. PostgreSQL 18 measures " + + "them directly instead." + : "This server reports no byte figures at all: op_bytes is absent and the measured " + + "columns PostgreSQL 18 replaced it with are not being collected. The counts and " + + "times above are unaffected."), combinations, }, McpHelpers.JsonOptions); } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgKernelStatsTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgKernelStatsTools.cs new file mode 100644 index 000000000..82a634e22 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgKernelStatsTools.cs @@ -0,0 +1,130 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for per-query OS resource usage, paired with the pg_stat_kcache-backed +/// pg_kernel_stats collector. +/// +/// +/// The other half of get_pg_top_queries. That tool reports ELAPSED time, which is CPU plus every +/// wait; this one reports the CPU inside it, split user and system, from the operating system's own +/// accounting rather than PostgreSQL's. A statement whose elapsed time is mostly CPU and one whose +/// elapsed time is mostly waiting are different problems with opposite fixes, and elapsed time alone +/// cannot separate them — which is why this pairs with get_pg_wait_sampling as naturally as it +/// does with the top-queries read. +/// +/// +/// +/// The byte counters mean something narrower than they look. They are bytes that reached the +/// DEVICE, so zero means the page cache served the read, not that nothing was read — the one reading of +/// this data that turns a healthy server into a mystery. The description says so, because a caller who +/// takes them as logical I/O will conclude a hot, entirely-cached workload does no reads at all. +/// +/// +[McpServerToolType] +public sealed class DarlingMcpPgKernelStatsTools +{ + [McpServerTool(Name = "get_pg_kernel_stats"), Description("Gets per-query-shape OPERATING SYSTEM resource usage from the pg_stat_kcache extension: CPU split into user and system time, bytes that reached the storage device, and major page faults. Use this together with get_pg_top_queries, which reports ELAPSED time: elapsed is CPU plus waiting, so comparing the two tells you whether a slow statement is burning CPU or waiting on something, which is the first split any tuning question needs. The byte counters are DEVICE reads and writes, so a zero means the operating system page cache served the request rather than that no data was read - do not read them as logical I/O. queryid joins get_pg_top_queries and get_pg_wait_sampling. Rows are ranked by total CPU.")] + public static async Task GetPgKernelStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgKernelStatsReader.GetPgKernelStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + /* pg_stat_kcache needs shared_preload_libraries and a restart, so "not installed" is the + likely answer and the precondition vocabulary names the fix. Ordered after the + capability check so a wrong-engine target is never told to install an extension. */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_kernel_stats") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_kernel_stats") + ?? McpHelpers.Status( + "empty", + $"No per-query OS resource usage for {resolved.ServerName} in the last " + + $"{hours_back} hour(s). These are per-interval deltas, so a single collection " + + "has nothing to difference against and the window fills on the second one."); + } + + var totalCpuMs = rows.Sum(r => r.TotalCpuMs); + var resetInWindow = rows.Any(r => r.CounterReset); + + var queries = rows.Select(r => new + { + /* String, like every other queryid on this surface — a signed 64-bit value would round in + a double-decoding JSON parser and produce an id that joins to nothing. */ + queryid = r.QueryId.ToString(CultureInfo.InvariantCulture), + database_name = r.DatabaseName, + cpu_ms = Math.Round(r.TotalCpuMs, 1), + /* Kept split rather than only summed: system time dominated by kernel work is a different + finding from user time dominated by the planner or by expression evaluation. */ + user_cpu_ms = Math.Round(r.ExecUserTimeMs, 1), + system_cpu_ms = Math.Round(r.ExecSystemTimeMs, 1), + pct_of_total_cpu = totalCpuMs > 0 ? Math.Round(r.TotalCpuMs / totalCpuMs * 100, 1) : 0, + device_read_bytes = r.ExecReadBytes, + device_write_bytes = r.ExecWriteBytes, + /* Bytes AND megabytes, the convention get_pg_index_usage already follows: an agent wants + the exact figure, a grid wants something readable, and deriving one from the other at + the display layer is where rounding disagreements start. */ + device_read_mb = Math.Round(r.ExecReadBytes / 1024.0 / 1024.0, 1), + device_write_mb = Math.Round(r.ExecWriteBytes / 1024.0 / 1024.0, 1), + /* A major fault is a page read from disk to satisfy a memory access — the signal that the + host is short of memory, which no PostgreSQL-side counter reports at all. */ + major_faults = r.MajorFaults, + counter_reset = r.CounterReset, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + total_cpu_ms = Math.Round(totalCpuMs, 1), + note = "CPU is measured by the OPERATING SYSTEM, not by PostgreSQL. device_read_bytes and " + + "device_write_bytes count bytes that reached the device, so zero means the page " + + "cache served it rather than that nothing was read." + + (resetInWindow + ? " At least one series was RESET inside this window, so its figures cover only " + + "the time since the reset." + : string.Empty), + queries, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL kernel stats failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPlanTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPlanTools.cs new file mode 100644 index 000000000..24e92184c --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPlanTools.cs @@ -0,0 +1,271 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// Captured PostgreSQL execution plans (#2567). +/// +/// It returns the plan, not a pointer to one. #2538 is explicit about this and it is the whole +/// point of the tool: an agent consuming the MCP has no viewer to follow a reference into, so a read that +/// answers with an id has not answered. +/// +/// Nothing here needs to redact anything, and nothing here should try. The plan JSON is +/// stripped at collection (#2566) — query text dropped, literals replaced — so there is no un-redacted copy +/// in the store for this read to leak. Re-deriving that logic here would create a second place for it to +/// drift out of agreement with the first. +/// +/// The empty answers are the work. "No plan" has three unrelated causes with three unrelated +/// remedies, and collapsing them into one sentence is how a missing grant reads as a healthy query. They are +/// separated here using facts the store already holds rather than prose that guesses between them, which is +/// what #2557 replaced on the Query Store side. +/// +[McpServerToolType] +public sealed class DarlingMcpPgPlanTools +{ + [McpServerTool(Name = "get_pg_plans"), Description("Gets execution plans captured from PostgreSQL by auto_explain, grouped by plan shape and ranked by total time. Returns the plan JSON ITSELF, not a reference to it. The plan is REDACTED at collection and that is not a limitation to work around: auto_explain emits the statement text and its filter literals verbatim, so the query text is dropped entirely and every literal inside the plan tree is replaced with a placeholder before the plan is ever stored — node types, relation names, costs, row estimates and the tree shape all survive intact, which is what a plan is read for. Join query_id to get_pg_top_queries for the statement's normalized text, its call count and its timing. queryid is returned as a STRING because it is a signed 64-bit value spread over the whole int8 range: most ids exceed what a JSON number survives and a numeric wire form is silently rounded by any parser decoding numbers as IEEE-754 doubles, after which it matches nothing. When there are no plans this tool distinguishes three genuinely different causes rather than reporting one vague absence: capture is not configured on the server (with the specific missing precondition), capture is configured but this window's statements never crossed the duration threshold, or plans existed and aged out of retention. This is PostgreSQL-only and separate from get_plan_xml, which covers SQL Server.")] + public static async Task GetPgPlans( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum plan shapes to return. Default 10.")] int limit = 10, + [Description("Only return plans for this queryid, as a string. Optional.")] string? query_id = null, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + /* Parsed from a string for the same reason it is returned as one: an agent that round-trips the id + through a JSON number has already lost it, and accepting a number here would make that silent. */ + long? wantedQueryId = null; + + if (!string.IsNullOrWhiteSpace(query_id)) + { + if (!long.TryParse(query_id, NumberStyles.Integer, CultureInfo.InvariantCulture, out var parsed)) + { + return McpHelpers.Status( + "error", + $"query_id '{query_id}' is not a 64-bit integer. PostgreSQL queryids are signed int8 " + + "values and must be passed as their exact decimal text — if this one arrived through a " + + "JSON number it has already been rounded and no longer matches anything."); + } + + wantedQueryId = parsed; + } + + try + { + var now = windowEnd; + var start = now.AddHours(-hours_back); + + var rows = await DarlingPgPlanCaptureReader.GetPgPlanCaptureAsync( + postgres, resolved.ServerId, start, now, wantedQueryId is null ? limit : limit * 10); + + if (wantedQueryId is not null) + { + rows = rows.Where(r => r.QueryId == wantedQueryId.Value).Take(limit).ToList(); + } + + if (rows.Count == 0) + { + return await NoPlansStatusAsync(postgres, resolved.ServerId, resolved.ServerName, wantedQueryId); + } + + return BuildPlansJson(resolved.ServerName, hours_back, rows, limit); + } + catch (Exception ex) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_plan_capture"); + if (gated != null) + { + return gated; + } + + return McpHelpers.Status("error", $"Reading PostgreSQL plans failed: {ex.Message}"); + } + } + + /// + /// Separates the three reasons there is no plan, using what the store already knows. + /// + /// The order matters and is not arbitrary. A server that cannot capture at all makes every other + /// explanation irrelevant, so the readiness facets are asked FIRST — and they are the only branch that + /// yields an actionable remedy. Only once capture is known to be possible does "this query never crossed + /// the threshold" become a true statement rather than a guess. + /// + private static async Task NoPlansStatusAsync( + NpgsqlDataSource postgres, int serverId, string serverName, long? wantedQueryId) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, serverId, serverName, "pg_plan_capture"); + if (gated != null) + { + return gated; + } + + var precondition = await DarlingRuntimePrecondition.StatusAsync( + postgres, serverId, serverName, "pg_plan_capture"); + if (precondition != null) + { + return precondition; + } + + /* The readiness collector (#2564) already measured every precondition and stored the remedy beside + it, so this names the specific missing step instead of listing everything that could be wrong. */ + var unmet = await UnsatisfiedFacetsAsync(postgres, serverId); + + if (unmet.Count > 0) + { + return McpHelpers.Status( + "precondition", + "This server cannot capture execution plans yet, so an empty result here says nothing about " + + "the query. Unsatisfied precondition(s), from pg_plan_capture_readiness: " + + string.Join(" | ", unmet) + + " — get_pg_plan_capture_readiness has the full detail and the remedy for each."); + } + + var subject = wantedQueryId is null ? "any statement" : "this statement"; + + return McpHelpers.Status( + "empty", + $"Capture is configured on this server and no plan was captured for {subject} in this window. " + + "Two things produce that and they are different: the statement never ran longer than " + + "auto_explain.log_min_duration, which is the healthy answer and means it is not the query to " + + "look at; or a plan was captured earlier and has aged out, which plan_content_retention_days " + + "governs — widen hours_back to tell those apart, because a plan that exists further back will " + + "reappear and one that never existed will not."); + } + + /// + /// The unsatisfied readiness facets, newest reading per facet. Read directly rather than through the + /// readiness tool so this stays a fact lookup rather than one MCP tool narrating another's prose. + /// + private static async Task> UnsatisfiedFacetsAsync(NpgsqlDataSource postgres, int serverId) + { + var unmet = new List(); + + const string sql = """ + SELECT facet, observed + FROM ( + SELECT DISTINCT ON (facet) facet, is_satisfied, observed, collection_time + FROM pg_plan_capture_readiness + WHERE server_id = $1 + ORDER BY facet, collection_time DESC + ) AS latest + WHERE is_satisfied IS NOT TRUE + ORDER BY facet + """; + + try + { + await using var command = postgres.CreateCommand(sql); + command.Parameters.AddWithValue(serverId); + + await using var reader = await command.ExecuteReaderAsync(); + + while (await reader.ReadAsync()) + { + var facet = reader.IsDBNull(0) ? "(unnamed)" : reader.GetString(0); + var observed = reader.IsDBNull(1) ? "(not reported)" : reader.GetString(1); + unmet.Add($"{facet} = {observed}"); + } + } + catch (PostgresException) + { + /* An older store without the readiness table. Silence is correct: the caller falls through to + the generic answer, which is less specific but not wrong. */ + } + + return unmet; + } + + /// + /// The response body, split out so the WIRE SHAPE can be asserted without a live store (#2548) — the + /// same reason BuildTopQueriesJson is separate. + /// + internal static string BuildPlansJson( + string serverName, + int hoursBack, + IReadOnlyList rows, + int limit) + { + var result = rows.Take(limit).Select(r => new + { + /* #2548: a STRING. queryid is a signed int8 spread over the whole 64-bit range, so most values + are past 2^53 and any parser decoding JSON numbers as doubles rounds one — and queryid is an + equality join key, which rounding loses outright rather than approximates. */ + queryid = r.QueryId.ToString(CultureInfo.InvariantCulture), + plan_hash = r.PlanHash, + top_node_type = r.TopNodeType, + node_count = r.NodeCount, + /* CAPTURES, not executions. The collector reads an overlapping tail of the server log, so one + execution can be seen twice; get_pg_top_queries.calls is the authority on how often a + statement actually ran. Named so the difference is visible on the wire. */ + captures = r.Captures, + total_duration_ms = Math.Round(r.TotalDurationMs, 3), + max_duration_ms = Math.Round(r.MaxDurationMs, 3), + avg_duration_ms = Math.Round(r.AvgDurationMs, 3), + last_seen = r.LastSeen, + /* The plan itself, already redacted at collection. Emitted as parsed JSON rather than as a + string so a consumer can walk the tree without a second parse. */ + plan = ParsePlan(r.PlanJson), + }).ToList(); + + return JsonSerializer.Serialize(new + { + server = serverName, + hours_back = hoursBack, + plan_shapes = result.Count, + note = "Plans are REDACTED at collection: statement text is dropped and literals inside the " + + "plan are replaced with placeholders. Node types, relation names, costs, row estimates " + + "and tree shape are intact. Join queryid to get_pg_top_queries for the statement text " + + "and its call counts.", + plans = result, + }, McpHelpers.JsonOptions); + } + + /// + /// Emits the stored plan as JSON when it parses and as a raw string when it does not, rather than + /// dropping it. A plan that survived collection but will not re-parse is worth showing to whoever has to + /// explain why — and an exception here would take out the whole response for one bad row. + /// + private static object? ParsePlan(string? planJson) + { + if (string.IsNullOrWhiteSpace(planJson)) + { + return null; + } + + try + { + return JsonSerializer.Deserialize(planJson); + } + catch (JsonException) + { + return planJson; + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPredicateTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPredicateTools.cs new file mode 100644 index 000000000..2e747579d --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgPredicateTools.cs @@ -0,0 +1,123 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for predicate selectivity, paired with the pg_qualstats-backed +/// pg_predicate_stats collector (#2629). +/// +/// +/// The evidence behind an index recommendation, which is the thing an index recommendation usually lacks. +/// It records which COLUMNS were actually filtered on, with which operators, how many rows each predicate +/// examined and how many it threw away — so "this column should be indexed" stops being a guess about the +/// workload and becomes a count taken from it. +/// +/// +/// +/// It is sampled. pg_qualstats.sample_rate decides what fraction of executions are recorded, +/// and the default is 1%. The row counts are therefore counts of what was SAMPLED; the sample rate travels +/// with every row so a caller can scale them, and deliberately is not applied here — multiplying by 100 +/// and presenting the product as a measurement would launder an estimate into a fact. +/// +/// +[McpServerToolType] +public sealed class DarlingMcpPgPredicateTools +{ + [McpServerTool(Name = "get_pg_predicate_stats"), Description("Gets PostgreSQL predicate selectivity from the pg_qualstats extension: which columns queries actually filter on, with which operator, how many rows each predicate evaluated and how many it filtered out. This is the evidence behind an index recommendation - a predicate that evaluates many rows and filters nearly all of them away is a column doing work an index could do instead. filtered_pct is that ratio. worst_estimate_error_ratio compares what the planner expected against what it got, so a large value marks a predicate the planner is misjudging, which is a statistics or correlated-column problem rather than an indexing one. IMPORTANT: these are SAMPLED counts - pg_qualstats records only a fraction of executions, given per row as sample_rate (commonly 0.01), and the counts are NOT scaled up here. queryid joins get_pg_top_queries.")] + public static async Task GetPgPredicateStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgPredicateStatsReader.GetPgPredicateStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_predicate_stats") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_predicate_stats") + ?? McpHelpers.Status( + "empty", + $"No predicate statistics for {resolved.ServerName} in the last {hours_back} " + + "hour(s). pg_qualstats SAMPLES executions — at the default 1% rate a low-traffic " + + "database can genuinely record nothing — and it needs shared_preload_libraries " + + "plus a restart to be active at all."); + } + + /* One rate for the server in the ordinary case; distinct() rather than First() because a rate + changed mid-window would otherwise be reported as whichever row sorted first. */ + var rates = rows.Select(r => r.SampleRate).Distinct().OrderBy(r => r).ToArray(); + + var predicates = rows.Select(r => new + { + database_name = r.DatabaseName, + schema_name = r.SchemaName, + table_name = r.TableName, + column_name = r.ColumnName, + @operator = r.Operator, + queryid = r.QueryId.ToString(CultureInfo.InvariantCulture), + sampled_executions = r.SampleCount, + rows_evaluated = r.RowsEvaluated, + rows_filtered = r.RowsFiltered, + /* Null when nothing was evaluated — a percentage of zero rows is not zero percent. */ + filtered_pct = r.FilteredPct, + worst_estimate_error_ratio = Math.Round(r.WorstEstimateErrorRatio, 2), + /* Per row, because it can change under the operator's hand mid-window and the counts + beside it are only interpretable against the rate that produced them. */ + sample_rate = r.SampleRate, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + predicate_count = rows.Count, + sample_rates = rates, + note = "Counts are SAMPLED and are not scaled up here: multiply by 1/sample_rate to " + + "estimate the true volume, and treat the product as an estimate. A predicate with " + + "a high filtered_pct over many rows is an indexing candidate; a high " + + "worst_estimate_error_ratio is a planner-estimate problem instead, which an index " + + "will not fix." + + (rates.Length > 1 + ? " The sample rate CHANGED inside this window, so rows are not directly comparable." + : string.Empty), + predicates, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL predicate stats failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgReplicationStatsTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgReplicationStatsTools.cs new file mode 100644 index 000000000..b469e014c --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgReplicationStatsTools.cs @@ -0,0 +1,112 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for CONNECTED replica health, paired with the pg_replication_stats collector +/// (#2629). +/// +/// +/// The other half of get_pg_replication_slots, and the two answer opposite questions. A slot is +/// what the primary RETAINS on a replica's behalf and exists whether or not anybody is attached — a slot +/// with no connection is exactly the shape of the disk-filling incident that motivates monitoring it at +/// all. This reads pg_stat_replication, which is the live connections: who is attached right now, +/// how far behind, and by how much time. +/// +/// +/// +/// The worst-in-window figures are the ones to read. A sample of replication lag catches whatever +/// the replica happened to be doing at that instant, and lag is spiky by nature — the peak is the fact +/// that matters, so it travels beside the latest value rather than being averaged into invisibility. +/// +/// +[McpServerToolType] +public sealed class DarlingMcpPgReplicationStatsTools +{ + [McpServerTool(Name = "get_pg_replication_stats"), Description("Gets the health of CONNECTED PostgreSQL replicas from pg_stat_replication: which replicas are attached, their state and sync state, how many bytes behind they are on send and on replay, and replay lag in milliseconds - with the WORST value seen in the window beside the latest, because lag is spiky and the peak is the fact that matters. This is the counterpart of get_pg_replication_slots and answers a different question: a slot describes what the primary RETAINS for a replica and exists even when nothing is attached, which is the disk-filling case; this describes replicas that are actually connected. A replica that disappears from this tool while its slot persists is the dangerous combination. sync_state distinguishes synchronous replicas, where lag is also commit latency on the primary, from asynchronous ones where it is not. Rows are SAMPLED, so a replica that connected and left between captures may not appear.")] + public static async Task GetPgReplicationStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgReplicationStatsReader.GetPgReplicationStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_replication_stats") + ?? McpHelpers.Status( + "empty", + $"No replica was connected to {resolved.ServerName} in the last {hours_back} " + + "hour(s). On a server with no replicas that is the expected answer. If a " + + "replica is SUPPOSED to be attached, check get_pg_replication_slots — a slot " + + "that persists with nothing connected to it retains WAL indefinitely, which is " + + "the case worth acting on."); + } + + var replicas = rows.Select(r => new + { + application_name = r.ApplicationName, + client_addr = r.ClientAddr, + state = r.State, + /* Synchronous lag is also commit latency on the primary; asynchronous lag is not. The + same number means two different things depending on this column. */ + sync_state = r.SyncState, + sent_bytes_behind = r.SentBytesBehind, + replay_bytes_behind = r.ReplayBytesBehind, + worst_replay_bytes_behind = r.WorstReplayBytesBehind, + replay_lag_ms = r.ReplayLagMs, + worst_replay_lag_ms = r.WorstReplayLagMs, + samples = r.Samples, + total_samples = r.TotalSamples, + backend_start = r.BackendStart, + last_seen = r.LastSeen, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + replica_count = rows.Count, + note = "Read the worst_* columns, not just the latest: lag is spiky and a sample catches " + + "one instant. `samples` against `total_samples` says how much of the window each " + + "replica was actually connected for — a replica present in only a few captures " + + "reconnected repeatedly, which the latest lag figure alone would hide.", + replicas, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL replication stats failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgServerStateTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgServerStateTools.cs new file mode 100644 index 000000000..22e38c239 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgServerStateTools.cs @@ -0,0 +1,507 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for four server-state reads that had a Viewer panel and nothing else (#2629): what is +/// resident in the buffer pool, which extensions the server has, what locking was sampled, and the +/// checkpoint / WAL write picture. +/// +[McpServerToolType] +public sealed class DarlingMcpPgServerStateTools +{ + /// + /// The major that removed buffers_backend and buffers_backend_fsync from + /// pg_stat_bgwriter (#2653). They moved nowhere in that view - the fact lives in + /// pg_stat_io from 17 on - so writes NULL for both at this + /// major and above, and this read names that rather than letting a structural absence read as a missing + /// measurement. + /// + private const int BuffersBackendRemovedInMajor = 17; + + [McpServerTool(Name = "get_pg_buffer_usage"), Description("Gets what is actually resident in the PostgreSQL shared buffer pool, per relation, from the pg_buffercache extension: how many buffers each table or index occupies, how many of those are dirty, and the average usage count that PostgreSQL's clock-sweep eviction reads. This answers which objects the cache is actually spent on, which is a different question from which objects are read most - a small hot table and a large one scanned once can produce similar read counts and completely different residency. High dirty counts concentrated in one relation point at write pressure. Note that scanning pg_buffercache takes a lock on the buffer mapping, so the collector runs it sparingly.")] + public static async Task GetPgBufferUsage( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgBufferUsageReader.GetPgBufferUsageAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_buffer_usage") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_buffer_usage") + ?? McpHelpers.Status( + "empty", + $"No buffer pool contents recorded for {resolved.ServerName} in the last " + + $"{hours_back} hour(s). This needs the pg_buffercache extension in the database " + + "the collector connects to."); + } + + var newest = rows[0]; + + var relations = rows.Select(r => new + { + database_name = r.DatabaseName, + relation_name = r.RelationName, + relation_kind = r.RelationKind, + buffers = r.Buffers, + buffer_mb = Math.Round(r.Buffers * 8.0 / 1024.0, 1), + dirty_buffers = r.DirtyBuffers, + pct_dirty = Math.Round(r.PctDirty, 1), + pct_of_pool = Math.Round(r.PctOfPool, 1), + /* The eviction signal: PostgreSQL's clock sweep decrements this on each pass and evicts at + zero, so a high average means the relation is being touched faster than it decays. */ + avg_usage_count = r.AvgUsageCount, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + relation_count = rows.Count, + /* 8 kB is PostgreSQL's near-universal block size but it IS a compile-time setting, so the + buffer COUNT is the measurement and the megabyte figure is a convenience beside it. */ + pool_buffers_total = newest.PoolBuffersTotal, + pool_buffers_used = newest.PoolBuffersUsed, + note = "Residency, not read volume — a small hot table and a large one scanned once can " + + "read alike and occupy the pool completely differently. avg_usage_count is what the " + + "clock-sweep eviction reads; buffer sizes assume the standard 8 kB block.", + relations, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL buffer usage failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_extensions"), Description("Gets which PostgreSQL extensions are installed, outdated, merely available, or absent, per database. Use this to answer why another read is empty: pg_stat_statements, pg_wait_sampling, pg_stat_kcache, pg_qualstats, pgstattuple, pg_buffercache and hypopg each back a specific tool, and 'absent' here is the reason that tool has nothing to show. State is one of installed, outdated, available or absent - 'available' means the files are on the server and CREATE EXTENSION would work, which is a different situation from absent and usually a one-line fix. IMPORTANT: extensions are per-DATABASE, so installed_version reflects the database the row names and not the cluster; an extension can be installed in the application database and absent from postgres.")] + public static async Task GetPgExtensions( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (7 days) - this collector runs daily.")] int hours_back = 168, + [Description("Maximum rows to return. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgExtensionAvailabilityReader.GetPgExtensionAvailabilityAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_extension_availability") + ?? McpHelpers.Status( + "empty", + $"No extension inventory for {resolved.ServerName} in the last {hours_back} " + + "hour(s). This collector runs DAILY, so a short window can be empty on a healthy " + + "server — widen it before concluding anything."); + } + + /* The result is capped at `limit`, so a summary count taken over the returned rows describes + the PAGE and reads as a fact about the SERVER. Measured while writing this: at limit=50 the + tool reported "installed: 10" for a server with more than that, because 50 was all it had + looked at. Suppressed rather than renamed — "installed_in_this_page" is a number nobody + wants. */ + var truncated = rows.Count >= limit; + + var extensions = rows.Select(r => new + { + database_name = r.DatabaseName, + extension_name = r.ExtensionName, + state = r.State, + installed_version = r.InstalledVersion, + default_version = r.DefaultVersion, + /* Whether THIS PRODUCT can use it, not merely whether it exists — the distinction that + makes this list actionable rather than an inventory. */ + monitoring_relevant = r.IsMonitoringRelevant, + comment = r.Comment, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + extension_count = rows.Count, + truncated, + installed = truncated + ? (int?)null + : rows.Count(r => string.Equals(r.State, "installed", StringComparison.OrdinalIgnoreCase)), + available_not_installed = truncated + ? (int?)null + : rows.Count(r => string.Equals(r.State, "available", StringComparison.OrdinalIgnoreCase)), + note = "Per DATABASE, not per cluster: installed_version describes the database the row " + + "names. 'available' means the files are present and CREATE EXTENSION would work — " + + "usually a one-line fix for an empty panel elsewhere." + + (truncated + ? " TRUNCATED at the row limit, so the state totals are WITHHELD: counting them " + + "over a capped result would describe this page rather than the server. Raise " + + "the limit for a complete inventory." + : string.Empty), + extensions, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL extension availability failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_lock_stats"), Description("Gets PostgreSQL lock activity as SAMPLED by the collector: lock type, mode, whether it was granted, the relation involved, how many captures saw it and the worst wait observed. This is a SAMPLE, not an event log - the collector periodically photographs pg_locks, so a lock held briefly between two samples is invisible here and an absence is not proof nothing was locked. Ungranted rows are the ones that matter: a granted lock is normal operation, while a lock waiting to be granted is a session blocked behind another. For a blocking chain with the root blocker attributed, use get_pg_blocking instead - this tool answers which lock modes and relations are contended over time rather than who is blocking whom right now.")] + public static async Task GetPgLockStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgLockStatsReader.GetPgLockStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_lock_stats") + ?? McpHelpers.Status( + "empty", + $"No lock activity sampled on {resolved.ServerName} in the last {hours_back} " + + "hour(s). This is a SAMPLE of pg_locks rather than an event log, so this is the " + + "healthy state on a server without sustained contention — and it is not proof " + + "that nothing was ever locked."); + } + + var truncated = rows.Count >= limit; + var ungranted = rows.Count(r => !r.Granted); + + var locks = rows.Select(r => new + { + database_name = r.DatabaseName, + lock_type = r.LockType, + mode = r.Mode, + granted = r.Granted, + relation_name = r.RelationName, + captures = r.Captures, + /* The denominator: how many captures happened at all, so `captures` can be read as a rate + of presence rather than an unanchored count. */ + total_captures = r.TotalCaptures, + max_backends = r.MaxBackends, + max_wait_ms = r.MaxWaitMs, + last_seen = r.LastSeen, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + lock_count = rows.Count, + truncated, + /* Counted over the returned rows, and reported only when they are all of them. */ + ungranted_count = truncated ? (int?)null : ungranted, + note = (truncated + ? "TRUNCATED at the row limit — the ungranted total is WITHHELD, because counting " + + "over a capped result would describe this page rather than the server. " + : string.Empty) + + "A SAMPLE of pg_locks, not an event log: a lock taken and released between two " + + "captures does not appear. Read `captures` against `total_captures` — that ratio is " + + "how much of the window the lock was present for. Ungranted rows are the contended " + + "ones; for who is blocking whom, use get_pg_blocking.", + locks, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL lock stats failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_write_stats"), Description("Gets PostgreSQL checkpoint and WAL write activity across the window: how many checkpoints were timed versus REQUESTED, how long they spent writing and syncing, how many buffers were written by checkpoints, by the background writer and by backends themselves, and the WAL record, full-page-image and byte totals. Requested checkpoints are the signal to look for - a timed checkpoint is the scheduled one, while a requested checkpoint means WAL filled max_wal_size before the interval elapsed, so a high requested share means checkpoints are being forced by write volume. buffers_backend counts writes a query had to do itself because no clean buffer was available, which is backpressure landing on user queries - PostgreSQL 17 removed that column from pg_stat_bgwriter, so it is null on 17 and later and the note says so. wal_fpi counts full-page images, which is why write volume spikes immediately after each checkpoint. Returns one row describing the whole window, not a series.")] + public static async Task GetPgWriteStats( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var row = await DarlingPgWriteStatsReader.GetPgWriteStatsAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd); + + if (row is null) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_write_stats") + ?? McpHelpers.Status( + "empty", + $"No checkpoint or WAL activity recorded for {resolved.ServerName} in the last " + + $"{hours_back} hour(s). These are differenced across snapshots, so a single " + + "collection has nothing to difference against and the window fills on the " + + "second one."); + } + + var timed = row.CheckpointsTimed ?? 0; + var requested = row.CheckpointsRequested ?? 0; + + /* #2653: 17 removed buffers_backend / buffers_backend_fsync from pg_stat_bgwriter with no + successor there, so the collector writes NULL for them and the fact now lives in pg_stat_io. + Without the registry's major this read cannot tell that structural absence from a measurement + that did not happen, and its note explained a column that will never have a value here. */ + var postgresMajor = await DarlingEngineCapability.PostgresMajorVersionAsync( + postgres, resolved.ServerId); + var backendCountersRemoved = postgresMajor >= BuffersBackendRemovedInMajor; + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + window_start = row.WindowStartUtc, + window_end = row.WindowEndUtc, + checkpoints_timed = row.CheckpointsTimed, + checkpoints_requested = row.CheckpointsRequested, + /* The whole reason both numbers are here. Null rather than 0 when there were no + checkpoints at all — 0% requested would claim a healthy result from no evidence. */ + pct_checkpoints_requested = timed + requested > 0 + ? Math.Round((double)requested / (timed + requested) * 100, 1) + : (double?)null, + checkpoint_write_time_ms = row.CheckpointWriteTimeMs, + checkpoint_sync_time_ms = row.CheckpointSyncTimeMs, + buffers_written_checkpoint = row.BuffersWrittenCheckpoint, + buffers_clean = row.BuffersClean, + /* Backpressure landing on user queries: a backend writing its own buffer is a query + paying for the write because no clean buffer was free. */ + buffers_backend = row.BuffersBackend, + buffers_backend_fsync = row.BuffersBackendFsync, + /* Named only when the registry actually carries a major that removed them: absent from the + payload otherwise, so a server whose version is unknown gets no claim rather than a + guessed one. */ + buffers_backend_availability = backendCountersRemoved + ? $"not_collected: removed from pg_stat_bgwriter in PostgreSQL 17 (this server runs {postgresMajor}); the fact now lives in pg_stat_io, read by get_pg_io_stats" + : null, + buffers_alloc = row.BuffersAlloc, + maxwritten_clean = row.MaxwrittenClean, + wal_records = row.WalRecords, + wal_fpi = row.WalFpi, + wal_bytes = row.WalBytes, + wal_buffers_full = row.WalBuffersFull, + wal_write_time_ms = row.WalWriteTimeMs, + wal_sync_time_ms = row.WalSyncTimeMs, + counter_reset = row.ResetDuringWindow, + note = "Requested checkpoints mean max_wal_size filled before the scheduled interval, so a " + + "high requested share means write volume is forcing them. " + + (backendCountersRemoved + ? "buffers_backend and buffers_backend_fsync are NULL because PostgreSQL 17 removed " + + "them from pg_stat_bgwriter, not because nothing was measured: the question they " + + "answered - whether backends are writing their own buffers - is answered on this " + + "server by pg_stat_io, which get_pg_io_stats reads. " + : "buffers_backend is a write a QUERY had to perform itself for want of a clean " + + "buffer. ") + + "wal_fpi counts full-page images, which is why WAL volume spikes just after each " + + "checkpoint." + + (row.ResetDuringWindow + ? " The counters were RESET inside this window, so these figures cover only the " + + "time since the reset." + : string.Empty), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL write stats failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_server_config"), Description("Gets the PostgreSQL server's configuration from pg_settings - what each parameter is set to, whether it differs from the compiled-in default, where the value came from (configuration file, command line, ALTER SYSTEM, per-database or per-role), and whether changing it needs a restart or only a reload. Non-default settings are listed FIRST, because a server has several hundred parameters and only the ones somebody chose are an answer. Reports pending_restart loudly: that means postgresql.conf was edited and reloaded but the running server is still using the old value, so the file and the server disagree with no symptom until the next restart. Session-scoped rows are excluded - pg_settings is a per-connection view and its client-source rows describe the monitoring connection, not the server. Snapshot from the most recent collection, not a window.")] + public static async Task GetPgServerConfig( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Maximum settings to return. Default 100.")] int limit = 100, + [Description("When true, include settings still at their default. Default false - the non-default ones are the answer.")] bool include_defaults = false) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var limitError = McpHelpers.ValidateTop(limit); + if (limitError != null) return McpHelpers.Status("error", limitError); + + try + { + var rows = await DarlingPgServerConfigReader.GetCurrentConfigAsync( + postgres, resolved.ServerId, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_server_config") + ?? McpHelpers.Status( + "empty", + $"No configuration snapshot has been collected for {resolved.ServerName} yet. " + + "This collector runs hourly, so a server registered in the last hour has not " + + "reached its first collection."); + } + + var shown = include_defaults ? rows : rows.Where(r => !r.IsDefault).ToList(); + var pendingRestart = rows.Where(r => r.PendingRestart).Select(r => r.Name).ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + status = "server_config", + /* Both counts, because they answer different questions and one without the other invites + the wrong conclusion: a small returned count is reassuring only if you know it was + filtered rather than truncated. */ + settings_returned = shown.Count, + non_default_count = rows.Count(r => !r.IsDefault), + truncated = rows.Count >= limit, + pending_restart_count = pendingRestart.Count, + pending_restart_settings = pendingRestart.Count > 0 ? pendingRestart : null, + note = pendingRestart.Count > 0 + ? "One or more settings are marked PENDING RESTART: the configuration file has been " + + "changed and reloaded, but the running server is still using the previous value. " + + "The file and the server disagree until the next restart, at which point the " + + "behaviour changes with no deployment to explain it." + : "Non-default settings first. 'source' says where the value came from; 'context' says " + + "what changing it would take - postmaster needs a restart, sighup a reload, user " + + "nothing. Session-scoped rows are excluded: pg_settings is a per-connection view " + + "and those describe the monitoring connection rather than the server.", + settings = shown.Select(r => new + { + name = r.Name, + setting = r.Setting, + unit = r.Unit, + /* The compiled-in default, so a reader can see what was moved away FROM without + needing a table of defaults that would rot at every major. */ + default_value = r.BootValue, + is_default = r.IsDefault, + source = r.Source, + context = r.Context, + requires_restart_to_change = string.Equals(r.Context, "postmaster", StringComparison.Ordinal), + pending_restart = r.PendingRestart, + category = r.Category, + description = r.ShortDescription, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL server config failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_server_config_changes"), Description("Gets PostgreSQL configuration parameters whose value CHANGED during the window, newest first, with the old and new value side by side. This is the read that answers 'this got slow sometime last month, what changed' - and nothing else in the stack can reconstruct it after the fact, because a configuration history that was not recorded cannot be recovered from the server. A setting appearing for the first time is deliberately NOT reported as a change: the first snapshot after an upgrade, or after an extension is loaded, would otherwise manufacture hundreds of changes nobody made. Session-scoped rows are excluded, so a monitoring reconnect does not read as a configuration change.")] + public static async Task GetPgServerConfigChanges( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (one week).")] int hours_back = 168, + [Description("Maximum changes to return. Default 100.")] int limit = 100, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + var limitError = McpHelpers.ValidateTop(limit); + if (limitError != null) return McpHelpers.Status("error", limitError); + + try + { + var rows = await DarlingPgServerConfigReader.GetConfigChangesAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_server_config") + ?? McpHelpers.Status( + "no_changes", + $"No configuration parameter changed value on {resolved.ServerName} in the last " + + $"{hours_back} hour(s). That is a real finding rather than missing data - this " + + "read compares consecutive snapshots, so an unchanged server produces no rows."); + } + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "config_changes", + change_count = rows.Count, + truncated = rows.Count >= limit, + note = "changed_at is the time of the snapshot that FIRST reported the new value, so the " + + "change happened at some point in the hour before it - this collector runs hourly. " + + "A setting appearing for the first time is not reported here.", + changes = rows.Select(r => new + { + changed_at = r.ChangedAtUtc, + name = r.Name, + old_value = r.OldValue, + new_value = r.NewValue, + unit = r.Unit, + source = r.Source, + context = r.Context, + description = r.ShortDescription, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL config changes failed: {ex.Message}"); + } + } + +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSessionStatesTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSessionStatesTools.cs new file mode 100644 index 000000000..45d670a49 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSessionStatesTools.cs @@ -0,0 +1,437 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Globalization; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for session states, paired with the pg_session_states collector (#2540). +/// The tool's job is the half the collector deliberately does not do: deciding what a long +/// idle-in-transaction session MEANS. The collector stores durations and a horizon age; the judgement that +/// separates "this is starving your vacuum" from "this is a connection pool doing nothing harmful" lives +/// here, and it is gated on peak_horizon_age rather than on the duration. +/// +[McpServerToolType] +public sealed class DarlingMcpPgSessionStatesTools +{ + /// + /// A session must have been seen holding the horizon in at least this share of the samples that saw it + /// before the tool calls it a sustained holder rather than a passing one. + /// Half. Every session with an open write transaction is momentarily the oldest holder at some + /// point — that is what a transaction IS — so a single sighting proves nothing beyond "this instance + /// does writes". Sustained holding is the finding, and the tool reports the raw sample counts beside the + /// verdict so a caller can disagree with the threshold rather than having to trust it. + /// + internal const double SustainedHolderSampleShare = 0.5; + + /// + /// The duration past which an idle-in-transaction session is worth a human's attention even when it + /// pins nothing. + /// Five minutes. It holds no snapshot and no transaction id, so it is costing vacuum nothing — but + /// it IS holding a connection and any locks the transaction took, and at five minutes it has stopped + /// looking like application latency and started looking like a code path that forgot to commit. The + /// finding says exactly that rather than borrowing the vacuum argument, which does not apply. + /// + internal const long IdleWithoutHorizonAttentionMs = 5 * 60 * 1000; + + /// + /// What this session was actually doing to the database, and the one place in this feature where the + /// causal claim is made or withheld. + /// + /// peak_horizon_age decides, not the duration. Four idle-in-transaction shapes were + /// measured on a live PostgreSQL 16.15 instance and two of them pin nothing: a READ COMMITTED + /// transaction that only read released its snapshot when the statement ended, and one whose UPDATE + /// matched zero rows never got a transaction id at all. Both sat idle in transaction indefinitely with + /// backend_xmin and backend_xid NULL. A tool that reasoned from the state string and the + /// clock would tell someone to kill those sessions to save their vacuum, and killing them would change + /// nothing about the horizon. + /// + internal static string HorizonFinding(DarlingPgSessionStatesReader.PgSessionStateRow r) + { + /* Redaction first, and it outranks everything: if the state columns were redacted then the inputs + to every other branch below are NULL, and a finding derived from them would be an artefact of a + missing GRANT rather than an observation about the database. */ + if (r.StateWasRedacted) + { + return "CANNOT SAY - the monitoring login lacks pg_monitor on this target, so PostgreSQL " + + "redacted this row. Only pid, application_name, database and user survive - state, " + + "state_change, xact_start, query_start, backend_start, wait_event, backend_type, " + + "client_addr and query_id all came back NULL and the query text was replaced. " + + "backend_xmin and backend_xid are NOT redacted, which is why an xid age can still " + + "appear above: the horizon still reads as pinned and nothing here can say by what. " + + "Grant pg_monitor to the monitoring role and this becomes answerable."; + } + + var sustained = r.SampleCount > 0 + && r.HorizonHolderSamples >= Math.Max(1, (int)Math.Ceiling(r.SampleCount * SustainedHolderSampleShare)); + + if (r.HorizonHolderSamples > 0 && sustained) + { + return "PINS THE XMIN HORIZON, and did so across most of the samples that saw it " + + $"({r.HorizonHolderSamples} of {r.SampleCount}). While this transaction stays open, " + + "VACUUM cannot reclaim any row version newer than it on ANY table in the cluster - which " + + "is why bloat grows on tables this session never touched, and why autovacuum keeps " + + "running and reporting success while reclaiming nothing. Cross-check the same window in " + + "get_pg_xmin_horizon: its 'session' row should name this pid. The fix is to make the " + + "transaction end, which is an application change, not a database one - terminating the " + + "backend releases the horizon immediately but the code path that opened it will do it " + + "again."; + } + + if (r.HorizonHolderSamples > 0) + { + var passing = "Held the oldest xmin in " + + $"{r.HorizonHolderSamples} of {r.SampleCount} sample(s) - a passing sighting " + + "rather than a sustained hold. Every write transaction is briefly the oldest " + + "holder, so the hold ON ITS OWN is normal traffic. It becomes a finding if the " + + "share climbs."; + + /* But the hold is not the only thing this row can be. A session that is ALSO idle in + transaction past the attention floor is a real finding whatever its share of the holds - + and returning only the sentence above for it was a genuine bug: the severity band said + Warning while this text said "not a finding", in the same JSON object. The two now agree, + and the reason they nearly did not is that is_horizon_holder means OLDEST rather than + HOLDS: several long transactions on one instance take turns being oldest, so none of them + reaches the sustained threshold and every one of them still needs saying. */ + return r.IdleInTransactionSamples > 0 + && r.PeakStateDurationMs >= IdleWithoutHorizonAttentionMs + ? passing + " It is a finding on the OTHER axis though: at " + + FormatDuration(r.PeakStateDurationMs) + " idle inside a transaction it is " + + "holding a connection and any locks the transaction took, and it did pin " + + "the horizon while it did so - just not as the oldest holder for most of " + + "the window. Being the oldest is a property of the INSTANCE, not of this " + + "session: on a busy one several long transactions take turns." + : passing; + } + + if (r.PeakHorizonAge < 0 && r.IdleInTransactionSamples > 0) + { + var note = "Idle in transaction, but pinned NOTHING - it held neither a snapshot " + + "(backend_xmin) nor a transaction id (backend_xid) in any sample. This is the " + + "ordinary shape of a READ COMMITTED transaction that has only read, or whose write " + + "matched no rows: the snapshot is released at the end of each statement. It is " + + "costing VACUUM nothing, so killing it will not help bloat or wraparound."; + + return r.PeakStateDurationMs >= IdleWithoutHorizonAttentionMs + ? note + " It is still worth looking at on its own terms: at " + + FormatDuration(r.PeakStateDurationMs) + " idle inside a transaction it is holding a " + + "connection and any locks the transaction already took, which is a forgotten " + + "commit rather than a vacuum problem." + : note; + } + + if (r.PeakHorizonAge < 0) + { + return "Pinned nothing - no snapshot and no transaction id in any sample. Held a transaction " + + "open, but not one that VACUUM had to wait behind."; + } + + return "Held a transaction id or snapshot at a peak age of " + + r.PeakHorizonAge.ToString("N0", CultureInfo.InvariantCulture) + + " transactions, but was never the OLDEST holder on the instance, so something else was " + + "setting the horizon. See get_pg_xmin_horizon for what."; + } + + /// + /// Severity band, server-computed so the browser and the WPF row styles paint the same thing from the + /// same decision rather than each re-deriving it. + /// The four words are the HOUSE vocabulary — Healthy, Warning, Critical, + /// Unknown, capitalised — the same set IndexSeverity and BloatSeverity return, and + /// the only set the shared sev-* CSS defines. A private vocabulary here would be worse than + /// wrong: sev-info and sev-none match no rule, so the badge would render unstyled rather + /// than failing, which is the kind of defect that ships. + /// + internal static string SessionSeverity(DarlingPgSessionStatesReader.PgSessionStateRow r) + { + if (r.StateWasRedacted) + { + /* Its own band, above every severity, for the same reason pg_table_bloat_stats gives + estimate_unavailable one: this row carries no trustworthy state at all, and painting it as + healthy or as critical would both be inventions. */ + return "Unknown"; + } + + var sustained = r.SampleCount > 0 + && r.HorizonHolderSamples >= Math.Max(1, (int)Math.Ceiling(r.SampleCount * SustainedHolderSampleShare)); + + if (r.HorizonHolderSamples > 0 && sustained) + { + return "Critical"; + } + + /* The Warning band is LONG IDLE IN TRANSACTION, and it is deliberately not conditioned on + whether the session pinned anything. Review suggested adding "&& PeakHorizonAge < 0" here to + match the viewer; the agreement was real and the direction was wrong. is_horizon_holder means + OLDEST on the instance rather than HOLDS, so on a busy instance several genuinely long idle + transactions each take turns being oldest, none of them clears the sustained threshold, and + that gate would have painted every one of them Healthy - the dangerous direction. The viewer + and the finding text were changed to agree with THIS instead. */ + if (r.IdleInTransactionSamples > 0 && r.PeakStateDurationMs >= IdleWithoutHorizonAttentionMs) + { + return "Warning"; + } + + /* A passing sighting collapses into Healthy rather than getting a band of its own. Every write + transaction is briefly the oldest xmin holder, so a distinct band there would paint ordinary + traffic and teach people to ignore the column. The sample counts are on the row for anyone who + wants to see the difference. */ + return "Healthy"; + } + + internal static string FormatDuration(long milliseconds) + { + if (milliseconds < 0) + { + return "not measured"; + } + + var span = TimeSpan.FromMilliseconds(milliseconds); + if (span.TotalMinutes < 1) + { + return span.TotalSeconds.ToString("0.#", CultureInfo.InvariantCulture) + "s"; + } + + return span.TotalHours < 1 + ? span.TotalMinutes.ToString("0.#", CultureInfo.InvariantCulture) + "m" + : span.TotalHours.ToString("0.#", CultureInfo.InvariantCulture) + "h"; + } + + [McpServerTool(Name = "get_pg_session_states")] + [Description( + "PostgreSQL sessions holding a transaction open - who is idle in transaction, for how long, and " + + "WHETHER THAT SESSION IS ACTUALLY PINNING THE XMIN HORIZON. That last part is the whole point and " + + "it is not the same question as the first two: measured on a live instance, an idle-in-transaction " + + "session under READ COMMITTED that has only read, or whose UPDATE matched zero rows, holds neither " + + "a snapshot nor a transaction id and starves VACUUM of nothing at all. Only a transaction that has " + + "written (holding backend_xid) or one under REPEATABLE READ (holding backend_xmin) pins anything. " + + "So peak_horizon_age, not the duration, is what supports a causal claim, and a peak_horizon_age of " + + "-1 means the session pinned NOTHING rather than something small. Pairs with get_pg_xmin_horizon, " + + "which names the CLASS of holder; this names the session. Rolled up per backend across the window, " + + "horizon holders first, then longest transaction. THIS IS A SAMPLE at the collection interval, not " + + "an event log - PostgreSQL records nothing about session state unless something asks, so a " + + "transaction that opened and closed between samples is invisible; the capture counts are reported " + + "so 'nothing found' is distinguishable from 'nobody looked'. No raw query text is stored: " + + "pg_stat_activity.query carries literal parameter values, so the normalised query_id and a " + + "whitelisted command keyword are stored instead - join query_id to get_pg_top_queries for the " + + "statement text with placeholders. Requires pg_monitor on the target; without it PostgreSQL " + + "silently returns rows with every state column NULL, which the read reports rather than hides.")] + public static async Task GetPgSessionStates( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum sessions to return, horizon holders first then longest transaction. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + var rows = await DarlingPgSessionStatesReader.GetPgSessionStatesAsync( + postgres, resolved.ServerId, start, end, limit); + + /* Fetched whether or not there are rows. On a surface where zero rows is the healthy answer, the + denominator is not an error path - it is what makes a healthy answer believable. */ + var captures = await DarlingPgSessionStatesReader.GetPgSessionStatesCaptureCountsAsync( + postgres, resolved.ServerId, start, end); + + if (rows.Count == 0) + { + return await EmptyAsync(postgres, resolved.ServerId, resolved.ServerName, hours_back, captures); + } + + var holders = rows.Count(r => r.HorizonHolderSamples > 0 && !r.StateWasRedacted); + var idleHolders = rows.Count(r => r.HorizonHolderSamples > 0 && r.IdleInTransactionSamples > 0 + && !r.StateWasRedacted); + var idleWithoutHorizon = rows.Count(r => r.IdleInTransactionSamples > 0 && r.PeakHorizonAge < 0 + && !r.StateWasRedacted); + var redacted = rows.Count(r => r.StateWasRedacted); + var truncated = rows.Any(r => r.CaptureWasTruncated); + + var sessions = rows.Select(r => new + { + pid = r.Pid, + /* The collector's synthetic (backend_start, pid) identity, and the reason a pid alone is not + the key: pids are reused, and this id is comparable to the one on get_pg_blocking's rows + for the same backend. */ + backend_id = r.BackendId, + severity = SessionSeverity(r), + database = r.DatabaseName, + username = r.Username, + application_name = r.ApplicationName, + client_addr = r.ClientAddr, + backend_type = r.BackendType, + last_state = r.LastState, + last_wait_event_type = r.LastWaitEventType, + last_wait_event = r.LastWaitEvent, + /* The leading SQL keyword, whitelisted at collection. Not a truncation of the statement - + see the collector for why a substring would be a data-leak vector. */ + last_command_tag = r.LastCommandTag, + /* Normalised statement identity. Join it to get_pg_top_queries for the text with $1 + placeholders where the literals were. */ + last_query_id = r.LastQueryId, + query_id_note = r.LastQueryId is null + ? "No query_id. PostgreSQL 13 does not have the column at all; on 14+ it is NULL when " + + "compute_query_id is off, and on a redacted row it is NULL with everything else " + + "privileged. Check the target's version before reading anything into this." + : null, + peak_xact_duration_ms = r.PeakXactDurationMs, + peak_xact_duration = FormatDuration(r.PeakXactDurationMs), + peak_state_duration_ms = r.PeakStateDurationMs, + peak_state_duration = FormatDuration(r.PeakStateDurationMs), + peak_query_duration_ms = r.PeakQueryDurationMs, + /* How long the BACKEND has existed, against how long it has held its transaction. The pair + separates two different bugs that look identical in the transaction duration alone: a + connection created ten minutes ago that has been idle in transaction for all ten is a + pool handing out a session nobody finished with, while a three-day-old worker that has + held one for ten minutes is a code path that forgot to commit. */ + backend_age_ms = r.PeakBackendDurationMs, + backend_age = FormatDuration(r.PeakBackendDurationMs), + /* -1 means pinned NOTHING. Not a small age - no age. */ + peak_horizon_age = r.PeakHorizonAge, + peak_xmin_age = r.PeakXminAge, + peak_xid_age = r.PeakXidAge, + pinned_the_horizon = r.PeakHorizonAge >= 0, + horizon_holder_samples = r.HorizonHolderSamples, + idle_in_transaction_samples = r.IdleInTransactionSamples, + sample_count = r.SampleCount, + first_seen_at = r.FirstSeenAt.ToString("o"), + last_seen_at = r.LastSeenAt.ToString("o"), + state_was_redacted = r.StateWasRedacted, + /* Capture context from this backend's most recent sample: two idle-in-transaction sessions + out of six connections is a different instance from two out of four thousand. */ + sessions_on_instance = r.TotalSessions, + active_sessions_on_instance = r.ActiveSessions, + idle_in_transaction_on_instance = r.IdleInTransactionSessions, + reportable_sessions_on_instance = r.ReportableSessions, + capture_was_truncated = r.CaptureWasTruncated, + finding = HorizonFinding(r), + }) + .ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "session_states", + session_count = sessions.Count, + horizon_holder_count = holders, + idle_in_transaction_holder_count = idleHolders, + /* Reported as its own number because it is the correction this tool exists to make: these + sessions look exactly like the harmful ones and are not. */ + idle_in_transaction_pinning_nothing = idleWithoutHorizon, + /* The denominator, always, on a sampled surface. */ + captures_in_window = captures.CapturesTotal, + captures_with_sessions = captures.CapturesWithSessions, + first_capture_at = captures.FirstCaptureAt?.ToString("o"), + last_capture_at = captures.LastCaptureAt?.ToString("o"), + redacted_row_count = redacted, + limit_reached = sessions.Count >= limit, + note = "This is a SAMPLE taken every collection cycle, not an event log. PostgreSQL records " + + "nothing about session state unless something asks it, so a transaction that opened " + + "and closed between two samples left no trace - " + + $"{captures.CapturesWithSessions} of {captures.CapturesTotal} capture(s) in this " + + "window found any reportable session at all. Duration alone is NOT evidence that a " + + "session is starving VACUUM: read peak_horizon_age, where -1 means the session pinned " + + "nothing." + + (idleWithoutHorizon > 0 + ? $" {idleWithoutHorizon} session(s) here were idle in transaction and pinned " + + "NOTHING - they held no snapshot and no transaction id, so terminating them " + + "would not reclaim a single dead row." + : string.Empty) + + (redacted > 0 + ? $" {redacted} row(s) are REDACTED: the monitoring login lacks pg_monitor on this " + + "target, so PostgreSQL returned the rows with every state column NULL rather " + + "than refusing the read. Those rows say nothing about session state and their " + + "severity is reported as unknown." + : string.Empty) + + (truncated + ? " At least one capture hit the collector's per-capture row cap, so the stored " + + "rows for it are a worst-first sample of a larger set - compare " + + "reportable_sessions_on_instance against the rows returned." + : string.Empty) + + (sessions.Count >= limit + ? $" The row limit of {limit} was REACHED, so the counts above cover only the " + + "sessions returned. Raise limit for the full picture." + : string.Empty), + sessions, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_session_states", ex); + } + } + + /// + /// Which KIND of nothing an empty result is. The engine question is asked FIRST (#2532): "no long + /// transactions" said about a SQL Server target is not a weak answer but a false one. + /// The denominator comes from collection_log and not from the data, because this is an + /// EXCEPTION surface like pg_blocking_edges: the collector stores nothing when every session is + /// behaving, so an absence of rows is the HEALTHY state and probing the table itself would report a + /// perfectly monitored server as uncollected (#2508). + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, int serverId, string serverName, int hoursBack, + DarlingPgSessionStatesReader.PgSessionStatesCaptureCounts captures) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, serverId, serverName, "pg_session_states"); + if (gated != null) + { + return gated; + } + + var hints = new + { + server = serverName, + hours_back = hoursBack, + captures_in_window = captures.CapturesTotal, + captures_with_sessions = captures.CapturesWithSessions, + first_capture_at = captures.FirstCaptureAt?.ToString("o"), + last_capture_at = captures.LastCaptureAt?.ToString("o"), + }; + + if (captures.CapturesTotal > 0) + { + return McpHelpers.Status( + "empty", + $"No session held a transaction open past the collector's floor on {serverName} in the last " + + $"{hoursBack} hour(s), across {captures.CapturesTotal} capture(s). This is the healthy " + + "state and a real all-clear rather than missing data - the collector stores nothing when " + + "every transaction is short. One caveat that is not a hedge: this samples at the " + + "collection interval, so a transaction that opened and closed between two samples is " + + "genuinely invisible here.", + hints); + } + + return McpHelpers.Status( + "unavailable", + $"No captures at all for {serverName} in the last {hoursBack} hour(s), so the window says " + + "nothing either way - this is NOT an all-clear. Check that collection is running and enabled " + + "for this server with get_collection_health, or widen hours_back.", + hints); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSlotTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSlotTools.cs index 25b5b509d..af0eb207a 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSlotTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgSlotTools.cs @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -47,17 +48,18 @@ internal static string Classify(string? walStatus, bool isActive, bool retainedW public static async Task GetPgReplicationSlots( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to analyze, used for the WAL growth comparison. Default 24.")] int hours_back = 24) + [Description("Hours of history to analyze, used for the WAL growth comparison. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgSlotReader.GetPgSlotsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); @@ -68,6 +70,16 @@ public static async Task GetPgReplicationSlots( a cluster-wide all-clear would be drawing the one conclusion this result cannot support. */ if (rows.Count == 0) { + /* The engine question comes first (#2532): "this instance has no replication slots" is a + real finding on a PostgreSQL target and a false one on a SQL Server target, where the + collector has never run and never will. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_replication_slots"); + if (gated != null) + { + return gated; + } + return JsonSerializer.Serialize(new { server = resolved.ServerName, diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgStatementTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgStatementTools.cs index e5367f298..0b84241d7 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgStatementTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgStatementTools.cs @@ -7,13 +7,16 @@ */ using System; +using System.Collections.Generic; using System.ComponentModel; +using System.Globalization; using System.Linq; using System.Text.Json; using System.Threading.Tasks; using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -24,96 +27,172 @@ namespace PerformanceMonitor.Darling.Service.Mcp; [McpServerToolType] public sealed class DarlingMcpPgStatementTools { - [McpServerTool(Name = "get_pg_top_queries"), Description("Gets the top PostgreSQL query shapes by total execution time over a time period, for Amazon Aurora PostgreSQL targets. Includes Aurora's I/O source breakdown, which stock PostgreSQL cannot provide: a block 'read' may have come from the distributed storage volume or from the local NVMe Optimized Reads cache, and the two have very different costs. Also reports peak memory per statement, the closest PostgreSQL equivalent of a memory grant, and WAL bytes generated, which has no SQL Server DMV counterpart. Returns query_text for each statement, captured hourly and keyed on queryid, or null when none has been captured yet (a statement first seen minutes ago, or a queryid minted by a major-version upgrade). queryid itself is stable within a major version but changes across a major upgrade — which is exactly why the text is STORED rather than fetched live: after an upgrade the live view no longer holds the old ids, so their text would otherwise be unrecoverable and the history would read as a list of integers. This is a separate tool from get_top_queries_by_cpu, which covers SQL Server.")] + [McpServerTool(Name = "get_pg_top_queries"), Description("Gets the top PostgreSQL query shapes by total execution time over a time period, for Amazon Aurora PostgreSQL targets. Includes Aurora's I/O source breakdown, which stock PostgreSQL cannot provide: a block 'read' may have come from the distributed storage volume or from the local NVMe Optimized Reads cache, and the two have very different costs. Also reports peak memory per statement, the closest PostgreSQL equivalent of a memory grant, and WAL bytes generated, which has no SQL Server DMV counterpart. Returns query_text for each statement, captured hourly and keyed on queryid, or null when none has been captured yet (a statement first seen minutes ago, or a queryid minted by a major-version upgrade). queryid itself is stable within a major version but changes across a major upgrade — which is exactly why the text is STORED rather than fetched live: after an upgrade the live view no longer holds the old ids, so their text would otherwise be unrecoverable and the history would read as a list of integers. queryid is returned as a STRING, not a number: it is a signed 64-bit value spread over the whole int8 range, so most ids exceed what a JSON number survives and a numeric wire form would be silently rounded by any parser that decodes numbers as IEEE-754 doubles — after which it matches nothing. Compare it and join on it as text. This is a separate tool from get_top_queries_by_cpu, which covers SQL Server.")] public static async Task GetPgTopQueries( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, - [Description("Maximum rows to return. Default 20.")] int limit = 20) + [Description("Maximum rows to return. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgStatementReader.GetPgTopQueriesAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) { - return McpHelpers.Status( - "unavailable", - "No PostgreSQL query statistics for this server and window. If this server is SQL " - + "Server, use get_top_queries_by_cpu instead. If it is Aurora PostgreSQL, check that " - + "pg_stat_statements is installed in the database the collector connects to — on some " - + "clusters it exists only in the application database, not in postgres."); + /* The capability answer settles the dialect branch and the stock-PostgreSQL branch + whenever the store knows the engine (#2532), so this sentence is reached in exactly two + states: an Aurora target that produced no rows, which has one genuinely diagnosable + cause, and a row whose engine_kind is NULL, where no claim can be made. It names those + two rather than repeating the ones that can no longer reach this line. */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_statement_stats") + /* #2546: the sentence below tells the reader to go and CHECK whether pg_stat_statements + is installed in the connected database. The collector has already checked — a 42P01 + against a database where the extension was never created classifies as a non-fatal + skip and stores the CREATE EXTENSION to run. Asking for it here is the whole reason + the precondition vocabulary exists on the PostgreSQL side (#2545): the extension is + mutable at runtime, so nothing decided at connect time could report it and then stop + reporting it when somebody acts on the advice. */ + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_statement_stats") + ?? McpHelpers.Status( + "unavailable", + "No PostgreSQL query statistics for this server and window. Check that " + + "pg_stat_statements is installed in the database the collector connects to — it is " + + "not installed by default, and on some clusters it exists only in the application " + + "database, not in postgres. Also expect an empty window from a SINGLE snapshot: the " + + "counters here are per-interval deltas, so the first collection after a restart has " + + "nothing to difference against and the window fills on the second. Otherwise the " + + "store has not recorded this server's engine yet, and a target it cannot classify may " + + "not be a PostgreSQL one at all — check list_servers."); } - var totalTimeMs = rows.Sum(r => r.TotalExecTimeMs); - - var result = rows.Take(limit).Select(r => - { - var storageAndCache = r.StorageBlocksRead + r.OrcacheBlocksHit; - return new - { - queryid = r.QueryId, - database_id = r.DatabaseId, - calls = r.Calls, - total_exec_time_ms = r.TotalExecTimeMs, - avg_exec_time_ms = r.Calls > 0 ? Math.Round((double)r.TotalExecTimeMs / r.Calls, 3) : 0, - max_exec_time_ms = Math.Round(r.MaxExecTimeMs, 3), - rows_returned = r.RowsReturned, - pct_of_total_time = totalTimeMs > 0 ? Math.Round((double)r.TotalExecTimeMs / totalTimeMs * 100, 1) : 0, - /* Aurora's I/O split, which is the point of using aurora_stat_statements over the - vanilla view. A high orcache share means the reads were cheap local NVMe hits; a - high storage share means network round trips to the cluster volume. The community - cache-hit ratio cannot distinguish these and so overstates the cost of one and - understates the other. */ - shared_blks_hit = r.SharedBlocksHit, - shared_blks_read = r.SharedBlocksRead, - storage_blks_read = r.StorageBlocksRead, - orcache_blks_hit = r.OrcacheBlocksHit, - orcache_hit_pct_of_reads = storageAndCache > 0 - ? Math.Round((double)r.OrcacheBlocksHit / storageAndCache * 100, 1) - : (double?)null, - /* Spills. temp blocks are sort/hash spill to disk, NOT temporary tables - those are - the local_blks_* family and a different problem. */ - temp_blks_read = r.TempBlocksRead, - temp_blks_written = r.TempBlocksWritten, - wal_bytes = r.WalBytes, - max_exec_peakmem_bytes = r.MaxPeakMemBytes, - // #2219: the statement text, or null when none has been captured for this queryid yet. - // Null is honest rather than a placeholder — text refreshes hourly, so a statement first - // seen minutes ago has none, and after a major-version upgrade re-keys queryid the new ids - // have none until the next refresh. An empty string would read as "the query is blank". - query_text = r.QueryText, - }; - }); - - return JsonSerializer.Serialize(new - { - server = resolved.ServerName, - hours_back, - total_exec_time_ms = totalTimeMs, - /* Every counter here covers the window, so a caller can safely divide one by another. - Only the two high-water marks are not counters, and saying which is cheaper than - letting someone assume max_exec_time_ms is a windowed total. */ - note = "All counters cover the requested window: calls, total_exec_time_ms and " - + "rows_returned from stored per-interval deltas, and the block and WAL figures " - + "differenced across the window's snapshots. max_exec_time_ms and " - + "max_exec_peakmem_bytes are high-water marks, not windowed totals.", - queries = result, - }, McpHelpers.JsonOptions); + return BuildTopQueriesJson(resolved.ServerName, hours_back, rows, limit); } catch (Exception ex) { + /* #2554: a THROW is a miss too, and until now it was the one miss the capability answer could + not reach. The gate sat inside `if (rows.Count == 0)`, so it only ever spoke when the query + SUCCEEDED and returned nothing — and a caller on a target that excludes pg_statement_stats got + a raw SQL error where its sibling get_pg_wait_stats gives the honest "does not run on that + engine, and never will". (That caller was on stock PostgreSQL until #2625, which gave the + collector a vanilla pg_stat_statements path; the gate now answers the DIALECT exclusion, and a + stock-PostgreSQL fault correctly reads as a fault.) + + Consulted HERE rather than moved ahead of the read, which is what it looks like it should be. + DarlingEngineCapability's contract is explicit that every call site asks AFTER its read came + back empty, never before: a server whose registry row says one engine while its collected rows + say another — a re-registration, a restored database — must still get its DATA rather than a + confident explanation of why it cannot have any. Asking first would trade this defect for that + one, across every read. On the throw path there is no data to prefer, so the same contract + points the other way. + + And this stays narrow rather than swallowing failures: NotCollectedStatusAsync returns null + unless the collector provably cannot run on this server's engine, so a transient error against + an Aurora target still surfaces as an error. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_statement_stats"); + if (gated != null) + { + return gated; + } + return McpHelpers.Status("error", $"Reading PostgreSQL query stats failed: {ex.Message}"); } } + + /// + /// The response body, split out from the tool so the WIRE SHAPE can be asserted directly (#2548). The + /// tool itself needs a live store and a resolved server, neither of which a serialization guard has any + /// business standing up — and a guard that re-implements the projection instead would keep passing while + /// the shipped one drifted underneath it. + /// + internal static string BuildTopQueriesJson( + string serverName, + int hoursBack, + IReadOnlyList rows, + int limit) + { + var totalTimeMs = rows.Sum(r => r.TotalExecTimeMs); + + var result = rows.Take(limit).Select(r => + { + /* Null when the source did not report the split at all (self-hosted PostgreSQL, #2625), which + is a different answer from zero and must not become a 0% cache-hit ratio. */ + var storageAndCache = r.StorageBlocksRead + r.OrcacheBlocksHit; + return new + { + /* #2548: a STRING, not a number. queryid is a signed int8 whose values are spread over the + whole 64-bit range, so most of them are past 2^53 and every parser that decodes JSON + numbers as IEEE-754 doubles — JSON.parse, json.loads, most agent tooling — silently + rounds one. That is not the same failure as a rounded metric: queryid is the ONLY + identity a PostgreSQL statement has, and every use of it (SELECT … WHERE queryid = …, + matching this row against the operator's own screen, quoting it in a ticket) is an + equality join, which a rounded key loses entirely rather than approximates. A string + costs a type change once; a number costs the value on every read. It also matches how + SQL Server's query_hash already reaches this surface — as text, never as an integer. + database_id stays a number beside it because an oid is unsigned 32-bit and so cannot + reach the rounding range. Invariant culture, because a negative id must render with + ASCII '-' whatever the host's locale would prefer. */ + queryid = r.QueryId.ToString(CultureInfo.InvariantCulture), + database_id = r.DatabaseId, + calls = r.Calls, + total_exec_time_ms = r.TotalExecTimeMs, + avg_exec_time_ms = r.Calls > 0 ? Math.Round((double)r.TotalExecTimeMs / r.Calls, 3) : 0, + max_exec_time_ms = Math.Round(r.MaxExecTimeMs, 3), + rows_returned = r.RowsReturned, + pct_of_total_time = totalTimeMs > 0 ? Math.Round((double)r.TotalExecTimeMs / totalTimeMs * 100, 1) : 0, + /* Aurora's I/O split, which is the point of using aurora_stat_statements over the + vanilla view. A high orcache share means the reads were cheap local NVMe hits; a + high storage share means network round trips to the cluster volume. The community + cache-hit ratio cannot distinguish these and so overstates the cost of one and + understates the other. */ + shared_blks_hit = r.SharedBlocksHit, + shared_blks_read = r.SharedBlocksRead, + storage_blks_read = r.StorageBlocksRead, + orcache_blks_hit = r.OrcacheBlocksHit, + orcache_hit_pct_of_reads = storageAndCache > 0 + ? Math.Round((double)r.OrcacheBlocksHit!.Value / storageAndCache.Value * 100, 1) + : (double?)null, + /* Spills. temp blocks are sort/hash spill to disk, NOT temporary tables - those are + the local_blks_* family and a different problem. */ + temp_blks_read = r.TempBlocksRead, + temp_blks_written = r.TempBlocksWritten, + wal_bytes = r.WalBytes, + max_exec_peakmem_bytes = r.MaxPeakMemBytes, + // #2219: the statement text, or null when none has been captured for this queryid yet. + // Null is honest rather than a placeholder — text refreshes hourly, so a statement first + // seen minutes ago has none, and after a major-version upgrade re-keys queryid the new ids + // have none until the next refresh. An empty string would read as "the query is blank". + query_text = r.QueryText, + }; + }); + + return JsonSerializer.Serialize(new + { + server = serverName, + hours_back = hoursBack, + total_exec_time_ms = totalTimeMs, + /* Every counter here covers the window, so a caller can safely divide one by another. + Only the two high-water marks are not counters, and saying which is cheaper than + letting someone assume max_exec_time_ms is a windowed total. */ + note = "All counters cover the requested window: calls, total_exec_time_ms and " + + "rows_returned from stored per-interval deltas, and the block and WAL figures " + + "differenced across the window's snapshots. max_exec_time_ms and " + + "max_exec_peakmem_bytes are high-water marks, not windowed totals.", + queries = result, + }, McpHelpers.JsonOptions); + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTableBloatTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTableBloatTools.cs new file mode 100644 index 000000000..2156a413d --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTableBloatTools.cs @@ -0,0 +1,444 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for per-table bloat, paired with the pg_table_bloat_stats collector (#2542). +/// +/// Everything this tool does is in service of one rule: never present an estimate as a +/// measurement. The number it reports is arithmetic over PostgreSQL's column-width statistics, not a +/// reading of the table, and the difference is the difference between "consider investigating this" and +/// "run VACUUM FULL on a production table". So the estimate is named _estimate in every field that +/// carries it, is SUPPRESSED rather than captioned when its inputs are missing, is downgraded when the +/// statistics behind it are stale, and always ships beside the exact command that would settle it. +/// +[McpServerToolType] +public sealed class DarlingMcpPgTableBloatTools +{ + /// + /// The threshold, re-exported from the ONE place the rule lives so this tool's tests can pin it without + /// reaching across projects. carries + /// the measurement it was set from. + /// + internal const double StaleStatisticsChurnRatio = DarlingPgTableBloatReader.StaleStatisticsChurnRatio; + + /// + /// The REASON the estimate may not be shown, or null when it is fit to publish. + /// + /// The decision itself is — shared + /// with the WPF grid, so the two surfaces cannot disagree about whether a number is publishable. What + /// lives here is only the prose, because a grid cell has room for three words and this has room to say + /// what to actually do about it. + /// + /// Any non-null return means the caller nulls the estimate fields rather than rendering them with + /// the reason attached — a footnote under a large red percentage is not a safeguard. + /// + internal static string? EstimateSuppressionReason(DarlingPgTableBloatReader.PgTableBloatRow r) + { + if (!DarlingPgTableBloatReader.EstimateIsUnpublishable(r)) + { + return null; + } + + if (r.EstimateUnavailable) + { + return "SUPPRESSED: the estimate has no basis for this table. Either the table has never been " + + "analyzed, or - much more likely in production - the monitoring login cannot SELECT from " + + "it, so pg_stats returned no column widths to compute from. pg_monitor does NOT confer " + + "SELECT on user tables, and in that state the estimator does not fail: it silently " + + "returns a large number. Measured against a pg_monitor-only role on a real target it " + + "reported 88.59% for a table whose true bloat was 0.50%. Grant pg_read_all_data " + + "(PostgreSQL 14+) or SELECT on the schema to the monitoring role, then re-read."; + } + + if (r.LiveTuples < 0) + { + return "SUPPRESSED: this table has never been analyzed, so PostgreSQL does not know how many " + + "rows it holds (reltuples is -1, its explicit 'unknown'). There is no row count to " + + "compare the page count against. Run ANALYZE on the table and re-read."; + } + + { + var pct = r.LiveTuples > 0 + ? Math.Round((double)r.ModsSinceAnalyze / r.LiveTuples * 100, 1) + : 0; + return "SUPPRESSED: the column-width statistics this estimate depends on are STALE - " + + pct.ToString("0.#", CultureInfo.InvariantCulture) + "% of this table's rows have been " + + "modified since it was last analyzed. This is the estimator's worst failure mode and it " + + "is not a small one: on a test pair differing only in this, the stale table's estimate " + + "was 92.64% against a true 10.93%, an error of 81 percentage points, with nothing in the " + + "arithmetic to show it. Run ANALYZE on the table and re-read."; + } + } + + /// + /// What to do about the number, once it is judged fit to show. The remedies are deliberately ordered + /// least-destructive first, and VACUUM FULL is never the lead recommendation — it takes an + /// ACCESS EXCLUSIVE lock for the duration and needs free space equal to the finished table. + /// + internal static string BloatFinding(decimal bloatPct, long bloatBytes, bool pgstattupleAvailable, string qualifiedName) + { + var confirm = pgstattupleAvailable + ? $"pgstattuple is installed on this database, so the exact figure is one query away: SELECT * FROM pgstattuple('{qualifiedName}'). It reads the WHOLE relation, so run it deliberately rather than on a schedule." + : $"pgstattuple is NOT installed on this database, so the exact figure is not reachable from here. CREATE EXTENSION pgstattuple, then SELECT * FROM pgstattuple('{qualifiedName}') - and note it reads the whole relation, which on a large table is expensive enough that it is why this collector estimates instead."; + + if (bloatPct < 20) + { + return "Estimated bloat is low. Nothing to do - and at this level the estimate's own error bar " + + "is a meaningful share of the number, so treat it as 'not a problem' rather than as a " + + "precise figure. " + confirm; + } + + if (bloatPct < 50) + { + return "Estimated bloat is moderate. This is usually autovacuum keeping up but running behind a " + + "steady write rate rather than a problem needing intervention - check " + + "get_pg_autovacuum_health for this table first, because if autovacuum is being BLOCKED " + + "(a long transaction, an idle-in-transaction session, an orphaned replication slot) then " + + "reclaiming the space now only buys time until the same thing happens again. " + + "get_pg_xmin_horizon names which of those it is. " + confirm; + } + + return "Estimated bloat is high - roughly " + + DescribeBytes(bloatBytes) + " of the heap is estimated to be reclaimable. CONFIRM IT BEFORE " + + "ACTING: this is arithmetic over column-width statistics, not a measurement of the table. " + + confirm + + " Once confirmed, the remedy order matters. Fix the CAUSE first (get_pg_xmin_horizon and " + + "get_pg_autovacuum_health) or the space comes straight back. To reclaim it, prefer VACUUM - " + + "which returns space to the table's free space map for reuse without a lock - or pg_repack, " + + "which rewrites the table with only a brief lock. VACUUM FULL returns space to the operating " + + "system but holds an ACCESS EXCLUSIVE lock for the entire rewrite, blocking every read and " + + "write against the table, and needs free disk equal to the finished table plus its indexes. " + + "It is the last option, not the first."; + } + + /// + /// The dead-tuple fraction — a NARROWER but genuinely measured quantity, from the server's own counters + /// rather than our arithmetic, and available whatever the grants are. This is what the tool falls back + /// to when the estimate is suppressed, so that a permissions gap degrades the answer instead of + /// removing it. + /// + internal static string DeadTupleFinding(long liveTuples, long deadTuples) + { + if (liveTuples < 0) + { + return "This table has never been analyzed, so PostgreSQL has no row count to compare its dead " + + "tuples against."; + } + + var total = liveTuples + deadTuples; + if (total <= 0) + { + return "No live or dead tuples recorded - the table is empty or its statistics have not been " + + "gathered yet."; + } + + var pct = Math.Round((double)deadTuples / total * 100, 2); + return pct.ToString("0.##", CultureInfo.InvariantCulture) + "% of this table's tuples are dead. " + + "This is MEASURED rather than estimated - it comes from the server's own counters and needs " + + "no width model, so it is the number to trust when the bloat estimate is suppressed. It is " + + "also NARROWER than bloat: it counts dead tuples awaiting vacuum, and NOT the free space a " + + "completed vacuum already left behind and did not return to the operating system, which is " + + "usually the larger share on a table that has been churning for a while."; + } + + /// + /// The row's severity band, computed HERE rather than in the browser — the web renderer's rule is that + /// it never re-derives a band, so a grid can only colour a row if the read hands it one. This is what + /// gives the web dashboard the same visual cue the WPF grid gets from its row-style triggers; without + /// it the two front ends disagree about which rows look urgent, which was a review finding. + /// + /// Unknown for a SUPPRESSED estimate, deliberately, and it is the whole reason this is not + /// a plain percentage band: the row has no trustworthy number, so colouring it by severity would assert + /// exactly what the suppression denies. It matches the neutral grey the WPF grid paints for the same + /// state. + /// + internal static string BloatSeverity(DarlingPgTableBloatReader.PgTableBloatRow r) + { + if (DarlingPgTableBloatReader.EstimateIsUnpublishable(r)) + { + return "Unknown"; + } + + return r.BloatPctEstimate >= 50m ? "Critical" + : r.BloatPctEstimate >= 20m ? "Warning" + : "Healthy"; + } + + /// Bytes for PROSE only; the payload carries raw _bytes and a numeric _mb. + internal static string DescribeBytes(long bytes) + { + if (bytes < 0) + { + return "an unmeasured size"; + } + + if (bytes >= 1024L * 1024 * 1024) + { + return Math.Round(bytes / 1024.0 / 1024 / 1024, 2).ToString("0.##", CultureInfo.InvariantCulture) + " GB"; + } + + if (bytes >= 1024L * 1024) + { + return Math.Round(bytes / 1024.0 / 1024, 1).ToString("0.#", CultureInfo.InvariantCulture) + " MB"; + } + + return Math.Round(bytes / 1024.0, 0).ToString("0", CultureInfo.InvariantCulture) + " KB"; + } + + [McpServerTool(Name = "get_pg_table_bloat")] + [Description( + "PostgreSQL per-table bloat - the damage autovacuum lag causes, next to the cause chain " + + "get_pg_autovacuum_health, get_pg_wraparound_risk and get_pg_xmin_horizon already report. THE BLOAT " + + "FIGURE IS AN ESTIMATE, NOT A MEASUREMENT: it is arithmetic over PostgreSQL's column-width " + + "statistics and never reads the table, which is what makes it cheap enough to collect hourly. " + + "Against pgstattuple on real tables with current statistics it was within about 2 percentage " + + "points; with STALE statistics it was wrong by 81 percentage points, so the tool suppresses the " + + "number outright rather than captioning it when the statistics behind it are stale, when the table " + + "has never been analyzed, or when the monitoring login cannot SELECT from the table (pg_monitor " + + "alone does not confer that, and in that state the estimator silently returns large numbers). " + + "Measured sizes, the dead-tuple fraction from the server's own counters, and the growth across the " + + "window are always reported, so a suppressed estimate degrades the answer rather than removing it. " + + "Every row ships the exact pgstattuple command that would settle the question. Collected hourly per " + + "database on writers only, for tables of at least 1 MB.")] + public static async Task GetPgTableBloat( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 168 (seven days), which shows whether the waste is growing or being reclaimed.")] int hours_back = 168, + [Description("Maximum tables to return, biggest estimated waste first. Default 25.")] int limit = 25, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + var rows = await DarlingPgTableBloatReader.GetPgTableBloatAsync( + postgres, resolved.ServerId, start, end, limit); + + if (rows.Count == 0) + { + return await EmptyAsync(postgres, resolved.ServerId, resolved.ServerName, hours_back, start, end); + } + + var suppressedCount = rows.Count(r => EstimateSuppressionReason(r) != null); + var trusted = rows.Where(r => EstimateSuppressionReason(r) == null).ToList(); + var trustedBloatBytes = trusted.Sum(r => Math.Max(r.BloatBytesEstimate, 0)); + var anyPgstattuple = rows.Any(r => r.PgstattupleAvailable); + var permissionsSuspected = rows.Count(r => r.EstimateUnavailable); + + var tables = rows.Select(r => + { + var suppression = EstimateSuppressionReason(r); + var qualified = (r.SchemaName ?? "public") + "." + (r.TableName ?? "?"); + var heapGrowth = r.HeapBytes >= 0 && r.FirstHeapBytes >= 0 + ? r.HeapBytes - r.FirstHeapBytes + : (long?)null; + + return new + { + database = r.DatabaseName, + schema = r.SchemaName, + table = r.TableName, + /* Server-computed, because the browser never re-derives a band. Drives the web grid's + cell colour and matches what the WPF row-style triggers paint. */ + severity = BloatSeverity(r), + + /* MEASURED, always reported. These are pg_relation_size readings and are true whatever + the statistics or the grants look like. */ + heap_bytes = r.HeapBytes, + heap_mb = r.HeapBytes >= 0 ? Math.Round(r.HeapBytes / 1024.0 / 1024, 2) : (double?)null, + toast_bytes = r.ToastBytes, + index_bytes = r.IndexBytes, + /* Beside the heap estimate, never inside it: the estimator models heap tuples and has + no model for TOAST, and pg_stattuple on a relation measures its heap and not its + TOAST either - so folding TOAST in would make the accuracy claim meaningless. */ + toast_note = r.ToastBytes > 0 + ? "This table has out-of-line TOAST storage. The bloat estimate covers the HEAP " + + "only - TOAST is reported here as a measured size because the estimator has no " + + "model for it, and a TOAST relation can be bloated independently of its heap." + : null, + + /* THE ESTIMATE - nulled, not captioned, when it cannot be trusted. */ + bloat_bytes_estimate = suppression is null ? r.BloatBytesEstimate : (long?)null, + bloat_mb_estimate = suppression is null + ? Math.Round(r.BloatBytesEstimate / 1024.0 / 1024, 2) + : (double?)null, + bloat_pct_estimate = suppression is null ? r.BloatPctEstimate : (decimal?)null, + estimate_is_an_estimate = true, + estimate_suppressed = suppression != null, + estimate_suppression_reason = suppression, + bloat_finding = suppression is null + ? BloatFinding(r.BloatPctEstimate, r.BloatBytesEstimate, r.PgstattupleAvailable, qualified) + : null, + + /* MEASURED fallback, always reported - this is what a suppressed estimate degrades to + rather than to nothing. */ + live_tuples = r.LiveTuples, + dead_tuples = r.DeadTuples, + dead_tuple_pct = r.LiveTuples >= 0 && r.LiveTuples + r.DeadTuples > 0 + ? Math.Round((double)r.DeadTuples / (r.LiveTuples + r.DeadTuples) * 100, 2) + : (double?)null, + dead_tuple_finding = DeadTupleFinding(r.LiveTuples, r.DeadTuples), + + /* The trust signals, reported whether or not they fired, so a reader can see WHY the + estimate was or was not published rather than having to take it on faith. */ + mods_since_analyze = r.ModsSinceAnalyze, + mods_since_analyze_pct_of_rows = r.LiveTuples > 0 + ? Math.Round((double)r.ModsSinceAnalyze / r.LiveTuples * 100, 1) + : (double?)null, + last_analyzed = r.LastAnalyzed?.ToString("o"), + fillfactor = r.FillFactor, + fillfactor_note = r.FillFactor < 100 + ? "This table has a non-default fillfactor, so it RESERVES free space on purpose for " + + "HOT updates. The estimate accounts for that reservation; pgstattuple's free_space " + + "does not, and will therefore report a larger figure for the same healthy table. " + + "Measured on a fillfactor-70 table: 3.79% estimated against 32.52% from " + + "pgstattuple, and the estimate is the more useful reading of the two." + : null, + + /* The estimator's own inputs, so the number can be argued with rather than believed. */ + estimated_tuple_bytes = r.EstimatedTupleBytes, + estimated_heap_pages = r.EstimatedHeapPages, + heap_pages = r.HeapPages, + alignment_bytes = r.AlignmentBytes, + + /* The escalation path, per table, with the command already written out. */ + pgstattuple_available = r.PgstattupleAvailable, + exact_measurement_command = r.PgstattupleAvailable + ? $"SELECT * FROM pgstattuple('{qualified}');" + : $"CREATE EXTENSION pgstattuple; SELECT * FROM pgstattuple('{qualified}');", + + /* The trend, on the MEASURED size, so it survives a suppressed estimate. */ + heap_bytes_growth_in_window = heapGrowth, + bloat_bytes_estimate_growth_in_window = suppression is null + ? r.BloatBytesEstimate - r.FirstBloatBytesEstimate + : (long?)null, + sample_count = r.SampleCount, + first_seen_at = r.FirstSeenAt.ToString("o"), + measured_at = r.MeasuredAt.ToString("o"), + }; + }) + .ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + status = "table_bloat", + table_count = tables.Count, + /* Named for what it is: a sum of ESTIMATES over the tables whose estimates were fit to + publish. Not "reclaimable bytes", which is a promise the arithmetic cannot make. */ + estimated_bloat_bytes_over_trusted_rows = trustedBloatBytes, + estimated_bloat_mb_over_trusted_rows = Math.Round(trustedBloatBytes / 1024.0 / 1024, 2), + trusted_estimate_count = trusted.Count, + suppressed_estimate_count = suppressedCount, + pgstattuple_available_anywhere = anyPgstattuple, + limit_reached = tables.Count >= limit, + /* Named at the top so it is read before any number below it. */ + figures_are_estimates_not_measurements = true, + note = "Every bloat figure here is an ESTIMATE computed from PostgreSQL's column-width " + + "statistics; the table itself is never read, which is what makes hourly collection " + + "affordable. Sizes, tuple counts and the dead-tuple fraction ARE measured. Against " + + "pgstattuple on real tables the estimate was within about 2 percentage points where " + + "the statistics were current, and wrong by 81 percentage points where they were not " + + "- so estimates whose inputs cannot be trusted are suppressed rather than captioned, " + + "and each row carries the pgstattuple command that would settle it exactly." + + (permissionsSuspected > 0 + ? $" {permissionsSuspected} table(s) have NO usable column statistics. The usual " + + "cause is permissions rather than a missing ANALYZE: pg_stats is filtered by " + + "SELECT privilege and pg_monitor does not grant it on user tables. Granting " + + "pg_read_all_data (PostgreSQL 14+) to the monitoring role fixes this for the " + + "whole instance; on PostgreSQL 13 that role does not exist and explicit GRANT " + + "SELECT is the only route." + : string.Empty) + + (tables.Count >= limit + ? $" The row limit of {limit} was REACHED, so the totals cover only the tables " + + "returned. Raise limit for the full picture." + : string.Empty), + tables, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_table_bloat", ex); + } + } + + /// + /// Which KIND of nothing an empty result is. The engine question is asked FIRST (#2532): "no bloat" is a + /// statement about a PostgreSQL instance and is simply false said about a SQL Server target. + /// The denominator is the DATA, on the same relation the read walks, because + /// pg_table_bloat_stats is a PERIODIC surface — a row exists per qualifying table every cycle, so + /// any stored sample proves somebody looked, and zero rows on a collected server means every table is + /// under the 1 MB floor. + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, int serverId, string serverName, int hoursBack, DateTime start, DateTime end) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, serverId, serverName, "pg_table_bloat_stats"); + if (gated != null) + { + return gated; + } + + var probe = await DarlingPgTableBloatReader.ProbePgTableBloatAsync(postgres, serverId, start, end); + + var hints = new + { + server = serverName, + hours_back = hoursBack, + rows_in_window = probe.RowsInWindow, + snapshots_in_window = probe.SnapshotsInWindow, + ever_collected = probe.RowsEver > 0, + }; + + if (probe.SnapshotsInWindow >= 1) + { + return McpHelpers.Status( + "empty", + $"No tables were recorded for {serverName} in the last {hoursBack} hour(s) even though " + + "collection ran. Every table on every collected database is under the collector's 1 MB " + + "floor, below which bloat is not an actionable amount of space. That is a genuine " + + "all-clear rather than missing data.", + hints); + } + + return McpHelpers.Status( + "unavailable", + probe.RowsEver > 0 + ? $"No bloat snapshots were collected for {serverName} in the last {hoursBack} hour(s), so " + + "the window says nothing either way. Collection HAS run for this server outside the " + + "window: widen hours_back, or use get_collection_health to find where it stopped." + : $"No bloat snapshots have EVER been collected for {serverName}, so there is nothing to " + + "read and this is NOT an all-clear. Check that collection is running and enabled for this " + + "server - and note this collector is gated OFF on read replicas, because the statistics " + + "it needs to judge its own accuracy read as zero there.", + hints); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTrendTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTrendTools.cs new file mode 100644 index 000000000..0fb09cf6e --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgTrendTools.cs @@ -0,0 +1,617 @@ +// Copyright (c) Erik Darling Data. All rights reserved. +// Licensed under the terms in the LICENSE file in the repository root. + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// Time-series reads for PostgreSQL (#2663). The service had fourteen trend reads and none worked on a +/// PostgreSQL target, so every PostgreSQL answer described one window and none of them answered "is this +/// getting worse". +/// +/// Every subject here is OPTIONAL and every read says which subject it chose. That is what makes them +/// usable as web panels with no drill-down plumbing, and it is what stops an automatic pick from being +/// mistaken for an answer about the event, query, I/O path or database the caller had in mind. +/// +[McpServerToolType] +public sealed class DarlingMcpPgTrendTools +{ + [McpServerTool(Name = "get_pg_wait_trend"), Description("Gets a time series for ONE PostgreSQL wait event: how much the server waited on it in each collection interval, normalised per second. Use get_pg_wait_sampling first to find which events dominate, then this to see whether one is growing. Summed across the queries that waited, because this answers a question about the SERVER - per-query attribution for a single window is what get_pg_wait_sampling already gives. The figures are estimates from a sampling profiler (samples x profile period), not measured durations, so treat the SHAPE as the finding rather than the absolute number. An interval where pg_wait_sampling's profile was reset - by pg_wait_sampling_reset_profile() or a server restart - is flagged, and reports everything since the reset rather than a misleadingly quiet interval.")] + public static async Task GetPgWaitTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("The exact wait event name, e.g. DataFileRead, WALWrite. Omit to follow whichever event dominates the window.")] string? wait_event = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var start = windowEnd.AddHours(-hours_back); + + /* With no event named, follow the one that actually dominates this server rather than a name + somebody guessed. Which one was chosen is reported, so the answer is never about a different + event than the reader thinks. */ + var chosen = string.IsNullOrWhiteSpace(wait_event) + ? await DarlingPgTrendReader.GetDominantWaitEventAsync(postgres, resolved.ServerId, start, windowEnd) + : wait_event.Trim(); + + if (string.IsNullOrWhiteSpace(chosen)) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wait_sampling") + ?? McpHelpers.Status( + "empty", + $"No wait event was sampled on {resolved.ServerName} in the last {hours_back} " + + "hour(s), so there is nothing to follow. pg_wait_sampling needs the extension " + + "loaded; get_pg_extensions reports whether it is."); + } + + var points = await DarlingPgTrendReader.GetWaitTrendAsync( + postgres, resolved.ServerId, chosen, start, windowEnd); + + if (points.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wait_sampling") + ?? McpHelpers.Status( + "empty", + $"No samples for wait event '{chosen}' on {resolved.ServerName} in the last " + + $"{hours_back} hour(s). A trend needs at least TWO snapshots to difference, so a " + + "window holding one collection is legitimately empty here even when the event is " + + "being sampled. get_pg_wait_sampling lists the events this server does record."); + } + + var resets = points.Count(p => p.CounterReset); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + wait_event = chosen, + /* Said plainly when it was not asked for, so a caller never reads this as an answer about + the event they had in mind. */ + wait_event_source = string.IsNullOrWhiteSpace(wait_event) + ? "chosen automatically: the most-sampled event in this window that is an actual WAIT. " + + "The CPU class - pg_wait_sampling's 'Running', meaning the backend was not waiting - " + + "is excluded from the choice because it dominates any healthy server's profile and " + + "would answer the opposite of the question. Name it explicitly to follow it." + : "as requested", + hours_back, + status = "wait_trend", + point_count = points.Count, + counter_reset_count = resets, + note = "estimated_wait_ms_per_second is samples x the profiler's period, over the interval's " + + "length - an estimate from a sampling profiler, not a measured duration, so the shape " + + "over time is the finding rather than the absolute value. Per SECOND because " + + "collection intervals are not uniform: a restart or a slow cycle stretches one, and a " + + "per-interval total would render that as a spike." + + (resets > 0 + ? $" {resets} interval(s) span a profile RESET - pg_wait_sampling_reset_profile() " + + "or a server restart - and report everything since the reset rather than a " + + "quiet interval, which is what clamping the difference at zero would have shown." + : string.Empty), + points = points.Select(p => new + { + collection_time = p.CollectionTimeUtc, + sample_count = p.SampleCount, + estimated_wait_ms_per_second = Math.Round(p.EstimatedWaitMsPerSecond, 3), + backend_count = p.BackendCount, + counter_reset = p.CounterReset, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading the PostgreSQL wait trend failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_query_duration_trend"), Description("Gets a time series for ONE PostgreSQL statement by queryid: what a single execution cost in each collection interval, how many times it ran, and its call rate. This is the regression read - a query whose mean execution time steps up and stays up has changed plan or lost an index, and the step is visible here where a single-window average hides it. Use get_pg_top_queries first to get a queryid. mean_exec_ms is null rather than zero for an interval where the statement did not run, because a mean over no calls is absent rather than fast.")] + public static async Task GetPgQueryDurationTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("The queryid from get_pg_top_queries; PostgreSQL query ids can be negative. Omit to follow the statement that spent the most time in the window.")] string? queryid = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + long parsedQueryId; + var requested = !string.IsNullOrWhiteSpace(queryid); + + /* Taken as TEXT and parsed here: a PostgreSQL query id is a signed 64-bit hash that routinely + exceeds what a JSON number survives intact, and a client that rounds one silently asks about a + statement that does not exist. */ + if (requested && !long.TryParse(queryid!.Trim(), NumberStyles.Integer, CultureInfo.InvariantCulture, out parsedQueryId)) + { + return McpHelpers.Status( + "error", + $"queryid '{queryid}' is not a 64-bit integer. PostgreSQL query ids are signed and often " + + "negative; pass the value from get_pg_top_queries exactly as it appears."); + } + + try + { + var start = windowEnd.AddHours(-hours_back); + + if (!requested) + { + var top = await DarlingPgTrendReader.GetTopQueryIdAsync(postgres, resolved.ServerId, start, windowEnd); + + if (top is null) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_statement_stats") + ?? McpHelpers.Status( + "empty", + $"No statement recorded execution time on {resolved.ServerName} in the last " + + $"{hours_back} hour(s), so there is nothing to follow."); + } + + parsedQueryId = top.Value; + } + else + { + parsedQueryId = long.Parse(queryid!.Trim(), NumberStyles.Integer, CultureInfo.InvariantCulture); + } + + var points = await DarlingPgTrendReader.GetQueryDurationTrendAsync( + postgres, resolved.ServerId, parsedQueryId, start, windowEnd); + + if (points.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_statement_stats") + ?? McpHelpers.Status( + "empty", + $"No samples for queryid {parsedQueryId} on {resolved.ServerName} in the last " + + $"{hours_back} hour(s). pg_stat_statements evicts statements under memory " + + "pressure, so a queryid that was there yesterday can be gone rather than idle - " + + "get_pg_top_queries shows what the server currently tracks."); + } + + var ran = points.Where(p => p.Calls > 0).ToList(); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + queryid = parsedQueryId.ToString(CultureInfo.InvariantCulture), + queryid_source = requested + ? "as requested" + : "chosen automatically: the statement with the most execution time in this window", + hours_back, + status = "query_duration_trend", + point_count = points.Count, + intervals_with_calls = ran.Count, + /* The two ends of the series that actually ran, so a step is visible without reading every + point - and null when nothing ran, rather than a shape invented from no executions. */ + first_mean_exec_ms = ran.Count > 0 ? Math.Round(ran[0].MeanExecMs, 3) : (double?)null, + last_mean_exec_ms = ran.Count > 0 ? Math.Round(ran[^1].MeanExecMs, 3) : (double?)null, + note = "mean_exec_ms is the interval's total execution time over its calls - what ONE " + + "execution cost then. It is null for an interval with no calls, because a mean over " + + "no executions is absent rather than zero. A step that persists is a plan or index " + + "change; a spike that recovers is usually contention, which get_pg_wait_trend and " + + "get_pg_blocking speak to.", + points = points.Select(p => new + { + collection_time = p.CollectionTimeUtc, + calls = p.Calls, + total_exec_ms = Math.Round(p.TotalExecMs, 3), + mean_exec_ms = p.Calls > 0 ? Math.Round(p.MeanExecMs, 3) : (double?)null, + calls_per_second = Math.Round(p.CallsPerSecond, 4), + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading the PostgreSQL query duration trend failed: {ex.Message}"); + } + } + + [McpServerTool(Name = "get_pg_io_trend"), Description("Gets a time series for ONE PostgreSQL (backend_type, context) pair from pg_stat_io: read, write and extend rates per second, the buffer-cache hit ratio for each interval, and per-operation latency where the server measures it. Use get_pg_io_stats first to see which combinations are busy, then this to see whether one is growing or slowing. The subject is a PAIR because a hit ratio summed across contexts is meaningless - bulkread is a sequential scan deliberately bypassing the buffer pool with a ring buffer, so averaging its misses with the normal context's understates both and they have opposite remedies. Object types are summed together, which cannot distort the ratio because WAL rows report no buffer hits. Reports whether the server tracks I/O TIMING at all: track_io_timing is OFF by default in PostgreSQL, and its zero read_time would otherwise divide out to a latency of 0.000 ms that reads as an impossibly fast disk rather than an unmeasured one. Also reports whether byte volumes are MEASURED (PostgreSQL 18's read_bytes/write_bytes) or ESTIMATED as operations x block size (pre-18) - the two are different quantities and 18 moves several blocks per operation, so the older estimate undercounts there. Write counters are null rather than zero on Amazon Aurora, where backends do not write data files. An interval spanning a pg_stat_reset_shared('io') or a restart is flagged and reports everything since the reset, rather than the quiet interval that clamping the difference at zero would show. Requires PostgreSQL 16 or later; valid on a standby.")] + public static async Task GetPgIoTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Which backend did the I/O, e.g. 'client backend', 'autovacuum worker', 'checkpointer'. Naming this alone follows the busiest CONTEXT for that backend; omit both to follow whichever pair moved the most I/O.")] string? backend_type = null, + [Description("Why the I/O happened: normal, bulkread, bulkwrite, vacuum, index, walreplay. Naming this alone follows the busiest BACKEND in that context; omit both to follow whichever pair moved the most I/O.")] string? context = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var start = windowEnd.AddHours(-hours_back); + + var askedBackend = string.IsNullOrWhiteSpace(backend_type) ? null : backend_type.Trim(); + var askedContext = string.IsNullOrWhiteSpace(context) ? null : context.Trim(); + var requested = askedBackend is not null && askedContext is not null; + string chosenBackend; + string chosenContext; + + if (requested) + { + chosenBackend = askedBackend!; + chosenContext = askedContext!; + } + else + { + /* HALF a subject is a real request, not an error and not an excuse to ignore the half that + was given: "the busiest context for the autovacuum worker" is a question somebody asks. + Whichever half was named CONSTRAINS the choice, and subject_source says which half was + chosen for them - the alternative was accepting a backend type and quietly answering + about a different one. */ + var dominant = await DarlingPgTrendReader.GetDominantIoSubjectAsync( + postgres, resolved.ServerId, start, windowEnd, askedBackend, askedContext); + + if (dominant is null) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_io_stats") + ?? McpHelpers.Status( + "empty", + (askedBackend ?? askedContext) is null + ? $"No backend type recorded any I/O or buffer activity on " + + $"{resolved.ServerName} in the last {hours_back} hour(s), so there is " + + "nothing to follow. Buffer HITS count here, not only physical reads and " + + "writes, so this is a genuinely idle window rather than a fully cached " + + "one. pg_stat_io also needs PostgreSQL 16 or later - on an older major " + + "the collector never runs, which get_collection_health reports." + : $"Nothing matching {(askedBackend is not null ? $"backend type '{askedBackend}'" : $"context '{askedContext}'")} " + + $"recorded any I/O or buffer activity on {resolved.ServerName} in the " + + $"last {hours_back} hour(s). get_pg_io_stats lists the combinations " + + "this server reports; omit both parameters to follow whichever pair is " + + "busiest."); + } + + chosenBackend = dominant.Value.BackendType; + chosenContext = dominant.Value.Context; + } + + var points = await DarlingPgTrendReader.GetIoTrendAsync( + postgres, resolved.ServerId, chosenBackend, chosenContext, start, windowEnd); + + if (points.Count == 0) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_io_stats") + ?? McpHelpers.Status( + "empty", + $"No I/O samples for '{chosenBackend}' in the '{chosenContext}' context on " + + $"{resolved.ServerName} in the last {hours_back} hour(s). A trend needs at least " + + "TWO snapshots to difference, so a window holding one collection is legitimately " + + "empty here even while the combination is being collected. get_pg_io_stats lists " + + "the combinations this server actually reports."); + } + + /* Asked of the server's OWN configuration rather than inferred from the zeros, because the two + readings a zero latency permits - "the disk is instant" and "nobody is timing it" - are not + distinguishable in the counters and only one of them is ever true. */ + var timingSetting = await DarlingPgTrendReader.GetIoTimingTrackedAsync( + postgres, resolved.ServerId, windowEnd); + var timingObserved = points.Any(p => p.ReadTimeMs > 0 || p.WriteTimeMs > 0); + var timingTracked = timingSetting ?? timingObserved; + + var writesTracked = points.Any(p => p.WriteCountersTracked); + var bytesMeasured = points.Any(p => p.BytesMeasured); + var bytesEstimated = !bytesMeasured && points.Any(p => p.BytesEstimable); + var resets = points.Count(p => p.CounterReset); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + backend_type = chosenBackend, + context = chosenContext, + /* Said plainly when it was not asked for, so a caller never reads this as an answer about + the combination they had in mind - and said per HALF, because half a subject can be + named and the other half then chosen inside it. */ + subject_source = requested + ? "as requested" + : (askedBackend is not null + ? "backend_type as requested; context chosen automatically within it: " + : askedContext is not null + ? "context as requested; backend_type chosen automatically within it: " + : "chosen automatically: ") + + "the pair that moved the most read, write and extend operations in this window. " + + "Ranked on OPERATIONS rather than on read time - which is what the single-window " + + "read ranks by - because track_io_timing is off by default, and over a store full " + + "of zeros that ranking degenerates into picking a name alphabetically. Name both " + + "parameters to pin the pair exactly.", + context_meaning = DarlingPgIoReader.ContextMeaning(chosenContext), + hours_back, + status = "io_trend", + point_count = points.Count, + counter_reset_count = resets, + total_reads = points.Sum(p => p.Reads), + total_writes = writesTracked ? points.Sum(p => p.Writes) : (long?)null, + total_extends = points.Sum(p => p.Extends), + total_hits = points.Sum(p => p.Hits), + total_read_bytes = bytesMeasured || bytesEstimated ? points.Sum(p => p.ReadBytes) : (decimal?)null, + total_write_bytes = (bytesMeasured || bytesEstimated) && writesTracked + ? points.Sum(p => p.WriteBytes) + : (decimal?)null, + peak_reads_per_second = Math.Round(points.Max(p => p.ReadsPerSecond), 3), + /* Named at the top rather than only per point: a caller has to know which of these are + measured before it draws any conclusion from a number below. */ + write_counters_tracked = writesTracked, + io_timing_tracked = timingTracked, + io_timing_source = timingSetting is null + ? "inferred from the data - this server's configuration has not been collected, so " + + "track_io_timing is unknown and the answer here is simply whether any non-zero I/O " + + "time appears in the window" + : "the target's own track_io_timing, as collected into pg_server_config", + bytes_source = bytesMeasured + ? "measured" + : (bytesEstimated ? "estimated_from_block_size" : "unavailable"), + note = "Every figure is the difference between CONSECUTIVE snapshots, normalised per SECOND " + + "by the interval's own length - collection cadence is not uniform, and a per-interval " + + "total would render a slow sweep as a spike in the data rather than in the server. " + + "cache_hit_pct is scoped to this pair, which is the only scope where it means " + + "anything." + + (writesTracked + ? string.Empty + : " This server tracks NO write counters - the signature of Amazon Aurora, where " + + "backends do not write data files and the storage layer does. The write fields " + + "are null rather than zero, because absent here means unmeasured.") + + (resets > 0 + ? $" {resets} interval(s) span a statistics RESET - pg_stat_reset_shared('io') or " + + "a restart - and report everything since the reset rather than the quiet " + + "interval that clamping the difference at zero would have shown." + : string.Empty), + timing_note = timingTracked + ? "avg_read_ms and avg_write_ms are measured per-operation latencies: this server has " + + "track_io_timing on." + : "avg_read_ms and avg_write_ms are NULL throughout because this server does not " + + "measure I/O time. track_io_timing is off by DEFAULT in PostgreSQL, so this is the " + + "ordinary configuration rather than a fault - but it means the volume figures below " + + "are the only I/O evidence here, and nothing in this store can say whether the " + + "storage is slow. Turning it on costs a clock read per operation; measure that on " + + "the platform before enabling it fleet-wide.", + bytes_note = bytesMeasured + ? "Byte rates are MEASURED, from PostgreSQL 18's read_bytes / write_bytes. They are not " + + "comparable with the figures a pre-18 server reports, which are operations x block " + + "size: 18 reads several blocks per operation, so the older estimate undercounts." + : (bytesEstimated + ? "Byte rates are ESTIMATED as operations x op_bytes, which is exact below " + + "PostgreSQL 18 because one operation moves one block. 18 removed op_bytes and " + + "measures the bytes directly instead." + : "This server reports no byte figures at all - neither op_bytes nor the measured " + + "columns PostgreSQL 18 replaced it with - so the byte rates are null. The " + + "operation counts and the hit ratio are unaffected."), + points = points.Select(p => new + { + collection_time = p.CollectionTimeUtc, + interval_seconds = Math.Round(p.IntervalSeconds, 1), + reads_per_second = Math.Round(p.ReadsPerSecond, 3), + writes_per_second = p.WriteCountersTracked ? Math.Round(p.WritesPerSecond, 3) : (double?)null, + extends_per_second = Math.Round(p.ExtendsPerSecond, 3), + /* The rate beside the ratio, because they answer different questions: cache_hit_pct + says what share the pool absorbed, hits_per_second says how much work there was to + absorb. A pair can hold 100% while doing almost nothing. */ + hits_per_second = Math.Round(p.HitsPerSecond, 3), + /* Ring-buffer REUSE is a different thing and is not folded in here: a bulk operation + recycling its own buffers is not pressure on the pool, and conflating the two is the + standard misreading of pg_stat_io. */ + evictions_per_second = p.IntervalSeconds > 0 + ? Math.Round(p.Evictions / p.IntervalSeconds, 3) + : 0, + cache_hit_pct = p.CacheHitPct is { } hit ? Math.Round(hit, 2) : (double?)null, + /* Nulled when the server does not time I/O, rather than passing the 0.000 the + arithmetic produces: that value is a statement about the configuration, not about + the disk, and printed as a latency it is the most reassuring wrong number here. */ + avg_read_ms = timingTracked && p.AvgReadMs is { } read ? Math.Round(read, 3) : (double?)null, + avg_write_ms = timingTracked && p.WriteCountersTracked && p.AvgWriteMs is { } write + ? Math.Round(write, 3) + : (double?)null, + read_bytes_per_second = bytesMeasured || bytesEstimated + ? Math.Round(p.ReadBytesPerSecond, 1) + : (double?)null, + write_bytes_per_second = (bytesMeasured || bytesEstimated) && p.WriteCountersTracked + ? Math.Round(p.WriteBytesPerSecond, 1) + : (double?)null, + counter_reset = p.CounterReset, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_io_trend", ex); + } + } + + [McpServerTool(Name = "get_pg_database_trend"), Description("Gets a time series for ONE PostgreSQL database from pg_stat_database: temp-file spills, the buffer-cache hit ratio, deadlocks and the rollback share, interval by interval. This is where the cache hit ratio becomes usable at all - pg_stat_database's counters are cumulative since the last reset, so the ratio computed from them raw is a lifetime average that barely moves, and a database that fell off a cliff an hour ago still reports 99% because of the weeks behind it. Differenced per interval, the cliff is visible. Temp files are the same story: 'this database spilled 40 GB this week' does not say whether it was one bad afternoon or a steady leak, and the two have different fixes. Deadlocks are reported as a COUNT per interval rather than a rate, because they are discrete server-recorded events. Omit database to follow the biggest temp-file spiller. PostgreSQL's shared-relations row - the cluster-wide catalog, which has a NULL database name - is followable by passing '(shared relations)' but is never chosen automatically. An interval spanning a pg_stat_reset or a crash restart is flagged and reports everything since the reset, rather than the quiet interval that clamping the difference at zero would show. Works on every PostgreSQL major and on a standby, where sorts spill exactly the way they do on a writer.")] + public static async Task GetPgDatabaseTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("The database to follow. Pass '(shared relations)' for PostgreSQL's cluster-wide catalog row. Omit to follow the biggest temp-file spiller in the window.")] string? database = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var start = windowEnd.AddHours(-hours_back); + var requested = !string.IsNullOrWhiteSpace(database); + string? chosen; + + if (requested) + { + var trimmed = database!.Trim(); + + /* The label the single-window read prints for PostgreSQL's NULL datname row is accepted back + as input. Without this the shared-relations series would be visible in one read and + unaskable in the other, which is the kind of gap that makes a reader assume the data is + missing rather than the parameter is. */ + chosen = string.Equals(trimmed, DarlingMcpPgDatabaseTools.SharedRelationsLabel, StringComparison.OrdinalIgnoreCase) + ? null + : trimmed; + } + else + { + chosen = await DarlingPgTrendReader.GetTopDatabaseAsync( + postgres, resolved.ServerId, start, windowEnd); + + if (chosen is null) + { + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_database_stats") + ?? McpHelpers.Status( + "empty", + $"No database recorded block accesses or temp files on {resolved.ServerName} in " + + $"the last {hours_back} hour(s), so there is nothing to follow."); + } + } + + var points = await DarlingPgTrendReader.GetDatabaseTrendAsync( + postgres, resolved.ServerId, chosen, start, windowEnd); + + var label = chosen ?? DarlingMcpPgDatabaseTools.SharedRelationsLabel; + + if (points.Count == 0) + { + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_database_stats"); + if (gated != null) + { + return gated; + } + + /* The denominator, from the SAME relation the trend walks: pg_stat_database is a periodic + surface, so any stored sample proves somebody looked, and TWO are needed before a + cumulative counter can be differenced at all. A server whose first cycle has just run is + in exactly that state, and telling its operator the database was quiet would be a + confident wrong answer for as long as it takes the second cycle to land. */ + var (samplesInWindow, everCollected) = await DarlingPgDatabaseReader.GetCoverageAsync( + postgres, resolved.ServerId, start, windowEnd); + + return McpHelpers.Status( + "empty", + $"No differenced intervals for database '{label}' on {resolved.ServerName} in the last " + + $"{hours_back} hour(s). " + + (samplesInWindow >= 2 + ? "The collector has samples in this window, so this database is either not one of " + + "the ones it reports or its name is spelled differently - get_pg_database_stats " + + "lists what the server actually reports." + : samplesInWindow == 1 + ? "Only ONE snapshot exists in this window and a cumulative counter needs two " + + "before it can be differenced, so this is too early rather than quiet - the " + + "next collection cycle fills it." + : everCollected + ? "No snapshot at all landed in this window, though this server has been " + + "collected before. get_collection_health says whether the sweep is " + + "running." + : "This server has never had pg_database_stats collected.")); + } + + var totalTempFiles = points.Sum(p => p.TempFiles); + var totalTempBytes = points.Sum(p => p.TempBytes); + var totalHits = points.Sum(p => p.BlksHit); + var totalReads = points.Sum(p => p.BlksRead); + var totalAccesses = totalHits + totalReads; + var totalCommits = points.Sum(p => p.XactCommit); + var totalRollbacks = points.Sum(p => p.XactRollback); + var spilled = points.Where(p => p.TempFiles > 0).ToList(); + var rated = points.Where(p => p.CacheHitPct.HasValue).ToList(); + var resets = points.Count(p => p.CounterReset); + + double? windowHitPct = totalAccesses > 0 + ? Math.Round((double)totalHits / totalAccesses * 100, 2) + : null; + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + database = label, + is_shared_relations = chosen is null, + database_source = requested + ? "as requested" + : "chosen automatically: the biggest temp-file spiller in this window, or the busiest by " + + "block access when nothing spilled. PostgreSQL's shared-relations row is never " + + "chosen automatically - it is the cluster-wide catalog and never the database in " + + "trouble - but it is followable by passing '(shared relations)'.", + hours_back, + status = "database_trend", + point_count = points.Count, + counter_reset_count = resets, + intervals_with_temp_files = spilled.Count, + total_temp_files = totalTempFiles, + total_temp_bytes = totalTempBytes, + total_deadlocks = points.Sum(p => p.Deadlocks), + /* The RANGE, not just the average. The average over the window is what the single-window + read already gives; the reason to difference per interval is that a database can average + 99% and still have spent twenty minutes at 40%, and only the floor says so. */ + cache_hit_pct_window = windowHitPct, + worst_interval_cache_hit_pct = rated.Count > 0 + ? Math.Round(rated.Min(p => p.CacheHitPct!.Value), 2) + : (double?)null, + worst_interval_at = rated.Count > 0 + ? rated.OrderBy(p => p.CacheHitPct!.Value).First().CollectionTimeUtc + : (DateTime?)null, + peak_temp_bytes_per_second = Math.Round(points.Max(p => p.TempBytesPerSecond), 1), + /* The same prose the single-window read serves, from the same method: the scale of a spill + decides whether work_mem is the answer or a plan is, and two copies of that judgement is + how it stops being one judgement. */ + spill_finding = DarlingMcpPgDatabaseTools.SpillFinding(totalTempFiles, totalTempBytes), + cache_finding = DarlingMcpPgDatabaseTools.CacheHitFinding(windowHitPct), + xact_commit = totalCommits, + xact_rollback = totalRollbacks, + rollback_finding = DarlingMcpPgDatabaseTools.RollbackFinding(totalCommits, totalRollbacks), + note = "Every figure is the difference between CONSECUTIVE snapshots, normalised per SECOND " + + "by the interval's own length where it is a rate. cache_hit_pct and rollback_pct are " + + "NULL rather than zero for an interval with no block accesses or no completed " + + "transactions, because a ratio over nothing is absent rather than bad. deadlocks is a " + + "COUNT for the interval, not a rate: they are discrete events, and per-second would " + + "render every real one as four leading zeros." + + (resets > 0 + ? $" {resets} interval(s) span a statistics RESET - pg_stat_reset(), or a crash " + + "restart discarding the statistics - and report everything since the reset " + + "rather than the quiet interval that clamping at zero would have shown." + : string.Empty), + points = points.Select(p => new + { + collection_time = p.CollectionTimeUtc, + transactions_per_second = Math.Round(p.TransactionsPerSecond, 3), + rollback_pct = p.RollbackPct is { } rb ? Math.Round(rb, 2) : (double?)null, + cache_hit_pct = p.CacheHitPct is { } hit ? Math.Round(hit, 2) : (double?)null, + temp_files = p.TempFiles, + temp_bytes = p.TempBytes, + temp_bytes_per_second = Math.Round(p.TempBytesPerSecond, 1), + deadlocks = p.Deadlocks, + counter_reset = p.CounterReset, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_pg_database_trend", ex); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitSamplingTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitSamplingTools.cs new file mode 100644 index 000000000..1384c3fbd --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitSamplingTools.cs @@ -0,0 +1,138 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Globalization; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The MCP surface for sampled PostgreSQL waits, paired with the pg_wait_sampling collector. +/// +/// +/// The stock-PostgreSQL answer to the question get_pg_wait_stats answers on Aurora, and #2629's +/// first entry because #2625 shipped a message that points here: a permanent-gap sentence now tells +/// a non-Aurora operator that pg_wait_sampling covers what Aurora's wait instrumentation cannot. +/// Until this tool existed that pointer led to a Windows-only WPF panel — so on a Linux host, and for any +/// agent anywhere, "why is this slow" had an answer that could not be read. +/// +/// +/// +/// It is a sampler, and the shape of the answer has to say so. Aurora's +/// aurora_stat_system_waits() measures accumulated wait TIME; this extension counts how many +/// periodic samples caught a backend in each state. Multiplying samples by the profile period estimates +/// milliseconds and the reader does, but it is an estimate whose error grows as the event gets rarer, so +/// the sample count travels beside it rather than being converted away. +/// +/// +/// +/// CPU is a row here, not a missing one. The collector maps a NULL event type to CPU/ +/// Running, so a backend that was not waiting is counted rather than dropped — which makes the +/// share column answer "waiting or working?" before it answers "waiting on what?". A profile that is +/// mostly CPU and a profile that is mostly IO call for opposite next steps, and a wait-only view cannot +/// tell them apart. +/// +/// +[McpServerToolType] +public sealed class DarlingMcpPgWaitSamplingTools +{ + [McpServerTool(Name = "get_pg_wait_sampling"), Description("Gets sampled PostgreSQL wait events attributed to query shapes, from the pg_wait_sampling extension. This is the stock-PostgreSQL counterpart of get_pg_wait_stats, which reads an Aurora-only source: use this tool on any self-hosted or non-Aurora PostgreSQL target. A sampling profiler periodically records what each backend is doing, so results are sample COUNTS, and estimated_wait_ms is samples multiplied by the sampling period rather than a measured duration - treat a rare event's estimate as approximate. Rows with event_type CPU mean the backend was running, not waiting, so the profile answers 'waiting or working' as well as 'waiting on what'. queryid joins get_pg_top_queries; queryid 0 is work belonging to no statement, such as a background worker. The profile is cluster-wide and carries no database attribution by design.")] + public static async Task GetPgWaitSampling( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, + [Description("Maximum rows to return. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + validation = McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var rows = await DarlingPgWaitSamplingReader.GetPgWaitSamplingAsync( + postgres, resolved.ServerId, windowEnd.AddHours(-hours_back), windowEnd, limit); + + if (rows.Count == 0) + { + /* Three distinct empty states, and only the last of them is "nothing happened". The + capability answer rules out a wrong-engine target; the precondition answer names an + extension that is not installed — which is the LIKELY one here, because + pg_wait_sampling needs shared_preload_libraries and a server restart, so an operator + who has not done that gets told what to do rather than shown a blank. */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wait_sampling") + ?? await DarlingRuntimePrecondition.StatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wait_sampling") + ?? McpHelpers.Status( + "empty", + $"No sampled waits for {resolved.ServerName} in the last {hours_back} hour(s). The " + + "figures here are per-interval deltas, so a single collection has nothing to " + + "difference against and the window fills on the second one. On a genuinely idle " + + "server this is the healthy state: the profiler samples backends, and an idle " + + "server has none to sample."); + } + + /* The denominator is SAMPLES, not the estimate. Every row's estimate is the same count times + the same period, so the two shares are arithmetically identical — and taking it from the + count keeps the percentage anchored to what was actually observed. */ + var totalSamples = rows.Sum(r => r.SampleCount); + var resetInWindow = rows.Any(r => r.CounterReset); + + var waits = rows.Select(r => new + { + event_type = r.EventType, + wait_event = r.Event, + /* String, like every other queryid on this surface: it is a signed 64-bit value and a + JSON number would round it in any double-decoding parser, silently producing an id + that joins to nothing. */ + queryid = r.QueryId.ToString(CultureInfo.InvariantCulture), + samples = r.SampleCount, + estimated_wait_ms = r.EstimatedWaitMs, + backends = r.BackendCount, + pct_of_samples = totalSamples > 0 + ? Math.Round((double)r.SampleCount / totalSamples * 100, 1) + : 0, + /* Per row, because a reset is per series: one query's profile can be reset while another + accumulated normally, and a single window-level flag would libel both. */ + counter_reset = r.CounterReset, + }); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + total_samples = totalSamples, + note = "Sample COUNTS from a periodic profiler, not measured durations. estimated_wait_ms " + + "is samples multiplied by the sampling period and is approximate — most so for rare " + + "events. event_type CPU means the backend was running rather than waiting." + + (resetInWindow + ? " At least one series was RESET inside this window, so its figures cover only " + + "the time since the reset." + : string.Empty), + waits, + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.Status("error", $"Reading PostgreSQL sampled waits failed: {ex.Message}"); + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitTools.cs index e0d9ea017..89e2e36ea 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWaitTools.cs @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -28,19 +29,20 @@ public static async Task GetPgWaitStats( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history to analyze. Default 24.")] int hours_back = 24, - [Description("Maximum rows to return. Default 20.")] int limit = 20) + [Description("Maximum rows to return. Default 20.")] int limit = 20, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgWaitReader.GetPgWaitStatsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now, limit); @@ -49,11 +51,20 @@ public static async Task GetPgWaitStats( so, rather than letting a caller read "no waits" as "no waiting". */ if (rows.Count == 0) { - return McpHelpers.Status( - "unavailable", - "No PostgreSQL wait data for this server and window. If this server is SQL Server, " - + "use get_wait_stats instead; if it is PostgreSQL but not Aurora, cumulative wait " - + "counters are not available (core PostgreSQL does not provide them)."); + /* Both halves of the old ambiguity are answerable when the store knows the engine (#2532): + a SQL Server target gets the dialect answer, a stock PostgreSQL one gets the Aurora-only + answer, and both say the gap is permanent instead of inviting a hunt. So the sentence + below is reached in exactly two states — an Aurora target with no rows in the window, or + a row whose engine_kind is NULL — and it names those rather than repeating the two the + capability answer has already ruled out. */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wait_stats") + ?? McpHelpers.Status( + "unavailable", + "No PostgreSQL wait data for this server and window. On Aurora, the pg_wait_stats " + + "collector may not have completed a cycle yet. Otherwise the store has not " + + "recorded this server's engine yet, and a target it cannot classify may not be a " + + "PostgreSQL one at all — check list_servers."); } var totalWaitMs = rows.Sum(r => r.TotalWaitTimeMs); diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWraparoundTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWraparoundTools.cs index 1c22cae69..5dfe9f793 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWraparoundTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgWraparoundTools.cs @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -78,27 +79,35 @@ grade calmer than the engine. */ public static async Task GetPgWraparoundRisk( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to analyze, used for the peak comparison. Default 24.")] int hours_back = 24) + [Description("Hours of history to analyze, used for the peak comparison. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgWraparoundReader.GetPgWraparoundAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) { - return McpHelpers.Status( - "unavailable", - "No PostgreSQL freeze-headroom data for this server and window. This collector runs on " - + "any PostgreSQL target, so an empty result means the server is SQL Server, or " - + "pg_wraparound_stats has not collected yet."); + /* This collector has no AppliesTo gate at all, so the capability answer above catches + every engine the store can classify (#2532). What is left is a PostgreSQL target that has + not collected yet, or a row whose engine_kind is NULL — naming "the server is SQL Server" + here would repeat a branch that can no longer reach this line. */ + return await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_wraparound_stats") + ?? McpHelpers.Status( + "unavailable", + "No PostgreSQL freeze-headroom data for this server and window. This collector runs on " + + "any PostgreSQL target, so either pg_wraparound_stats has not collected yet, or the " + + "store has not recorded this server's engine — and a target it cannot classify may " + + "not be a PostgreSQL one at all. Check list_servers."); } var databases = rows diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgXminTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgXminTools.cs index f33499358..1448af572 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgXminTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPgXminTools.cs @@ -14,6 +14,7 @@ using ModelContextProtocol.Server; using Npgsql; using PerformanceMonitor.Common; +using PerformanceMonitor.Darling.Storage; namespace PerformanceMonitor.Darling.Service.Mcp; @@ -59,17 +60,18 @@ public sealed class DarlingMcpPgXminTools public static async Task GetPgXminHorizon( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history to analyze, used for the persistence figures. Default 24.")] int hours_back = 24) + [Description("Hours of history to analyze, used for the persistence figures. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPgXminReader.GetPgXminHorizonAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); @@ -78,6 +80,16 @@ public static async Task GetPgXminHorizon( holder" is a real finding that redirects the investigation rather than a dead end. */ if (rows.Count == 0) { + /* ...but only when this server COULD have a horizon. On a SQL Server target "nothing is + holding back the xmin horizon" is a confident all-clear about a mechanism that does not + exist there, which is the same defect one engine over (#2532). */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync( + postgres, resolved.ServerId, resolved.ServerName, "pg_xmin_horizon"); + if (gated != null) + { + return gated; + } + return JsonSerializer.Serialize(new { server = resolved.ServerName, diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCacheSchedulerTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCacheSchedulerTools.cs index 4876466b3..a0797d68b 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCacheSchedulerTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCacheSchedulerTools.cs @@ -34,21 +34,23 @@ public sealed class DarlingMcpPlanCacheSchedulerTools public static async Task GetPlanCacheBloat( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of data to analyze. Default 24.")] int hours_back = 24) + [Description("Hours of data to analyze. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingPlanCacheSchedulerReader.GetPlanCacheBloatAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No plan cache statistics available in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "plan_cache_stats") + ?? McpHelpers.Status("unavailable", "No plan cache statistics available in the requested time range."); var totalPlans = rows.Sum(r => (long)r.TotalPlans); var totalSingleUse = rows.Sum(r => (long)r.SingleUsePlans); @@ -103,7 +105,8 @@ public static async Task GetCpuSchedulerPressure( { var item = await DarlingPlanCacheSchedulerReader.GetCpuSchedulerPressureAsync(postgres, resolved.ServerId); if (item == null) - return McpHelpers.Status("unavailable", "No CPU scheduler data available. The scheduler collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "cpu_scheduler_stats") + ?? McpHelpers.Status("unavailable", "No CPU scheduler data available. The scheduler collector may not have run yet."); var workerUtilizationPercent = item.MaxWorkersCount > 0 ? Math.Round(item.TotalCurrentWorkersCount * 100.0 / item.MaxWorkersCount, 2) diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCorrectionTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCorrectionTools.cs index 1e3e3b10e..72ffa013e 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCorrectionTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanCorrectionTools.cs @@ -35,29 +35,31 @@ public static async Task GetPlanCorrections( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 24.")] int hours_back = 24, - [Description("Maximum recommendation rows. Default 50.")] int limit = 50) + [Description("Maximum recommendation rows. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var tuning = await DarlingPlanCorrectionReader.GetLatestAutomaticTuningAsync(postgres, resolved.ServerId); var rows = await DarlingPlanCorrectionReader.GetPlanCorrectionsAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (tuning.Count == 0 && rows.Count == 0) { - return McpHelpers.Status("empty", - "No plan correction data collected for this server. The collector runs against SQL Server 2017+ " + - "(sys.dm_db_tuning_recommendations); a server that has never produced a row here either predates " + - "that or has no databases with Query Store on."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "plan_correction") + ?? McpHelpers.Status("empty", + "No plan correction data collected for this server. The collector runs against SQL Server 2017+ " + + "(sys.dm_db_tuning_recommendations); a server that has never produced a row here either predates " + + "that or has no databases with Query Store on."); } var recommendations = rows.Take(limit).Select(r => new diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanTools.cs index 620e67403..9467a62f0 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPlanTools.cs @@ -60,9 +60,10 @@ public static async Task AnalyzeQueryPlan( var xml = await DarlingStoredPlanReader.GetQueryStatsPlanXmlByHashAsync( postgres, resolved.ServerId, query_hash, database_name); if (string.IsNullOrEmpty(xml)) - return McpHelpers.Status( - "unavailable", - $"No stored plan found for query_hash '{query_hash}'{DbSuffix(database_name)}. The plan collector may not have captured a plan for this query, or the row has aged out of the store."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_stats") + ?? McpHelpers.Status( + "unavailable", + $"No stored plan found for query_hash '{query_hash}'{DbSuffix(database_name)}. The plan collector may not have captured a plan for this query, or the row has aged out of the store."); return McpPlanAnalysisFormatter.BuildAnalysisResult(xml, resolved.ServerName, "query_stats", query_hash); } @@ -89,9 +90,10 @@ public static async Task AnalyzeProcedurePlan( var xml = await DarlingStoredPlanReader.GetProcedurePlanXmlBySqlHandleAsync( postgres, resolved.ServerId, sql_handle); if (string.IsNullOrEmpty(xml)) - return McpHelpers.Status( - "unavailable", - $"No stored plan found for sql_handle '{sql_handle}'. The plan collector may not have captured a plan for this procedure, or the row has aged out of the store."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "procedure_stats") + ?? McpHelpers.Status( + "unavailable", + $"No stored plan found for sql_handle '{sql_handle}'. The plan collector may not have captured a plan for this procedure, or the row has aged out of the store."); return McpPlanAnalysisFormatter.BuildAnalysisResult(xml, resolved.ServerName, "procedure_stats", sql_handle); } @@ -120,9 +122,10 @@ public static async Task AnalyzeQueryStorePlan( var xml = await DarlingStoredPlanReader.GetQueryStorePlanTextAsync( postgres, resolved.ServerId, database_name, query_id, plan_id); if (string.IsNullOrEmpty(xml)) - return McpHelpers.Status( - "unavailable", - $"No stored Query Store plan found for query_id {query_id} in database '{database_name}'{PlanSuffix(plan_id)}. Query Store plan capture may be disabled for this database, or the plan has been purged."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_store") + ?? McpHelpers.Status( + "unavailable", + $"No stored Query Store plan found for query_id {query_id} in database '{database_name}'{PlanSuffix(plan_id)}. Query Store plan capture may be disabled for this database, or the plan has been purged."); var identifier = plan_id is null ? $"{database_name}:{query_id}" : $"{database_name}:{query_id}:{plan_id}"; return McpPlanAnalysisFormatter.BuildAnalysisResult(xml, resolved.ServerName, "query_store", identifier); @@ -171,7 +174,8 @@ public static async Task GetPlanXml( var xml = await DarlingStoredPlanReader.GetQueryStatsPlanXmlByHashAsync( postgres, resolved.ServerId, query_hash, database_name); if (string.IsNullOrEmpty(xml)) - return McpHelpers.Status("unavailable", $"No stored plan found for query_hash '{query_hash}'{DbSuffix(database_name)}."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_stats") + ?? McpHelpers.Status("unavailable", $"No stored plan found for query_hash '{query_hash}'{DbSuffix(database_name)}."); return McpHelpers.Truncate(xml, 512_000) ?? "No plan XML available."; } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPvsTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPvsTools.cs index 3c78cdd77..7ff673528 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPvsTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpPvsTools.cs @@ -51,9 +51,10 @@ public static async Task GetPvsStats( var rows = await DarlingPvsReader.GetPvsStatsLatestAsync(postgres, resolved.ServerId); if (rows.Count == 0) { - return McpHelpers.Status("empty", - "No PVS data collected for this server. The collector reads sys.dm_tran_persistent_version_store_stats " + - "(SQL Server 2019+); a server with no rows either predates ADR or has not completed a pvs_stats cycle yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "pvs_stats") + ?? McpHelpers.Status("empty", + "No PVS data collected for this server. The collector reads sys.dm_tran_persistent_version_store_stats " + + "(SQL Server 2019+); a server with no rows either predates ADR or has not completed a pvs_stats cycle yet."); } var databases = rows.Select(r => new diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryHeatmapTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryHeatmapTools.cs new file mode 100644 index 000000000..48db3ca5f --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryHeatmapTools.cs @@ -0,0 +1,207 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; + +#pragma warning disable CA1707 // MCP tools use snake_case naming convention + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// get_query_heatmap (#2484) — the viewer's Queries > Query Heatmap tab, the last of the ten viewer +/// surfaces that had no /api/read endpoint. +/// +/// The interactive plot is desktop-only by design (#2484 group (c)); the READ behind it is not, and a +/// bucketed table is the same answer in a shape a browser and an agent can both take. What it adds over +/// every other query read is the TIME axis: get_top_queries_by_cpu ranks a whole window and cannot +/// show that the window had a quiet half and a bad half, which is the question anyone asks first about an +/// incident that has already ended. +/// +/// A STORED read over , no live monitored-server hit. +/// +[McpServerToolType] +public sealed class DarlingMcpQueryHeatmapTools +{ + /// The web panel's cap and this tool's default: 500 cells, which is a full day of 5-minute bins + /// on a server whose queries land in two or three magnitude buckets per bin. + public const int DefaultCellLimit = 500; + + [McpServerTool(Name = "get_query_heatmap"), Description("Draws the desktop viewer's Query Heatmap as a table: how many distinct queries fell into each (time bin x log-magnitude bucket) cell over a window, plus the most-executed query in each cell. It answers when a server was slow and how slow at the same time - get_top_queries_by_cpu ranks queries over a whole window and cannot show that the window had two very different halves. Bins are 5 minutes wide by default because that is exactly what the desktop viewer uses, so a browser, an agent and a desktop pointed at the same server draw the same picture; raise bucket_minutes for a longer window, which is also the lever that fits more of the window inside the cell cap. Magnitude buckets are the viewer's seven, in the metric's own unit: under 1, 1-10, 10-100, 100-1K, 1K-10K, 10K-100K and over 100K.")] + public static async Task GetQueryHeatmap( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("How far back to look, in hours. Default 24.")] int hours_back = 24, + [Description("Which per-execution metric to bucket by: duration, cpu, logical_reads, logical_writes or execution_count. Default duration.")] string? metric = null, + [Description("Limit to one database. Omit for all databases.")] string? database_name = null, + [Description("Width of each time bin, in minutes. Default 5 - the desktop viewer's own bin width, so the two surfaces agree. Raise it to cover a longer window in fewer cells.")] int bucket_minutes = DarlingQueryHeatmapReader.ViewerBucketMinutes, + [Description("Maximum CELLS to return, most recent bins first. Default 500. A full day of 5-minute bins can reach 2,016 cells on a busy server; raise bucket_minutes rather than the cap to see the whole window.")] int limit = DefaultCellLimit, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd) + ?? McpHelpers.ValidateTop(limit) + ?? ValidateBucketMinutes(bucket_minutes); + if (validation != null) return validation; + + /* A metric we do not know is REFUSED, not quietly turned into duration: a caller who asked for CPU + and silently got elapsed time would read the wrong grid with nothing to tell them so. */ + if (!DarlingQueryHeatmapReader.TryParseMetric(metric, out var parsedMetric)) + return InvalidMetric(metric!); + + try + { + /* + #2495: the anchor, and it earns its place here more than on most reads. A heatmap IS a time + axis, so "the 4 hours ending Tuesday 03:00" is the shape of every question anyone brings to + it — and widening hours_back until an old incident falls inside is not the same question, + because the extra hours land as extra COLUMNS that push the incident's own columns past the + cell cap. + */ + var end = windowEnd; + var start = end.AddHours(-hours_back); + + /* + Over-fetch by one. Comparing the row count to the cap reports truncation for a server that + happens to have exactly `limit` cells and nothing more, which is a false positive in the one + field whose whole reason for existing is that the cap should not have to be inferred. + */ + var rows = await DarlingQueryHeatmapReader.GetQueryHeatmapAsync( + postgres, resolved.ServerId, parsedMetric, start, end, database_name, bucket_minutes, limit + 1); + + if (rows.Count == 0) + return await EmptyAsync(postgres, resolved.ServerName, resolved.ServerId, start, end, hours_back); + + var truncated = rows.Count > limit; + var cells = rows.Take(limit).ToList(); + + if (truncated) + { + /* + The rows arrive newest-bin-first so the cap keeps the RECENT end of the window, which is + what anyone looking at an incident wants. The cost is that the cap can land in the middle + of the oldest bin it reached, handing back a column missing its low buckets — and a column + with holes reads as "nothing fast ran then" rather than "we stopped looking", which is the + kind of quiet wrong answer a grid makes very easy to believe. So the partial column goes. + Kept when it is the ONLY column (a cap below one bin's seven cells has nothing to fall + back to); `truncated` and last_time_bin still say what happened. + */ + var oldestReached = cells[^1].TimeBucket; + var wholeColumns = cells.Where(c => c.TimeBucket > oldestReached).ToList(); + if (wholeColumns.Count > 0) cells = wholeColumns; + } + + /* + Back into reading order: time ascending, and buckets ascending WITHIN each bin. A plain + Reverse() would hand back the bins in the right order with each bin's buckets upside down, + because the SQL sorts time DESC and bucket ASC. The DESC exists only so the cap cuts the + right end of the window; it should not leak into the shape of the grid. + */ + cells = cells.OrderBy(c => c.TimeBucket).ThenBy(c => c.BucketIndex).ToList(); + + var labels = DarlingQueryHeatmapReader.BucketLabels[parsedMetric]; + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + metric = DarlingQueryHeatmapReader.MetricName(parsedMetric), + metric_unit = DarlingQueryHeatmapReader.MetricUnit(parsedMetric), + database_name, + /* The window that was QUERIED, anchored or not — the caller reads the bins against it. */ + window_start = start.ToString("o"), + window_end = end.ToString("o"), + bucket_minutes, + /* The same bin width the desktop viewer hardcodes, so the two surfaces cannot disagree + about the same server over the same window. */ + bucket_minutes_matches_desktop_viewer = bucket_minutes == DarlingQueryHeatmapReader.ViewerBucketMinutes, + /* A bare bucket_index is unreadable, and the labels differ by metric family: duration and + CPU are milliseconds, the other three are counts. */ + magnitude_buckets = labels.Select((label, index) => new { bucket_index = index, label }), + time_bin_count = cells.Select(c => c.TimeBucket).Distinct().Count(), + cell_count = cells.Count, + /* Which slice of [window_start, window_end] actually came back. When truncated is true the + read dropped the OLDEST bins, not the least interesting cells, so these two are the only + way to see how much of the window is missing. */ + first_time_bin = cells[0].TimeBucket.ToString("o"), + last_time_bin = cells[^1].TimeBucket.ToString("o"), + truncated, + cells = cells.Select(c => new + { + time_bucket = c.TimeBucket.ToString("o"), + bucket_index = c.BucketIndex, + bucket_label = labels[Math.Clamp(c.BucketIndex, 0, DarlingQueryHeatmapReader.BucketCount - 1)], + /* Distinct queries in the cell, NOT executions — a cell of 40 is forty different queries + that ran at that speed, which is a different finding from one query running 40 times. */ + query_count = c.QueryCount, + top_query_hash = c.TopQueryHash, + top_query_text = c.TopQueryText, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_query_heatmap", ex); + } + } + + /// + /// What an empty grid means, which is three different things. + /// The probe reads the DATA rather than counting SUCCESS runs in collection_log, and that is + /// a judgement about which kind of table this is. query_stats is PERIODIC: the collector writes + /// rows every cycle for whatever is in the plan cache, so an empty history really does mean nobody + /// looked. An edge table — blocking, deadlocks — would need the opposite treatment, because there zero + /// rows is the healthy answer and a data probe sends someone to fix collection that works. + /// The third branch is the one only this read has: collection ran, the window has captures, and + /// every one of them recorded zero executions. That is an IDLE server, not a broken one, and telling a + /// caller to widen the window there would be advice pointed at the wrong problem. + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, string serverName, int serverId, DateTime start, DateTime end, int hours_back) + { + var (hasAny, hasInWindow) = await DarlingQueryHeatmapReader.GetCoverageAsync(postgres, serverId, start, end); + + if (!hasAny) + { + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, serverId, serverName, "query_stats") + ?? McpHelpers.Status( + "unavailable", + $"No query stats have EVER been collected for {serverName}, so this is NOT a report of a quiet server — there is nothing to draw. query_stats is a PERIODIC table rather than an edge table: the collector writes rows every cycle for whatever is in the plan cache, so an empty history means nobody looked. Check get_collection_health for this server."); + } + + if (!hasInWindow) + { + return McpHelpers.Status( + "empty", + $"{serverName} has query stats from outside this window but nothing collected IN the last {hours_back} hour(s), so the grid has no columns rather than no hot cells. Widen hours_back, or check get_collection_health — a collector that stopped looks exactly like this."); + } + + return McpHelpers.Status( + "empty", + $"Query stats WERE collected for {serverName} in the last {hours_back} hour(s), but no capture recorded an execution: every row carried a zero execution delta, so nothing lands on the grid. A server that is up and idle looks exactly like this, and so does a database_name filter matching nothing collected. Delta-based collection also needs a SECOND cycle before the first non-zero row exists."); + } + + /// The bin-width bound. Refuses out of range rather than clamping, for the same reason the row + /// cap does: a silently rewritten bin width draws a different grid than the one that was asked for. + private static string? ValidateBucketMinutes(int bucket_minutes) => + bucket_minutes >= 1 && bucket_minutes <= DarlingQueryHeatmapReader.MaxBucketMinutes + ? null + : $"Invalid bucket_minutes value '{bucket_minutes}'. Must be between 1 and 1440 (one day). The desktop viewer's Query Heatmap uses 5, which is this read's default."; + + private static string InvalidMetric(string metric) => + $"Invalid metric '{metric}'. Valid values: duration, cpu, logical_reads, logical_writes, execution_count."; +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryStoreRegressionTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryStoreRegressionTools.cs new file mode 100644 index 000000000..572e16c34 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpQueryStoreRegressionTools.cs @@ -0,0 +1,160 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.ComponentModel; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using ModelContextProtocol.Server; +using Npgsql; +using PerformanceMonitor.Common; + +#pragma warning disable CA1707 // MCP tools use snake_case naming convention + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// get_query_store_regressions (#2484) - the viewer's Queries > Query Store Regressions tab, which was the +/// only tab in the per-server page entirely unreachable from a browser or an agent rather than merely +/// reduced. +/// +/// Every other Query Store read answers "what is expensive". This answers "what got WORSE", which is +/// a different question and not derivable from the first: the most expensive query on a server is usually +/// the one that has always been the most expensive, and the one that changed last Tuesday can sit well +/// down the list. A STORED read over , no live +/// monitored-server hit. +/// +[McpServerToolType] +public sealed class DarlingMcpQueryStoreRegressionTools +{ + [McpServerTool(Name = "get_query_store_regressions"), Description("Finds queries whose Query Store performance got WORSE, by comparing each (database, query_id) group's averages inside a recent window against its baseline - every capture BEFORE that window. Returns baseline vs recent duration, CPU and logical reads with the regression percent for each, the execution-count-weighted extra duration (the ranking key: a 5 ms regression executed a million times outranks a 5-second one executed twice), the plan counts on both sides, and a duration-driven severity band. get_query_store_top answers what is EXPENSIVE; the most expensive query is usually the one that always was. This answers what CHANGED. Rows are kept only where average CPU regressed by more than 25%.")] + public static async Task GetQueryStoreRegressions( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Size of the RECENT window, in hours back from now. Everything collected before it is the baseline. Default 24.")] int hours_back = 24, + [Description("Limit to one database. Omit for all databases.")] string? database_name = null, + [Description("Maximum rows to return, worst first. Default 50 (the number the desktop viewer shows).")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd) ?? McpHelpers.ValidateTop(limit); + if (validation != null) return validation; + + try + { + var end = windowEnd; + var start = end.AddHours(-hours_back); + + /* + Over-fetch by one. Comparing the row count to the cap reports truncation for a server that + happens to have exactly `limit` regressions and nothing more, which is a false positive in + the one field whose whole reason for existing is that the cap should not have to be inferred. + */ + var rows = await DarlingQueryStoreRegressionReader.GetQueryStoreRegressionsAsync( + postgres, resolved.ServerId, start, end, database_name, limit + 1); + + if (rows.Count == 0) + return await EmptyAsync(postgres, resolved.ServerName, resolved.ServerId, start, end, hours_back); + + var truncated = rows.Count > limit; + var shown = rows.Take(limit); + + return JsonSerializer.Serialize(new + { + server = resolved.ServerName, + hours_back, + database_name, + /* + Named so the caller cannot mistake which side is which. "baseline" is NOT a fixed + lookback: it is everything collected before the window, so a longer hours_back makes + the recent window bigger AND the baseline shorter. + */ + recent_window_start = start.ToString("o"), + recent_window_end = end.ToString("o"), + baseline_is = "every Query Store capture collected BEFORE recent_window_start", + gate = "average CPU regressed by more than 25%", + regression_count = Math.Min(rows.Count, limit), + truncated, + regressions = shown.Select(r => new + { + database_name = r.DatabaseName, + query_id = r.QueryId, + severity = r.Severity, + baseline_duration_ms = r.BaselineDurationMs, + recent_duration_ms = r.RecentDurationMs, + duration_regression_percent = r.DurationRegressionPercent, + baseline_cpu_ms = r.BaselineCpuMs, + recent_cpu_ms = r.RecentCpuMs, + cpu_regression_percent = r.CpuRegressionPercent, + baseline_reads = r.BaselineReads, + recent_reads = r.RecentReads, + io_regression_percent = r.IoRegressionPercent, + /* The ranking key, and the one number that says whether this regression MATTERS: a + 5 ms regression executed a million times outranks a 5-second one executed twice. */ + additional_duration_ms = r.AdditionalDurationMs, + baseline_exec_count = r.BaselineExecCount, + recent_exec_count = r.RecentExecCount, + /* A plan count that moved between the two sides is the first thing to check: a query + that regressed while gaining a plan is usually a plan-choice problem, not a data one. */ + baseline_plan_count = r.BaselinePlanCount, + recent_plan_count = r.RecentPlanCount, + last_execution_time = r.LastExecutionTime?.ToString("o"), + query_text = r.QueryTextSample, + }), + }, McpHelpers.JsonOptions); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_query_store_regressions", ex); + } + } + + /// + /// What zero regressions actually means, which is four different things. + /// Only ONE of them is good news, and the other three all look identical to it in a bare empty + /// array. The dangerous one is a server whose entire collected history sits INSIDE the requested + /// window: it has no BEFORE, so it can never show a regression however badly it regressed, and + /// answering "no regressions" there is a confident wrong answer rather than a missing one. One probe, + /// two booleans, run only on this path. + /// + private static async Task EmptyAsync( + NpgsqlDataSource postgres, string serverName, int serverId, DateTime start, DateTime end, int hours_back) + { + var (hasBaseline, hasRecent) = await DarlingQueryStoreRegressionReader.GetCoverageAsync( + postgres, serverId, start, end); + + if (!hasBaseline && !hasRecent) + { + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, serverId, serverName, "query_store") + ?? McpHelpers.Status( + "unavailable", + $"No Query Store data has EVER been collected for {serverName}, so this is NOT a report of zero regressions — there is nothing to compare. Query Store may be OFF on this server's databases, which get_query_store_health will say; otherwise check that collection is running for this server."); + } + + if (!hasBaseline) + { + return McpHelpers.Status( + "unavailable", + $"Every Query Store capture for {serverName} falls INSIDE the last {hours_back} hour(s), so there is no baseline to compare against and no regression can be detected however badly one regressed. This is NOT a clean bill of health. Shorten hours_back so more of the collected history falls before the window, or wait until this server has history older than it."); + } + + if (!hasRecent) + { + return McpHelpers.Status( + "empty", + $"{serverName} has Query Store history from before this window but nothing collected IN the last {hours_back} hour(s), so there is a recent side missing rather than nothing to report. Widen hours_back, or check get_collection_health — a collector that stopped looks exactly like this."); + } + + return McpHelpers.Status( + "empty", + $"No query on {serverName} regressed in the last {hours_back} hour(s). Both a baseline and this window were collected and no query's average CPU is more than 25% worse than its baseline — this IS the all-clear for this read."); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpSessionTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpSessionTools.cs index e00d59fa8..af2f78bc7 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpSessionTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpSessionTools.cs @@ -49,7 +49,8 @@ public static async Task GetSessionStats( { var rows = await DarlingSessionReader.GetLatestSessionStatsAsync(postgres, resolved.ServerId); if (rows.Count == 0) - return McpHelpers.Status("unavailable", "No session statistics available. The session collector may not have run yet."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "session_stats") + ?? McpHelpers.Status("unavailable", "No session statistics available. The session collector may not have run yet."); var totalConnections = rows.Sum(r => r.ConnectionCount); var totalRunning = rows.Sum(r => r.RunningCount); @@ -93,23 +94,25 @@ public static async Task GetActiveQueries( [Description("Hours of data to retrieve. Default 1.")] int hours_back = 1, [Description("Filter to a specific database.")] string? database_name = null, [Description("Show only queries involved in blocking (blocking_session_id > 0 or is a head blocker).")] bool blocking_only = false, - [Description("Maximum number of rows to return. Default 50.")] int limit = 50) + [Description("Maximum number of rows to return. Default 50.")] int limit = 50, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingSessionReader.GetActiveQueriesAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No active query snapshots found in the requested time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_snapshots") + ?? McpHelpers.Status("empty", "No active query snapshots found in the requested time range."); IEnumerable filtered = rows; @@ -166,23 +169,25 @@ public static async Task GetWaitingTasks( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of history. Default 1.")] int hours_back = 1, - [Description("Maximum rows. Default 30.")] int limit = 30) + [Description("Maximum rows. Default 30.")] int limit = 30, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateTop(limit); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var rows = await DarlingSessionReader.GetWaitingTasksAsync( postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (rows.Count == 0) - return McpHelpers.Status("empty", "No waiting tasks captured in the specified time range."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "waiting_tasks") + ?? McpHelpers.Status("empty", "No waiting tasks captured in the specified time range."); var result = rows.Take(limit).Select(r => new { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTools.cs index e4217f5bd..e4addaa31 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTools.cs @@ -36,23 +36,30 @@ namespace PerformanceMonitor.Darling.Service.Mcp; [McpServerToolType] public sealed class DarlingMcpTools { - [McpServerTool(Name = "analyze_server"), Description("Runs the diagnostic inference engine against a server's collected data. Scores wait stats, blocking, memory, config, and other facts, then traverses a relationship graph to build evidence-backed stories about what's wrong and why. Anomaly detection compares the analysis window against 30-day time-bucketed baselines (hour-of-day x day-of-week) to identify deviations that are unusual for this specific time slot, not just unusual overall. Returns structured findings with severity scores, evidence chains, baseline context for anomalies, and recommended next tools to call. A remediable finding also carries remediation_command: the full copy-paste T-SQL remediation (identical to the viewer card), including a two-sided risk-disclosure comment header on destructive changes; it is advisory only and never executed. A force-plan remediation additionally carries structured_remediation: the same decision as machine-readable fields — eligible, named blockers (parameter_sensitivity_cofired, secondary_replica_evidence), evidence numbers, and split force_sql/unforce_sql/verify_sql artifacts — so agents consume the verdict as data instead of parsing comment prose.")] + [McpServerTool(Name = "analyze_server"), Description("Runs the diagnostic inference engine against a server's collected data. Scores wait stats, blocking, memory, config, and other facts, then traverses a relationship graph to build evidence-backed stories about what's wrong and why. Anomaly detection compares the analysis window against 30-day time-bucketed baselines (hour-of-day x day-of-week) to identify deviations that are unusual for this specific time slot, not just unusual overall. Returns structured findings with severity scores, evidence chains, baseline context for anomalies, and recommended next tools to call. A remediable finding also carries remediation_command: the full copy-paste T-SQL remediation (identical to the viewer card), including a two-sided risk-disclosure comment header on destructive changes; it is advisory only and never executed. A force-plan remediation additionally carries structured_remediation: the same decision as machine-readable fields — eligible, named blockers (parameter_sensitivity_cofired, secondary_replica_evidence), evidence numbers, and split force_sql/unforce_sql/verify_sql artifacts — so agents consume the verdict as data instead of parsing comment prose. Set as_of to analyze a PAST window instead of the present — hours_back stays the window's LENGTH, and the anomaly baseline moves with it, so the findings are the ones that window deserves rather than today's findings over older rows. An anchored run is EXPLORATORY: its findings are returned in full but deliberately NOT written to the store, because a finding row is stamped with the time the analysis RAN and would then be read as this server's current state by get_analysis_findings and by the viewer. The result says so in persisted / persistence_note.")] public static async Task AnalyzeServer( DarlingAnalysisService analysisService, NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of data to analyze. Default 4. Longer windows give more stable results but may miss recent spikes.")] int hours_back = 4) + [Description("Hours of data to analyze. Default 4. Longer windows give more stable results but may miss recent spikes.")] int hours_back = 4, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; + /* Null when the caller sent no anchor, and that distinction is load-bearing here rather than + cosmetic: ValidateWindow hands back "now" for an absent as_of, and passing THAT through would + make every ordinary run look anchored to the engine — which is exactly the set of runs that + must still persist. AnalysisContext.AsOfUtc means "anchored", not "the window ends somewhere". */ + var anchor = string.IsNullOrWhiteSpace(as_of) ? (DateTime?)null : windowEnd; + try { var findings = await analysisService.AnalyzeAsync( - resolved.ServerId, resolved.ServerName, hours_back); + resolved.ServerId, resolved.ServerName, hours_back, asOfUtc: anchor); if (analysisService.InsufficientDataMessage != null) { @@ -64,6 +71,15 @@ public static async Task AnalyzeServer( }, McpHelpers.JsonOptions); } + /* #2506: whether this run's findings reached the store, and why not when they did not. + Reported rather than left to the documentation because the caller cannot otherwise tell: + an anchored run returns a complete, correct set of findings that simply does not exist in + analysis_findings, and an agent that assumed otherwise would tell someone to "check the + persisted findings" for a run that never wrote any. */ + var persistenceNote = anchor is null + ? null + : "as_of was supplied, so this analysis ran over a PAST window and is exploratory: the findings below are complete but were NOT written to the store. A finding row carries the time the analysis RAN, and the reads that consume those rows (get_analysis_findings, the viewer's Recommendations tab) treat the newest analysis_time as this server's CURRENT state — so persisting a backdated run would make last week's findings today's headline and would inflate the occurrence stats of any live incident sharing a story path. Re-run without as_of to analyze and persist the present."; + if (findings.Count == 0) { /* A successful analysis that found nothing wrong: a true negative ("all clear"), @@ -71,7 +87,12 @@ surfaced with the shared miss vocabulary so callers branch on it uniformly. */ return McpHelpers.Status( "empty", "No significant findings. All metrics are within normal ranges.", - new { analysis_time = analysisService.LastAnalysisTime?.ToString("o") }); + new + { + analysis_time = analysisService.LastAnalysisTime?.ToString("o"), + persisted = anchor is null, + persistence_note = persistenceNote + }); } // Correlate-and-focus slice 1 (review §1d): each finding's "what else fired this window". @@ -85,6 +106,10 @@ surfaced with the shared miss vocabulary so callers branch on it uniformly. */ status = "findings", finding_count = findings.Count, analysis_time = analysisService.LastAnalysisTime?.ToString("o"), + persisted = anchor is null, + /* Null on the ordinary unanchored run — nothing needs saying when the answer is the + one every caller already assumed. */ + persistence_note = persistenceNote, time_range = new { start = findings[0].TimeRangeStart?.ToString("o"), @@ -155,18 +180,23 @@ public static async Task GetAnalysisFacts( [Description("Server name or display name.")] string? server_name = null, [Description("Hours of data to analyze. Default 4.")] int hours_back = 4, [Description("Filter to a specific source category: waits, blocking, config, memory. Omit for all.")] string? source = null, - [Description("Minimum severity to include. Default 0 (all facts). Use 0.5 to see only significant facts.")] double min_severity = 0) + [Description("Minimum severity to include. Default 0 (all facts). Use 0.5 to see only significant facts.")] double min_severity = 0, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; + /* Null for an absent anchor — see analyze_server's note. Nothing here persists, so the + distinction costs nothing; it is kept so AnalysisContext.AsOfUtc means one thing everywhere. */ + var anchor = string.IsNullOrWhiteSpace(as_of) ? (DateTime?)null : windowEnd; + try { var facts = await analysisService.CollectAndScoreFactsAsync( - resolved.ServerId, resolved.ServerName, hours_back); + resolved.ServerId, resolved.ServerName, hours_back, asOfUtc: anchor); if (facts.Count == 0) { @@ -227,12 +257,13 @@ public static async Task CompareAnalysis( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours back for the comparison (recent) period. Default 4.")] int hours_back = 4, - [Description("Hours back for the baseline period start, measured from now. Default 28 (yesterday same time). The baseline period will be the same duration as the comparison period.")] int baseline_hours_back = 28) + [Description("Hours back for the baseline period start, measured from the end of the comparison window (now, or as_of). Default 28 (yesterday same time). The baseline period will be the same duration as the comparison period.")] int baseline_hours_back = 28, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; validation = McpHelpers.ValidateHoursBack(baseline_hours_back); if (validation != null) return validation; @@ -242,11 +273,13 @@ public static async Task CompareAnalysis( try { - var now = DateTime.UtcNow; - var comparisonEnd = now; - var comparisonStart = now.AddHours(-hours_back); - var baselineEnd = now.AddHours(-baseline_hours_back + hours_back); - var baselineStart = now.AddHours(-baseline_hours_back); + /* BOTH windows hang off the anchor, not just the comparison one — baseline_hours_back has + always been measured from the comparison window's end, and moving only that end would + silently change what the two windows are relative to each other. */ + var comparisonEnd = windowEnd; + var comparisonStart = windowEnd.AddHours(-hours_back); + var baselineEnd = windowEnd.AddHours(-baseline_hours_back + hours_back); + var baselineStart = windowEnd.AddHours(-baseline_hours_back); var (baselineFacts, comparisonFacts) = await analysisService.ComparePeriodsAsync( resolved.ServerId, resolved.ServerName, @@ -279,9 +312,48 @@ public static async Task CompareAnalysis( .OrderByDescending(c => Math.Abs(c.severity_delta)) .ToList(); + if (comparisons.Count == 0) + { + /* + Neither window produced a single fact, and the old payload said that with all-zero + counters and facts: [] -- which reads as "nothing changed" when it actually means + "there was nothing to compare". Those are opposite conclusions about the same server. + No probe is needed to tell them apart: comparisons is the UNION of both windows' keys, + so zero entries is exactly "both fact sets were empty" and the fact_counts already in + hand are the whole answer. + */ + return McpHelpers.Status( + "unavailable", + $"No analysis facts were collected for {resolved.ServerName} in EITHER window, so there is nothing to compare — this is NOT a report that nothing changed. Fact collection needs collected data in the window it scores; check that collection covered both periods (get_collection_log) before drawing any conclusion from this comparison.", + new + { + server = resolved.ServerName, + baseline_start = baselineStart.ToString("o"), + baseline_end = baselineEnd.ToString("o"), + comparison_start = comparisonStart.ToString("o"), + comparison_end = comparisonEnd.ToString("o"), + }); + } + + /* + One window empty and the other populated is the OTHER way this read lies, and it lies + loudly: every fact in the populated window lands in new_issues or resolved_issues purely + because it has nothing to be compared against. "47 resolved issues" on a server whose recent + window simply was not collected is a worse answer than no answer. Data-bearing results keep + their own shape rather than the status envelope, so the warning rides in the payload. + */ + var caveat = + baselineFacts.Count == 0 + ? "The BASELINE window produced no facts at all, so every fact below counts as a new issue only because there was nothing to compare it against. Confirm collection covered the baseline window (get_collection_log) before reading new_issues as a regression." + : comparisonFacts.Count == 0 + ? "The COMPARISON window produced no facts at all, so every fact below counts as a resolved issue only because there is nothing in the recent window to compare against. Confirm collection is running (get_collection_log) before reading resolved_issues as an improvement." + : null; + return JsonSerializer.Serialize(new { server = resolved.ServerName, + /* Null when both windows produced facts — the ordinary case, where nothing needs saying. */ + caveat, baseline = new { start = baselineStart.ToString("o"), @@ -536,20 +608,30 @@ public static async Task GetAnalysisFindings( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, [Description("Hours of finding history to retrieve. Default 24.")] int hours_back = 24, - [Description("If true, each finding carries drill_down: the persisted evidence rows (e.g. the parameter-sensitive plans, top spill queries) behind the chain's latest occurrence. Default false - the rows can be bulky and the summary usually suffices.")] bool include_drilldown = false) + [Description("If true, each finding carries drill_down: the persisted evidence rows (e.g. the parameter-sensitive plans, top spill queries) behind the chain's latest occurrence. Default false - the rows can be bulky and the summary usually suffices.")] bool include_drilldown = false, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; + /* Null for an absent anchor — see analyze_server's note. */ + var anchor = string.IsNullOrWhiteSpace(as_of) ? (DateTime?)null : windowEnd; + try { /* #2000: the window-covering limit, not the store default 100 — occurrence stats - computed over a silently-truncated read would lie about first_seen/occurrences. */ + computed over a silently-truncated read would lie about first_seen/occurrences. + + #2506: the window is on ANALYSIS TIME — when the scheduled pass ran — so an anchor here + asks "what was analysis saying about this server then", which is a different question + from "analyze that window now" (that is analyze_server with the same anchor). Both are + worth having: this one is the historical record and cannot change, the other recomputes + from whatever rows the store still holds. */ var findings = await analysisService.GetRecentFindingsAsync( - resolved.ServerId, hours_back, FindingOccurrences.WindowCoveringLimit); + resolved.ServerId, hours_back, FindingOccurrences.WindowCoveringLimit, asOfUtc: anchor); if (findings.Count == 0) { diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTrendTools.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTrendTools.cs index 57a7a9fa0..fd24fbbcb 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTrendTools.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpTrendTools.cs @@ -7,6 +7,7 @@ */ using System; +using System.Collections.Generic; using System.ComponentModel; using System.Linq; using System.Text.Json; @@ -21,7 +22,8 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// /// The windowed-trend data-read MCP tools — get_memory_trend / get_perfmon_trend / get_file_io_trend / -/// get_query_trend / get_query_duration_trend — served over Darling's Postgres store. These are the trend +/// get_query_trend / get_query_duration_trend / get_procedure_duration_trend / +/// get_query_store_duration_trend — served over Darling's Postgres store. These are the trend /// siblings of the merged core data-read tools (): each is a per-second / /// per-collection time-series over the respective collected table, the SAME shape a client already sees on /// Lite / the Dashboard. Every tool body mirrors LITE's Mcp*Tools trend tools field-for-field @@ -34,9 +36,10 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// windowed BOTH-sides on the naive-UTC collection_time, byte-identical to the viewer's proven /// chart reads. get_perfmon_trend reproduces Lite's miss vocabulary (the intentionally-uncollected Page /// Life Expectancy special-case + the collected-counters hint); get_query_trend reproduces Lite's per-key -/// "empty" miss; the three unkeyed trends return the #1224 "unavailable" miss on an empty window, matching -/// the merged get_tempdb_trend sibling. A response-shape change here must land in Lite's Mcp*Tools too, and -/// vice versa. +/// "empty" miss; the three unkeyed trends now distinguish the two kinds of nothing (#2485) rather than +/// collapsing both into one "unavailable" — a quiet window answers "empty" and tells the caller to widen +/// it, while a server the collector has never sampled answers "unavailable" and says so outright. A +/// response-shape change here must land in Lite's Mcp*Tools too, and vice versa. /// /// [McpServerToolType] @@ -46,20 +49,42 @@ public sealed class DarlingMcpTrendTools public static async Task GetMemoryTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var points = await DarlingTrendReader.GetMemoryTrendAsync(postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (points.Count == 0) - return McpHelpers.Status("unavailable", "No memory trend data available."); + { + /* + "No memory trend data available" was true of two opposite states and told the caller + neither. A server that collected fine and was simply quiet in THIS window wants the + window widened; a server the collector has never touched wants somebody to go look at + collection, and widening will never fill it. Probed only here, on the path that already + found nothing, against the SAME source the trend read. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "memory_stats"); + if (gated != null) + { + return gated; + } + + return await DarlingTrendReader.HasAnyMemoryStatAsync(postgres, resolved.ServerId) + ? McpHelpers.Status( + "empty", + $"No memory samples recorded for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS collected memory stats before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.") + : McpHelpers.Status( + "unavailable", + $"No memory stats have EVER been recorded for {resolved.ServerName}. This is not an empty window — the memory_stats collector has stored nothing at all for this server. Check that collection is running and that the server is enabled; get_memory_stats will be equally empty until it does."); + } var result = points.Select(p => new { @@ -92,21 +117,32 @@ public static async Task GetPerfmonTrend( NpgsqlDataSource postgres, [Description("The exact counter name, e.g. 'Batch Requests/sec'.")] string counter_name, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var start = now.AddHours(-hours_back); var points = await DarlingTrendReader.GetPerfmonTrendAsync(postgres, resolved.ServerId, counter_name, start, now); if (points.Count == 0) { + /* The engine question comes BEFORE the distinct-counter probe, not after it. Both are on + the miss path, so either order keeps the property that matters — but a permanently gated + engine takes this branch on every call, forever, and neither the probe nor the PLE branch + below could tell it anything. Asking first makes that case one query instead of two. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "perfmon_stats"); + if (gated != null) + { + return gated; + } + /* No points can mean three different things to a caller. Distinguish them so an LLM doesn't read a bad counter name as "this metric looks fine" — Lite's get_perfmon_trend miss vocabulary. */ @@ -169,20 +205,39 @@ by the configured 60 s is wrong by whatever the jitter was (#2233, #2234). */ public static async Task GetFileIoTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var points = await DarlingTrendReader.GetFileIoLatencyTrendAsync(postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (points.Count == 0) - return McpHelpers.Status("unavailable", "No I/O trend data available."); + { + /* Same two states as the memory trend, same probe discipline. The quiet-window sentence + carries one extra clause the others do not need: this read's top_files CTE requires + delta_reads or delta_writes above zero, so a genuinely idle file set is empty here even + on a server whose file_io_stats collector ran every cycle. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "file_io_stats"); + if (gated != null) + { + return gated; + } + + return await DarlingTrendReader.HasAnyFileIoStatAsync(postgres, resolved.ServerId) + ? McpHelpers.Status( + "empty", + $"No file I/O samples recorded for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS collected file I/O stats before, so this window is genuinely quiet rather than broken — widen hours_back, or read it as no measurable read or write activity on any file in this window.") + : McpHelpers.Status( + "unavailable", + $"No file I/O stats have EVER been recorded for {resolved.ServerName}. This is not an empty window — the file_io_stats collector has stored nothing at all for this server. Check that collection is running and that the server is enabled; get_file_io_stats will be equally empty until it does."); + } var result = points.Select(p => new { @@ -211,17 +266,18 @@ public static async Task GetQueryTrend( [Description("The query_hash value from get_top_queries_by_cpu or get_query_store_top.")] string query_hash, [Description("The database name the query belongs to.")] string database_name, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var history = await DarlingTrendReader.GetQueryHistoryAsync(postgres, resolved.ServerId, database_name, query_hash, now.AddHours(-hours_back), now); var rows = history.Points; if (rows.Count == 0) @@ -230,12 +286,13 @@ public static async Task GetQueryTrend( last N hours" over a span the read never covered — for a query whose history had aged out of the raw tier that is a false statement, not an incomplete one, and an agent acts on it by concluding the query did not run. */ - return McpHelpers.Status( - "empty", - $"No history found for query_hash '{query_hash}' in database '{database_name}' in the " + - $"{history.Source} tier over the last {hours_back} hours. This means nothing was recorded " + - "for that query_hash in that window in the tier searched — confirm the hash and database " + - "with get_top_queries_by_cpu before concluding the query did not run."); + return await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_stats") + ?? McpHelpers.Status( + "empty", + $"No history found for query_hash '{query_hash}' in database '{database_name}' in the " + + $"{history.Source} tier over the last {hours_back} hours. This means nothing was recorded " + + "for that query_hash in that window in the tier searched — confirm the hash and database " + + "with get_top_queries_by_cpu before concluding the query did not run."); } /* The hourly rollup keeps executions, CPU and elapsed and nothing else. Those columns arrive as @@ -294,41 +351,181 @@ concluding the query did not run. */ public static async Task GetQueryDurationTrend( NpgsqlDataSource postgres, [Description("Server name or display name.")] string? server_name = null, - [Description("Hours of history. Default 24.")] int hours_back = 24) + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) { var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); if (error != null) return error; - var validation = McpHelpers.ValidateHoursBack(hours_back); + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); if (validation != null) return validation; try { - var now = DateTime.UtcNow; + var now = windowEnd; var points = await DarlingTrendReader.GetQueryDurationTrendAsync(postgres, resolved.ServerId, now.AddHours(-hours_back), now); if (points.Count == 0) - return McpHelpers.Status("unavailable", "No query duration trend data available."); + { + /* Same two states again. The probe reads the BASE query_stats table because this trend + does — v_query_stats is the payload-resolving view on a V38+ store, and probing a + different relation from the one the read walks is how an existence probe ends up + reporting the wrong branch. */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_stats"); + if (gated != null) + { + return gated; + } + + return await DarlingTrendReader.HasAnyQueryStatAsync(postgres, resolved.ServerId) + ? McpHelpers.Status( + "empty", + $"No query samples recorded for {resolved.ServerName} in the last {hours_back} hour(s). This server HAS collected query stats before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.") + : McpHelpers.Status( + "unavailable", + $"No query stats have EVER been recorded for {resolved.ServerName}. This is not an empty window — the query_stats collector has stored nothing at all for this server. Check that collection is running and that the server is enabled; get_top_queries_by_cpu will be equally empty until it does."); + } - var result = points.Select(p => new + /* The two siblings below serialize through the SAME helper, so the three Performance-Trends + reads cannot advertise three different field sets for one shape. */ + return SerializeTrend(resolved.ServerName, hours_back, points); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_query_duration_trend", ex); + } + } + + [McpServerTool(Name = "get_procedure_duration_trend"), Description("Gets a time-series of stored-procedure elapsed time per second and executions per second over time, summed across every procedure. The sibling of get_query_duration_trend, and NOT a duplicate of it: query_stats attributes a procedure's work to the individual statements inside it, so a procedure that got slower is smeared across however many statements it runs. This charges the whole call to the procedure. Read the two together to tell an ad-hoc SQL regression from a procedure regression.")] + public static async Task GetProcedureDurationTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var now = windowEnd; + var points = await DarlingTrendReader.GetProcedureDurationTrendAsync( + postgres, resolved.ServerId, now.AddHours(-hours_back), now); + + if (points.Count == 0) { - time = p.CollectionTime.ToString("o"), - value = p.Value, - execution_count = p.ExecutionCount - }); + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "procedure_stats"); + if (gated != null) + { + return gated; + } + + return await EmptyTrendAsync( + DarlingTrendReader.HasAnyProcedureStatAsync(postgres, resolved.ServerId), + resolved.ServerName, hours_back, "stored-procedure", + "Check that collection is running and that the server is enabled. A server that genuinely runs no stored procedures also lands here, and that is a real answer rather than a fault."); + } - return JsonSerializer.Serialize(new + return SerializeTrend(resolved.ServerName, hours_back, points); + } + catch (Exception ex) + { + return McpHelpers.FormatError("get_procedure_duration_trend", ex); + } + } + + [McpServerTool(Name = "get_query_store_duration_trend"), Description("Gets a time-series of Query Store duration per second and executions per second over time, summed across every query. Where get_query_duration_trend reads the plan cache and loses everything an eviction or a restart takes with it, this reads Query Store, which persists per interval - so it is the series that survives a failover and the one to reach for when a regression is older than the cache. Each interval is counted once, at the hour the work ran.")] + public static async Task GetQueryStoreDurationTrend( + NpgsqlDataSource postgres, + [Description("Server name or display name.")] string? server_name = null, + [Description("Hours of history. Default 24.")] int hours_back = 24, + [Description(McpHelpers.AsOfDescription)] string? as_of = null) + { + var (resolved, error) = await DarlingServerResolver.ResolveOrErrorAsync(postgres, server_name); + if (error != null) return error; + + var validation = McpHelpers.ValidateWindow(hours_back, as_of, out var windowEnd); + if (validation != null) return validation; + + try + { + var now = windowEnd; + var points = await DarlingTrendReader.GetQueryStoreDurationTrendAsync( + postgres, resolved.ServerId, now.AddHours(-hours_back), now); + + if (points.Count == 0) { - server = resolved.ServerName, - hours_back, - trend = result - }, McpHelpers.JsonOptions); + /* + The one empty answer here that is NOT about the collector: Query Store can be off on + every database on the instance. A server with no Query Store data is not a server with + no slow queries, so the message names that cause first. + */ + var gated = await DarlingEngineCapability.NotCollectedStatusAsync(postgres, resolved.ServerId, resolved.ServerName, "query_store"); + if (gated != null) + { + return gated; + } + + return await EmptyTrendAsync( + DarlingTrendReader.HasAnyQueryStoreStatAsync(postgres, resolved.ServerId), + resolved.ServerName, hours_back, "Query Store", + "Query Store may be OFF on this server's databases — that, not an absence of slow queries, is the usual cause. Check QUERY_STORE = ON per database, then that collection is running for this server."); + } + + return SerializeTrend(resolved.ServerName, hours_back, points); } catch (Exception ex) { - return McpHelpers.FormatError("get_query_duration_trend", ex); + return McpHelpers.FormatError("get_query_store_duration_trend", ex); } } + /// + /// The one payload shape the three Performance-Trends siblings share, so a caller can chart them on one + /// axis without learning three field names. + /// execution_count and executions_per_second are the SAME quantity. The first + /// shipped truncated to an integer, which on a quiet server turns 0.4 executions a second into a + /// reported ZERO - an idle server, when the truth was a slow one. It is kept so a consumer reading it + /// does not break; read executions_per_second. + /// + private static string SerializeTrend( + string serverName, int hours_back, List points) => + JsonSerializer.Serialize(new + { + server = serverName, + hours_back, + trend = points.Select(p => new + { + time = p.CollectionTime.ToString("o"), + value = p.Value, + execution_count = p.ExecutionCount, + executions_per_second = p.ExecutionsPerSecond, + }), + }, McpHelpers.JsonOptions); + + /// + /// The two-branch empty answer the two new Performance-Trends siblings share (#2484), phrased to match + /// the one get_query_duration_trend already ships (#2485) so the three reads tell one story. + /// Zero points is two facts wanting opposite responses. A server that HAS been sampled and was + /// quiet in this window wants the window widened; a server that has never been sampled wants somebody + /// to go look at why. Both are literally "no trend data", and the probe — one LIMIT 1 against the same + /// table the trend reads, run only on this path — is what separates them. + /// + private static async Task EmptyTrendAsync( + Task probe, string serverName, int hours_back, string what, string checkThis) + { + var everSampled = await probe; + return everSampled + ? McpHelpers.Status( + "empty", + $"No {what} samples were recorded for {serverName} in the last {hours_back} hour(s). This server HAS been sampled before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.") + : McpHelpers.Status( + "unavailable", + $"No {what} samples have EVER been recorded for {serverName}. This is not an empty window — nothing at all has been stored for this server, so it is NOT a quiet server. {checkThis}"); + } + /// /// True when the caller asked for Page Life Expectancy by any common spelling. Matches the full /// counter name (case-insensitive) or an exact "PLE" — but not "PLE" as a substring, so counters diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingObjectStatsReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingObjectStatsReader.cs index 9ccbb2545..6780abfbc 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingObjectStatsReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingObjectStatsReader.cs @@ -167,7 +167,15 @@ public static async Task> GetObjectSizeGrowthAsync( /// /// Per-index usage at the latest snapshot — Lite's GetIndexUsageAsync ported to Postgres: /// seeks/scans/lookups/updates with the Unused / Write-only / Active classification, unused-first then - /// largest reserved. Counters are cumulative since the last restart. $1 server_id, $2 cap. + /// largest reserved. Counters are cumulative since the last restart. + /// $1 server_id, $2 database filter (NULL = every database), $3 cap. + /// #2636: the ordering and the cap interact badly and a field report found the sharp edge. Unused + /// sorts ahead of everything SERVER-WIDE, so on an instance with 200+ unused indexes concentrated in one + /// legacy database, the entire capped result is consumed by that database and every Active index in every + /// other database is invisible — with nothing in the answer to say so. The reporter's database had + /// healthy collection, full retention and zero returned rows, which reads exactly like a collection + /// failure. The database filter is what makes the question answerable; the count below is what stops the + /// answer being read as complete. /// public const string IndexUsageSql = """ SELECT @@ -194,18 +202,40 @@ END AS classification FROM v_index_object_stats WHERE server_id = $1 AND collection_time = (SELECT MAX(collection_time) FROM v_index_object_stats WHERE server_id = $1) + AND ($2::text IS NULL OR database_name = $2::text) ORDER BY CASE WHEN COALESCE(user_seeks, 0) + COALESCE(user_scans, 0) + COALESCE(user_lookups, 0) = 0 THEN 0 ELSE 1 END, reserved_mb DESC - LIMIT $2 + LIMIT $3 """; + /// + /// How many rows the same filter MATCHES, before the cap (#2636). Read alongside the rows so the answer + /// can say it was truncated and by how much — the reporter's complaint was not that a cap exists, it was + /// that nothing distinguished "not returned" from "not collected". + /// A second query rather than a window function over the first: the count has to be of the whole + /// match, and a COUNT(*) OVER () inside a LIMITed statement returns the count of what survived the + /// LIMIT — which is the exact mistake this is here to report. + /// + public const string IndexUsageMatchCountSql = """ + SELECT count(*) + FROM v_index_object_stats + WHERE server_id = $1 + AND collection_time = (SELECT MAX(collection_time) FROM v_index_object_stats WHERE server_id = $1) + AND ($2::text IS NULL OR database_name = $2::text) + """; + + /// + /// The rows the cap allowed. null means every database — the shape the + /// tool had before #2636, kept so callers that genuinely want a server-wide sweep still get one. + /// public static async Task> GetIndexUsageAsync( - NpgsqlDataSource postgres, int serverId, int top, CancellationToken cancellationToken = default) + NpgsqlDataSource postgres, int serverId, int top, string? databaseName = null, CancellationToken cancellationToken = default) { var rows = new List(); await using var command = postgres.CreateCommand(IndexUsageSql); DarlingMcpReadParameters.AddInt(command, serverId); + command.Parameters.AddWithValue((object?)databaseName ?? DBNull.Value); DarlingMcpReadParameters.AddInt(command, top); await using var reader = await command.ExecuteReaderAsync(cancellationToken); while (await reader.ReadAsync(cancellationToken)) @@ -230,6 +260,20 @@ public static async Task> GetIndexUsageAsync( return rows; } + /// + /// How many index rows the same server and database filter match at the latest snapshot, ignoring the + /// cap (#2636). + /// + public static async Task GetIndexUsageMatchCountAsync( + NpgsqlDataSource postgres, int serverId, string? databaseName = null, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(IndexUsageMatchCountSql); + DarlingMcpReadParameters.AddInt(command, serverId); + command.Parameters.AddWithValue((object?)databaseName ?? DBNull.Value); + + return await command.ExecuteScalarAsync(cancellationToken) is long count ? count : 0; + } + /* ─────────────────────────── object locking ─────────────────────────── */ /// diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs index 8c7fa48fa..c99080835 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs @@ -234,7 +234,7 @@ internal static string ResolutionMissDisclosure(Snapshot snapshot, string? reque if (matching.Count > 0) { /* Two peers can legitimately both claim a name (overlapping `matches`, e.g. "use1" and - "prod-pos"), so the follow-on sentence agrees in number rather than saying "That is a SEPARATE + "prod-sql"), so the follow-on sentence agrees in number rather than saying "That is a SEPARATE store" about a list of two. */ var single = matching.Count == 1; text.Append(subject) diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryHeatmapReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryHeatmapReader.cs new file mode 100644 index 000000000..a7eae9c98 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryHeatmapReader.cs @@ -0,0 +1,292 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using NpgsqlTypes; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// Which per-execution metric the Query Heatmap buckets rows by — the viewer's +/// HeatmapMetric, which is itself Lite's. Copied rather than referenced because the headless +/// service does not (and should not) reference the WPF viewer. +public enum HeatmapMetric +{ + Duration, + Cpu, + LogicalReads, + LogicalWrites, + ExecutionCount, +} + +/// +/// The service-side read behind get_query_heatmap (#2484) — the viewer's +/// ViewerDataService.BuildQueryHeatmapSql, which is itself the Postgres port of Lite's DuckDB +/// heatmap. Copied VERBATIM apart from the two things a desktop chart does not need and an MCP read does: +/// the bin width is a bound parameter instead of the literal INTERVAL '5 minutes' (defaulting to +/// that same 5), and the tail carries ORDER BY time_bin DESC + LIMIT so a capped call keeps +/// the most RECENT bins rather than the oldest ones. Every other clause — the magnitude CASE, the +/// delta_execution_count > 0 and metric IS NOT NULL filters, the LEFT(query_text, 120) +/// preview, the top-1 window that replaced DuckDB's ARG_MAX — is the viewer's. +/// +/// The bucketing is the viewer's, deliberately. The whole point of the web/MCP surface is that +/// it answers the same question the desktop does, so a browser and a desktop pointed at the same server over +/// the same window must not draw different pictures. The viewer's bucketing turns out to be a CONSTANT +/// (5 minutes) rather than something derived from the window length, so there is nothing to reproduce — just +/// a default to honor. It is exposed as bucket_minutes because an agent asking about a 7-day window +/// wants coarser columns, not 2,016 of them, and because widening the bin is the one lever that covers more +/// window inside the row cap. +/// +/// The origin is load-bearing once the width is a parameter. Postgres date_bin takes an +/// explicit origin and this read passes the Unix epoch, matching the viewer. DuckDB's time_bucket +/// defaults to a DIFFERENT origin (2000-01-03), which is invisible at 5 minutes — the two origins are +/// 15,780,960 minutes apart, and 5 divides that exactly — and visible at 7, where the two SKUs bin the same +/// row one minute apart. Lite's twin therefore passes the epoch explicitly too. Verified across +/// {1, 5, 7, 13, 60, 90, 360, 1440}-minute strides on PostgreSQL 17 and DuckDB 1.5.5: identical bins. +/// +/// A STORED read — no live monitored-server hit. query_stats is a PERIODIC table, not an edge +/// table: the collector writes rows every cycle for whatever is in the plan cache, so "no rows" here means +/// nobody looked, and an existence probe on the data is the right denominator (see +/// ). +/// +/// The SQL is built by a public method so the tests can pin the dialect and the shape without a live +/// Postgres. +/// +internal static class DarlingQueryHeatmapReader +{ + /// The desktop viewer's bin width, and therefore this read's default. Not derived from the + /// window length — the viewer hardcodes INTERVAL '5 minutes' whatever range is on screen. + public const int ViewerBucketMinutes = 5; + + /// The widest bin a caller may ask for: one day. Past this the "heatmap" is one column. + public const int MaxBucketMinutes = 1440; + + /// The seven log-magnitude rows of the grid — the viewer's, not a new banding. + public const int BucketCount = 7; + + /// One heatmap cell: the query count in a (time bin x magnitude bucket) plus the most-executed + /// query in it, which is what the desktop shows on hover. + public sealed record HeatmapCellRow( + DateTime TimeBucket, + int BucketIndex, + long QueryCount, + string TopQueryHash, + string TopQueryText); + + /// The viewer's per-metric magnitude labels, verbatim. Duration and CPU are milliseconds per + /// execution; the rest are plain counts, so the two families label the same seven buckets differently. + /// Returned with every result — a bare bucket_index is unreadable without them. + public static readonly IReadOnlyDictionary BucketLabels = + new Dictionary + { + [HeatmapMetric.Duration] = new[] { "0-1ms", "1-10ms", "10-100ms", "100ms-1s", "1-10s", "10-100s", ">100s" }, + [HeatmapMetric.Cpu] = new[] { "0-1ms", "1-10ms", "10-100ms", "100ms-1s", "1-10s", "10-100s", ">100s" }, + [HeatmapMetric.LogicalReads] = new[] { "0-1", "1-10", "10-100", "100-1K", "1K-10K", "10K-100K", ">100K" }, + [HeatmapMetric.LogicalWrites] = new[] { "0-1", "1-10", "10-100", "100-1K", "1K-10K", "10K-100K", ">100K" }, + [HeatmapMetric.ExecutionCount] = new[] { "0-1", "1-10", "10-100", "100-1K", "1K-10K", "10K-100K", ">100K" }, + }; + + /// + /// The per-execution metric expression, byte-identical with the viewer's HeatmapMetricExpr and + /// Lite's GetMetricColumn — the / 1000.0 microsecond-to-millisecond scaling, the + /// NULLIF(delta_execution_count, 0) per-execution average, the CAST(... AS DOUBLE PRECISION). + /// Internal constants only (no caller text reaches this), so string-composing it into the SQL is + /// injection-safe; the caller's metric string is mapped through first + /// and never appears in the query. + /// + public static string MetricExpression(HeatmapMetric metric) => metric switch + { + HeatmapMetric.Duration => "(delta_elapsed_time / 1000.0) / NULLIF(delta_execution_count, 0)", + HeatmapMetric.Cpu => "(delta_worker_time / 1000.0) / NULLIF(delta_execution_count, 0)", + HeatmapMetric.LogicalReads => "CAST(delta_logical_reads AS DOUBLE PRECISION) / NULLIF(delta_execution_count, 0)", + HeatmapMetric.LogicalWrites => "CAST(delta_logical_writes AS DOUBLE PRECISION) / NULLIF(delta_execution_count, 0)", + HeatmapMetric.ExecutionCount => "CAST(delta_execution_count AS DOUBLE PRECISION)", + _ => "(delta_elapsed_time / 1000.0) / NULLIF(delta_execution_count, 0)", + }; + + /// The snake_case name a caller passes, and the one echoed back in the result. + public static string MetricName(HeatmapMetric metric) => metric switch + { + HeatmapMetric.Duration => "duration", + HeatmapMetric.Cpu => "cpu", + HeatmapMetric.LogicalReads => "logical_reads", + HeatmapMetric.LogicalWrites => "logical_writes", + HeatmapMetric.ExecutionCount => "execution_count", + _ => "duration", + }; + + /// What one cell's magnitude actually measures. Without it a caller reading "1-10" cannot tell + /// a per-execution average from a per-interval total, and execution_count is the one metric that is a + /// TOTAL rather than an average. + public static string MetricUnit(HeatmapMetric metric) => metric switch + { + HeatmapMetric.Duration => "milliseconds of elapsed time per execution", + HeatmapMetric.Cpu => "milliseconds of CPU per execution", + HeatmapMetric.LogicalReads => "logical reads per execution", + HeatmapMetric.LogicalWrites => "logical writes per execution", + HeatmapMetric.ExecutionCount => "executions in the collection interval (a total, not a per-execution average)", + _ => "milliseconds of elapsed time per execution", + }; + + /// Maps the caller's metric string onto the enum. Returns false rather than silently + /// falling back to duration: a caller who asked for CPU and got duration would read the wrong grid with + /// no sign anything went wrong. Lite's twin accepts exactly these five names. + public static bool TryParseMetric(string? metric, out HeatmapMetric parsed) + { + parsed = HeatmapMetric.Duration; + if (string.IsNullOrWhiteSpace(metric)) return true; + + switch (metric.Trim().ToLowerInvariant()) + { + case "duration": parsed = HeatmapMetric.Duration; return true; + case "cpu": parsed = HeatmapMetric.Cpu; return true; + case "logical_reads": parsed = HeatmapMetric.LogicalReads; return true; + case "logical_writes": parsed = HeatmapMetric.LogicalWrites; return true; + case "execution_count": parsed = HeatmapMetric.ExecutionCount; return true; + default: return false; + } + } + + /// + /// The viewer's heatmap read for one metric. $1 server_id, $2 window start, $3 window end, $4 database + /// filter (text[] or NULL), $5 bin width in minutes, $6 cell cap. + /// Ordered newest bin first ONLY so the cap keeps the recent end of the window; the tool re-sorts + /// chronologically before returning. The viewer needs no cap and orders ascending. + /// + public static string BuildQueryHeatmapSql(HeatmapMetric metric) + { + var metricExpr = MetricExpression(metric); + return $""" + WITH base AS ( + SELECT + date_bin(($5::integer * INTERVAL '1 minute'), collection_time, TIMESTAMP '1970-01-01 00:00:00') AS time_bin, + {metricExpr} AS metric_value, + query_hash, + LEFT(query_text, 120) AS query_preview, + delta_execution_count + FROM v_query_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND ($4::text[] IS NULL OR database_name = ANY($4)) + AND delta_execution_count > 0 + AND {metricExpr} IS NOT NULL + ), + binned AS ( + SELECT + time_bin, + CASE + WHEN metric_value < 1 THEN 0 + WHEN metric_value < 10 THEN 1 + WHEN metric_value < 100 THEN 2 + WHEN metric_value < 1000 THEN 3 + WHEN metric_value < 10000 THEN 4 + WHEN metric_value < 100000 THEN 5 + ELSE 6 + END AS bucket_index, + query_hash, + query_preview, + delta_execution_count + FROM base + ), + ranked AS ( + SELECT + time_bin, + bucket_index, + query_hash, + query_preview, + COUNT(*) OVER (PARTITION BY time_bin, bucket_index) AS query_count, + ROW_NUMBER() OVER (PARTITION BY time_bin, bucket_index ORDER BY delta_execution_count DESC) AS rn + FROM binned + ) + SELECT + time_bin, + bucket_index, + query_count, + query_hash AS top_query_hash, + query_preview AS top_query_text + FROM ranked + WHERE rn = 1 + ORDER BY time_bin DESC, bucket_index + LIMIT $6 + """; + } + + /// + /// Whether this server has query stats AT ALL, and whether it has any inside the window. + /// One round trip for the two facts that decide what an empty grid means, run only on the empty + /// path. This probes the DATA rather than SUCCESS rows in collection_log, and that is a judgement + /// about which kind of table this is: query_stats is PERIODIC, not an edge table. The collector + /// writes a row every cycle for whatever sits in the plan cache, so a server with zero rows in its whole + /// history is a server nobody collected — unlike blocking or deadlocks, where zero rows is the healthy + /// answer and a data probe would send someone to fix collection that works. + /// Probes v_query_stats, the same relation the read itself uses, so the probe cannot + /// disagree with the read about which rows exist. $1 server_id, $2 window start, $3 window end. + /// + public const string HeatmapCoverageSql = """ + SELECT + EXISTS ( + SELECT 1 + FROM v_query_stats + WHERE server_id = $1 + ) AS has_any, + EXISTS ( + SELECT 1 + FROM v_query_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + ) AS has_in_window + """; + + /// Runs . Rows come back newest bin first. + public static async Task> GetQueryHeatmapAsync( + NpgsqlDataSource postgres, int serverId, HeatmapMetric metric, DateTime startUtc, DateTime endUtc, + string? databaseName, int bucketMinutes, int limit, CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(BuildQueryHeatmapSql(metric)); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + command.Parameters.Add(new NpgsqlParameter + { + NpgsqlDbType = NpgsqlDbType.Array | NpgsqlDbType.Text, + Value = string.IsNullOrWhiteSpace(databaseName) ? DBNull.Value : new[] { databaseName }, + }); + command.Parameters.Add(new NpgsqlParameter { TypedValue = bucketMinutes }); + command.Parameters.Add(new NpgsqlParameter { TypedValue = limit }); + + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new HeatmapCellRow( + reader.GetDateTime(0), + reader.IsDBNull(1) ? 0 : Convert.ToInt32(reader.GetValue(1)), + reader.IsDBNull(2) ? 0 : Convert.ToInt64(reader.GetValue(2)), + reader.IsDBNull(3) ? "" : reader.GetString(3), + reader.IsDBNull(4) ? "" : reader.GetString(4))); + } + + return rows; + } + + /// Runs . + public static async Task<(bool HasAny, bool HasInWindow)> GetCoverageAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HeatmapCoverageSql); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + if (!await reader.ReadAsync(cancellationToken)) + return (false, false); + return (reader.GetBoolean(0), reader.GetBoolean(1)); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryStoreRegressionReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryStoreRegressionReader.cs new file mode 100644 index 000000000..e929f27a7 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingQueryStoreRegressionReader.cs @@ -0,0 +1,267 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using NpgsqlTypes; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The service-side read behind get_query_store_regressions (#2484) - the viewer's +/// ViewerDataService.QueryStoreRegressionsSql, which is itself the Postgres port of the Dashboard's +/// report.query_store_regressions inline TVF. Copied VERBATIM apart from the row cap, which the +/// viewer hardcodes at the TVF's TOP (50) and this binds as a parameter so a caller can ask for fewer or +/// more; the gate and the ranking are untouched, so the first 50 rows of any call are the viewer's 50. +/// +/// A STORED read - no live monitored-server hit. Two windowed passes over the SAME +/// query_store_stats table: BASELINE is every capture BEFORE the window start, RECENT is the window +/// itself. Both are DEDUPED first, and that is correctness rather than performance: Query Store rows are +/// CUMULATIVE per-interval snapshots and the collector re-fetches an open interval every cycle, so the same +/// interval is stored repeatedly with a growing execution_count. This read is the most exposed of any to +/// that, because the baseline arm is UNBOUNDED (potentially months) while the recent arm is a short window: +/// the two arms have systematically different re-collection density per interval, which alone moves the +/// averages the regression percent is computed from and the 25% CPU gate - manufacturing and hiding +/// regressions for reasons that have nothing to do with the query. +/// +/// The SQL is a public const so the tests can pin the dialect and the shape without a live Postgres. +/// +internal static class DarlingQueryStoreRegressionReader +{ + /// One regression row - the viewer's ViewerQueryStoreRegressionRow, without the + /// display-formatting members. Durations and CPU are ms (converted from the stored microseconds); + /// reads are raw pages; the percents are plain deltas. + public sealed record RegressionRow( + string DatabaseName, + long QueryId, + double BaselineDurationMs, + double RecentDurationMs, + double DurationRegressionPercent, + double BaselineCpuMs, + double RecentCpuMs, + double CpuRegressionPercent, + double BaselineReads, + double RecentReads, + double IoRegressionPercent, + double AdditionalDurationMs, + long BaselineExecCount, + long RecentExecCount, + int BaselinePlanCount, + int RecentPlanCount, + string Severity, + string QueryTextSample, + DateTime? LastExecutionTime); + + /// + /// The viewer's regression read. $1 server_id, $2 window start (the baseline is everything < $2), + /// $3 window end, $4 database filter (text[] or NULL), $5 row cap. + /// + public const string QueryStoreRegressionsSql = """ + WITH deduped_baseline AS ( + /* LOAD-BEARING (correctness, not just perf) — #1841. The rows are CUMULATIVE per-interval + snapshots and the collector re-fetches the OPEN interval every cycle, so the SAME interval + (same first_execution_time) is stored repeatedly with a growing execution_count. Keep the + LATEST snapshot per interval before aggregating. */ + SELECT + database_name, + query_id, + plan_id, + execution_count, + avg_duration_us, + avg_cpu_time_us, + avg_logical_io_reads, + ROW_NUMBER() OVER + ( + PARTITION BY database_name, query_id, plan_id, runtime_stats_interval_id, first_execution_time, execution_type_desc, replica_role + ORDER BY collection_time DESC, execution_count DESC + ) AS rn + FROM query_store_stats + WHERE server_id = $1 + AND collection_time < $2 + AND ($4::text[] IS NULL OR database_name = ANY($4)) + ), + deduped_recent AS ( + SELECT + database_name, + query_id, + plan_id, + query_text, + execution_count, + avg_duration_us, + avg_cpu_time_us, + avg_logical_io_reads, + last_execution_time, + ROW_NUMBER() OVER + ( + PARTITION BY database_name, query_id, plan_id, runtime_stats_interval_id, first_execution_time, execution_type_desc, replica_role + ORDER BY collection_time DESC, execution_count DESC + ) AS rn + FROM query_store_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND ($4::text[] IS NULL OR database_name = ANY($4)) + ), + baseline_performance AS ( + SELECT + database_name, + query_id, + AVG(CAST(avg_duration_us AS double precision)) / 1000.0 AS avg_duration_ms, + AVG(CAST(avg_cpu_time_us AS double precision)) / 1000.0 AS avg_cpu_time_ms, + AVG(CAST(avg_logical_io_reads AS double precision)) AS avg_logical_io_reads, + CAST(SUM(execution_count) AS bigint) AS exec_count, + CAST(COUNT(DISTINCT plan_id) AS integer) AS plan_count + FROM deduped_baseline + WHERE rn = 1 + GROUP BY database_name, query_id + ), + recent_performance AS ( + SELECT + database_name, + query_id, + MAX(query_text) AS query_text_sample, + AVG(CAST(avg_duration_us AS double precision)) / 1000.0 AS avg_duration_ms, + AVG(CAST(avg_cpu_time_us AS double precision)) / 1000.0 AS avg_cpu_time_ms, + AVG(CAST(avg_logical_io_reads AS double precision)) AS avg_logical_io_reads, + CAST(SUM(execution_count) AS bigint) AS exec_count, + CAST(COUNT(DISTINCT plan_id) AS integer) AS plan_count, + MAX(last_execution_time) AS last_execution_time + FROM deduped_recent + WHERE rn = 1 + GROUP BY database_name, query_id + ) + SELECT + r.database_name, + r.query_id, + b.avg_duration_ms AS baseline_duration_ms, + r.avg_duration_ms AS recent_duration_ms, + (r.avg_duration_ms - b.avg_duration_ms) * 100.0 / NULLIF(b.avg_duration_ms, 0) AS duration_regression_percent, + b.avg_cpu_time_ms AS baseline_cpu_ms, + r.avg_cpu_time_ms AS recent_cpu_ms, + (r.avg_cpu_time_ms - b.avg_cpu_time_ms) * 100.0 / NULLIF(b.avg_cpu_time_ms, 0) AS cpu_regression_percent, + b.avg_logical_io_reads AS baseline_reads, + r.avg_logical_io_reads AS recent_reads, + (r.avg_logical_io_reads - b.avg_logical_io_reads) * 100.0 / NULLIF(b.avg_logical_io_reads, 0) AS io_regression_percent, + (r.avg_duration_ms - b.avg_duration_ms) * r.exec_count AS additional_duration_ms, + b.exec_count AS baseline_exec_count, + r.exec_count AS recent_exec_count, + b.plan_count AS baseline_plan_count, + r.plan_count AS recent_plan_count, + CASE + WHEN (r.avg_duration_ms - b.avg_duration_ms) * 100.0 / NULLIF(b.avg_duration_ms, 0) > 100 THEN 'CRITICAL' + WHEN (r.avg_duration_ms - b.avg_duration_ms) * 100.0 / NULLIF(b.avg_duration_ms, 0) > 50 THEN 'HIGH' + WHEN (r.avg_duration_ms - b.avg_duration_ms) * 100.0 / NULLIF(b.avg_duration_ms, 0) > 25 THEN 'MEDIUM' + ELSE 'LOW' + END AS severity, + /* #2150: text comes from collect.query_store_text now, and this projection's grain is exactly + that table's key, so it resolves here with a keyed join rather than inside the aggregate. + The MAX(query_text) sample below it stays as the fallback: it is where text lived before the + cutover, and it is what keeps the regression rows readable for existing history. */ + COALESCE(x.query_sql_text, r.query_text_sample) AS query_text_sample, + r.last_execution_time + FROM recent_performance AS r + JOIN baseline_performance AS b + ON b.database_name = r.database_name + AND b.query_id = r.query_id + LEFT JOIN query_store_text AS x + ON x.server_id = $1 + AND x.database_name = r.database_name + AND x.query_id = r.query_id + WHERE (r.avg_cpu_time_ms - b.avg_cpu_time_ms) * 100.0 / NULLIF(b.avg_cpu_time_ms, 0) > 25 + ORDER BY additional_duration_ms DESC + LIMIT $5 + """; + + /// + /// Whether this server has Query Store rows BEFORE the window, and whether it has any INSIDE it. + /// One round trip for the two facts that decide what an empty result means, and it is run only + /// on the empty path. Zero regressions is four different states here, not two. No baseline and no + /// recent rows means nothing was ever collected. A baseline with an empty window means collection may + /// have stopped. And - the one this read has that its siblings do not - RECENT rows with no baseline + /// means there is no BEFORE to compare against: a regression needs one, and a server whose entire + /// collected history sits inside the requested window can never show a regression however bad it got. + /// Reporting that as "no regressions" is the failure this exists to prevent. + /// Probes the base query_store_stats table, the same source the read itself uses. + /// $1 server_id, $2 window start, $3 window end. + /// + public const string RegressionCoverageSql = """ + SELECT + EXISTS ( + SELECT 1 + FROM query_store_stats + WHERE server_id = $1 + AND collection_time < $2 + ) AS has_baseline, + EXISTS ( + SELECT 1 + FROM query_store_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + ) AS has_recent + """; + + /// Runs . + public static async Task> GetQueryStoreRegressionsAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + string? databaseName, int limit, CancellationToken cancellationToken = default) + { + var rows = new List(); + await using var command = postgres.CreateCommand(QueryStoreRegressionsSql); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + command.Parameters.Add(new NpgsqlParameter + { + NpgsqlDbType = NpgsqlDbType.Array | NpgsqlDbType.Text, + Value = string.IsNullOrWhiteSpace(databaseName) ? DBNull.Value : new[] { databaseName }, + }); + command.Parameters.Add(new NpgsqlParameter { TypedValue = limit }); + + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + rows.Add(new RegressionRow( + reader.IsDBNull(0) ? "" : reader.GetString(0), + reader.IsDBNull(1) ? 0 : reader.GetInt64(1), + reader.IsDBNull(2) ? 0 : Convert.ToDouble(reader.GetValue(2)), + reader.IsDBNull(3) ? 0 : Convert.ToDouble(reader.GetValue(3)), + reader.IsDBNull(4) ? 0 : Convert.ToDouble(reader.GetValue(4)), + reader.IsDBNull(5) ? 0 : Convert.ToDouble(reader.GetValue(5)), + reader.IsDBNull(6) ? 0 : Convert.ToDouble(reader.GetValue(6)), + reader.IsDBNull(7) ? 0 : Convert.ToDouble(reader.GetValue(7)), + reader.IsDBNull(8) ? 0 : Convert.ToDouble(reader.GetValue(8)), + reader.IsDBNull(9) ? 0 : Convert.ToDouble(reader.GetValue(9)), + reader.IsDBNull(10) ? 0 : Convert.ToDouble(reader.GetValue(10)), + reader.IsDBNull(11) ? 0 : Convert.ToDouble(reader.GetValue(11)), + reader.IsDBNull(12) ? 0 : reader.GetInt64(12), + reader.IsDBNull(13) ? 0 : reader.GetInt64(13), + reader.IsDBNull(14) ? 0 : reader.GetInt32(14), + reader.IsDBNull(15) ? 0 : reader.GetInt32(15), + reader.IsDBNull(16) ? "" : reader.GetString(16), + reader.IsDBNull(17) ? "" : reader.GetString(17), + reader.IsDBNull(18) ? null : reader.GetDateTime(18))); + } + + return rows; + } + + /// Runs . + public static async Task<(bool HasBaseline, bool HasRecent)> GetCoverageAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(RegressionCoverageSql); + DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + if (!await reader.ReadAsync(cancellationToken)) + return (false, false); + return (reader.GetBoolean(0), reader.GetBoolean(1)); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingRuntimePrecondition.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingRuntimePrecondition.cs new file mode 100644 index 000000000..0e1393cbe --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingRuntimePrecondition.cs @@ -0,0 +1,274 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +/// +/// The store half of the #2546 runtime-precondition answer: reads what collection already recorded about a +/// collector's most recent run — or about the target's Query Store configuration — and, when it names a +/// precondition somebody can satisfy, returns the precondition envelope both SKUs emit. The DECISION +/// and the WORDS come from ; nothing about what counts as a +/// precondition, and nothing about how it is described, lives in this file or in its Lite twin +/// (McpRuntimePrecondition). +/// +/// Why the store and not a probe. The MCP surface never touches a monitored server — every read +/// is served from collected rows — so "is the session running right now" has to be answered from what the +/// last collection cycle found. That is not a compromise: the runners already classify these outcomes +/// (SESSION_MISSING for an absent capture session, the non-fatal bucket for a denied grant, a missing source +/// object or a disabled feature) and already write the actionable sentence into +/// collection_log.error_message. The store has held the answer all along; no read reported it. +/// +/// Called after the capability answer, and only on the miss path. A permanent engine gap wins: +/// a precondition message on an engine that can never have the surface would re-introduce the defect #2511 +/// closed. And like its capability sibling, this is asked AFTER the read came back empty, never before, so a +/// server whose recorded state disagrees with its collected rows still gets its DATA. +/// +/// A registry or log read that FAILS answers null, deliberately and for the same reason the +/// capability probe does: this runs on a path that has already found no data, and turning a diagnostic into +/// a read error would replace one honest miss with a worse one. +/// +internal static class DarlingRuntimePrecondition +{ + /// + /// Whether one collector has EVER run for one server, beside when that server last collected anything at + /// all. Both halves in one round trip, because the inference needs both and they must describe the same + /// instant: a collector with no rows on a server that is collecting normally is gated off, while a + /// collector with no rows on a server that has collected nothing is just a collection outage. + /// + /// EXISTS rather than a count or a MAX on the collector side: the question is presence, and + /// on a hypertable holding months of rows for dozens of collectors an existence probe stops at the first + /// hit while an aggregate reads the partition. $1 server_id, $2 collector. + /// + public const string CollectorEverRanSql = @" +SELECT EXISTS ( + SELECT 1 + FROM collection_log + WHERE server_id = $1 + AND collector_name = $2 + ) AS collector_ever_ran, + ( + SELECT MAX(collection_time) + FROM collection_log + WHERE server_id = $1 + ) AS server_last_collected"; + + /// + /// The most recent run of one collector for one server. Ordered by log_id rather than + /// collection_time because the id is monotonic per insert while two runs inside the same cycle can + /// share a timestamp — and "the latest run" is the whole claim this makes. $1 server_id, $2 collector. + /// + public const string LatestCollectorOutcomeSql = @" +SELECT status, error_message, collection_time +FROM collection_log +WHERE server_id = $1 +AND collector_name = $2 +ORDER BY log_id DESC +LIMIT 1"; + + /// + /// The latest per-database Query Store configuration snapshot, optionally narrowed to the one database a + /// read was scoped to. Reads v_query_store_health — the same relation + /// get_query_store_health answers from, so this can never claim a state that tool would not + /// confirm. $1 server_id, $2 database name or NULL for every database. + /// + /// The database filter is part of the EVIDENCE, not a convenience: a caller who asked + /// get_query_store_top for one database and got nothing needs to know about THAT database's Query + /// Store, and a server-wide answer would hide an off database behind a healthy sibling. + /// + public const string LatestQueryStoreStatesSql = @" +SELECT database_name, actual_state, capture_time +FROM v_query_store_health +WHERE server_id = $1 +AND capture_time = (SELECT MAX(capture_time) FROM v_query_store_health WHERE server_id = $1) +AND ($2::text IS NULL OR database_name = $2) +ORDER BY database_name"; + + /// + /// The precondition envelope when the collector serving this read recorded one on its most recent + /// run, or null when it did not — in which case the caller falls through to its own + /// empty/unavailable miss, unchanged. + /// + public static async Task StatusAsync( + NpgsqlDataSource postgres, + int serverId, + string serverName, + string collectorName, + CancellationToken cancellationToken = default) + { + string? status; + string? errorMessage; + DateTime? observedUtc; + + try + { + (status, errorMessage, observedUtc) = + await ReadLatestOutcomeAsync(postgres, serverId, collectorName, cancellationToken); + } + catch (Exception) + { + return null; + } + + var message = CollectorRuntimePrecondition.CollectionOutcomeMessage( + serverName, collectorName, status, errorMessage, observedUtc); + + return message is null ? null : McpHelpers.Status(CollectorRuntimePrecondition.StatusWord, message); + } + + /// + /// The precondition envelope when the collector serving this read has never run against this + /// server while the server is collecting normally — i.e. its AppliesTo gate is off — or + /// null otherwise. + /// + /// Call this AFTER , never instead of it: a collector that ran and + /// recorded a denial has a specific sentence from the monitored server itself, which is strictly better + /// evidence than this inference from an absence. + /// + /// Operator-facing text naming what could switch this collector off. The + /// deciding facts are not persisted, so the caller supplies the candidates rather than this method + /// guessing which one applies. + public static async Task GatedOffStatusAsync( + NpgsqlDataSource postgres, + int serverId, + string serverName, + string collectorName, + string gateCandidates, + CancellationToken cancellationToken = default) + { + bool everRan; + DateTime? serverLastCollectedUtc; + + try + { + (everRan, serverLastCollectedUtc) = + await ReadCollectorEverRanAsync(postgres, serverId, collectorName, cancellationToken); + } + catch (Exception) + { + /* Same rule as the rest of this class: a diagnostic that throws must not turn an honest miss + into a read error. */ + return null; + } + + var message = CollectorRuntimePrecondition.GatedOffMessage( + serverName, collectorName, gateCandidates, everRan, serverLastCollectedUtc); + + return message is null ? null : McpHelpers.Status(CollectorRuntimePrecondition.StatusWord, message); + } + + /// + /// The precondition envelope when this server's recorded Query Store configuration says the + /// databases the read covered are not recording runtime statistics, or null when it says nothing + /// of the kind (no snapshot yet, or at least one database in scope is READ_WRITE). + /// + public static async Task QueryStoreStatusAsync( + NpgsqlDataSource postgres, + int serverId, + string serverName, + string? databaseName, + CancellationToken cancellationToken = default) + { + List states; + DateTime? observedUtc; + + try + { + (states, observedUtc) = + await ReadQueryStoreStatesAsync(postgres, serverId, databaseName, cancellationToken); + } + catch (Exception) + { + return null; + } + + var message = CollectorRuntimePrecondition.QueryStoreDisabledMessage(serverName, states, observedUtc); + return message is null ? null : McpHelpers.Status(CollectorRuntimePrecondition.StatusWord, message); + } + + private static async Task<(string? Status, string? ErrorMessage, DateTime? ObservedUtc)> ReadLatestOutcomeAsync( + NpgsqlDataSource postgres, + int serverId, + string collectorName, + CancellationToken cancellationToken) + { + await using var command = postgres.CreateCommand(LatestCollectorOutcomeSql); + DarlingMcpReadParameters.AddInt(command, serverId); + DarlingMcpReadParameters.AddText(command, collectorName); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + + if (!await reader.ReadAsync(cancellationToken)) + { + return (null, null, null); + } + + return ( + reader.IsDBNull(0) ? null : reader.GetString(0), + reader.IsDBNull(1) ? null : reader.GetString(1), + reader.IsDBNull(2) ? null : DateTime.SpecifyKind(reader.GetDateTime(2), DateTimeKind.Utc)); + } + + private static async Task<(bool EverRan, DateTime? ServerLastCollectedUtc)> ReadCollectorEverRanAsync( + NpgsqlDataSource postgres, + int serverId, + string collectorName, + CancellationToken cancellationToken) + { + await using var command = postgres.CreateCommand(CollectorEverRanSql); + DarlingMcpReadParameters.AddInt(command, serverId); + DarlingMcpReadParameters.AddText(command, collectorName); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + + if (!await reader.ReadAsync(cancellationToken)) + { + /* No row is impossible for this shape (both halves are scalar subqueries), but treating it as + "we cannot tell" keeps the caller on its existing miss rather than asserting a gate. */ + return (true, null); + } + + return ( + !reader.IsDBNull(0) && reader.GetBoolean(0), + reader.IsDBNull(1) ? null : DateTime.SpecifyKind(reader.GetDateTime(1), DateTimeKind.Utc)); + } + + private static async Task<(List States, DateTime? ObservedUtc)> + ReadQueryStoreStatesAsync( + NpgsqlDataSource postgres, + int serverId, + string? databaseName, + CancellationToken cancellationToken) + { + var states = new List(); + DateTime? observedUtc = null; + + await using var command = postgres.CreateCommand(LatestQueryStoreStatesSql); + DarlingMcpReadParameters.AddInt(command, serverId); + DarlingMcpReadParameters.AddNullableText(command, string.IsNullOrWhiteSpace(databaseName) ? null : databaseName); + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + + while (await reader.ReadAsync(cancellationToken)) + { + states.Add(new CollectorRuntimePrecondition.QueryStoreDatabaseState( + reader.IsDBNull(0) ? string.Empty : reader.GetString(0), + reader.IsDBNull(1) ? null : reader.GetString(1))); + + observedUtc ??= reader.IsDBNull(2) + ? null + : DateTime.SpecifyKind(reader.GetDateTime(2), DateTimeKind.Utc); + } + + return (states, observedUtc); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingSystemHealthReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingSystemHealthReader.cs index 5b66ef3b0..4bf162294 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingSystemHealthReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingSystemHealthReader.cs @@ -91,6 +91,37 @@ public static async Task> ReadEventXmlAsync( return xmls; } + /// + /// Whether this server has EVER recorded a system_health event of one type, ignoring any window. + /// Lets an empty parse-on-read result say WHICH kind of nothing it found. Zero significant rows + /// is true both of a healthy window and of a server whose system_health events were never collected, + /// and the two want opposite responses -- widen the window, versus go find out why nothing is being + /// captured. Probes v_system_health_events, the SAME source + /// reads, so it cannot report a server as captured for rows + /// the read itself can never see. Scoped to the event_type because that is the granularity the caller + /// asked about: a server capturing sp_server_diagnostics but no wait_info has not been sampled for + /// waits, whatever its other categories hold. LIMIT 1, so it stops at the first row. + /// $1 server_id, $2 event_type. + /// + public const string HasAnyEventOfTypeSql = """ + SELECT 1 + FROM v_system_health_events + WHERE server_id = $1 + AND event_type = $2 + AND event_xml IS NOT NULL + LIMIT 1 + """; + + /// Runs . + public static async Task HasAnyEventOfTypeAsync( + NpgsqlDataSource postgres, int serverId, string eventType, CancellationToken cancellationToken = default) + { + await using var command = postgres.CreateCommand(HasAnyEventOfTypeSql); + DarlingMcpReadParameters.AddInt(command, serverId); + DarlingMcpReadParameters.AddText(command, eventType); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } + /// Loads the server's latest database_id → database_name map for Severe Errors DB resolution. public static async Task> GetDatabaseNameMapAsync( NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingTrendReader.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingTrendReader.cs index 7c613c070..e308b2776 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingTrendReader.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingTrendReader.cs @@ -64,9 +64,17 @@ public sealed record PerfmonTrendPoint(DateTime CollectionTime, long Value, long public sealed record FileIoLatencyTrendPoint( DateTime CollectionTime, string DatabaseName, double AvgReadLatencyMs, double AvgWriteLatencyMs); - /// One query-duration / execution-count trend point: the per-second rate (elapsed ms/sec) plus - /// executions/sec (Lite's QueryTrendPoint shape). - public sealed record QueryDurationTrendPoint(DateTime CollectionTime, double Value, long ExecutionCount); + /// + /// One query-duration / execution-count trend point, shared by the three Performance-Trends siblings: + /// the per-second rate (elapsed ms/sec) plus executions/sec (Lite's QueryTrendPoint shape). + /// ExecutionCount and ExecutionsPerSecond are the SAME quantity - executions per + /// second - and both are here because the first one shipped truncated to a long. On a server doing + /// three executions a second that rounds harmlessly; on a quiet one doing 0.4 it reports ZERO, which + /// reads as an idle server rather than a slow one. The long is kept so a consumer reading it does not + /// break; new readers should take ExecutionsPerSecond. + /// + public sealed record QueryDurationTrendPoint( + DateTime CollectionTime, double Value, long ExecutionCount, double ExecutionsPerSecond); /// One point of a single query's per-collection history (Lite's QueryStatsHistoryRow, /// the columns get_query_trend surfaces): the interval deltas + DOP spread + the plan hash. Time metrics @@ -345,19 +353,156 @@ FROM raw ORDER BY collection_time """; - public static async Task> GetQueryDurationTrendAsync( + /// Runs . + public static Task> GetQueryDurationTrendAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + => ReadDurationTrendAsync(QueryDurationTrendSql, postgres, serverId, startUtc, endUtc, cancellationToken); + + /* --------------------- procedure + Query Store duration trends (#2484) --------------------- */ + + /// + /// The procedure-stats duration trend - the viewer's ProcedureDurationTrendSql, verbatim apart + /// from the database filter the MCP copy of its query-stats twin already drops. + /// Not a duplicate of the query-stats trend, and the difference is the point: query_stats + /// attributes a procedure's work to the individual statements inside it, so a procedure that got slower + /// shows up smeared across however many statements it runs. This charges the whole call to the + /// procedure. When both are available, the pair answers "did ad-hoc SQL regress, or did a procedure?" - + /// which one series alone never can. $1 server_id, $2/$3 window (naive UTC). + /// + public const string ProcedureDurationTrendSql = """ + WITH raw AS + ( + SELECT + collection_time, + SUM(delta_elapsed_time) / 1000.0 AS total_elapsed_ms, + SUM(delta_execution_count) AS total_executions, + extract(epoch FROM (date_trunc('second', collection_time) - date_trunc('second', LAG(collection_time) OVER (ORDER BY collection_time)))) AS interval_seconds + FROM procedure_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + GROUP BY collection_time + ) + SELECT + collection_time, + CASE WHEN interval_seconds > 0 THEN total_elapsed_ms / interval_seconds ELSE 0 END AS elapsed_ms_per_second, + CASE WHEN interval_seconds > 0 THEN CAST(total_executions AS DOUBLE PRECISION) / interval_seconds ELSE 0 END AS executions_per_second + FROM raw + ORDER BY collection_time + """; + + /// + /// The Query Store duration trend - the viewer's QueryStoreDurationTrendSql, verbatim apart from + /// the database filter, INCLUDING the #1841 tier-2 interval placement and its legacy arm. + /// Copied rather than simplified. The two arms are not decoration: arm 1 dedups each runtime + /// interval to its final cumulative snapshot and places it at interval_start_time_utc - the hour + /// the work actually RAN - because Query Store has no delta columns and re-fetches an open interval's + /// running totals every cycle, so charging every fetch to its collection time triple-counts the same + /// work. Arm 2 keeps rows collected before that fix on their old un-deduped treatment, because no + /// interval start exists for them and none can be reconstructed. The arms split on + /// interval_start_time_utc IS NULL, so they partition the rows with no overlap and no gap. + /// Rewriting either arm here would make the browser and the desktop viewer disagree about the same + /// hour. $1 server_id, $2/$3 window (naive UTC). + /// + public const string QueryStoreDurationTrendSql = """ + WITH placed AS + ( + /* Arm 1 (#1841 tier 2) - rows carrying the interval identity. Dedup to the interval's FINAL + cumulative snapshot, then place it at interval_start_time_utc: the hour the work ran, not + the cycle that last fetched it. */ + SELECT + interval_start_time_utc AS point_time, + execution_count, + avg_duration_us + FROM + ( + SELECT + interval_start_time_utc, + execution_count, + avg_duration_us, + ROW_NUMBER() OVER + ( + PARTITION BY database_name, query_id, plan_id, runtime_stats_interval_id, first_execution_time, execution_type_desc, replica_role + ORDER BY collection_time DESC, execution_count DESC + ) AS rn + FROM query_store_stats + WHERE server_id = $1 + /* Windowed on the column this arm PLACES its points at (#1892). Filtering on + collection_time here put a point outside the range the caller asked for, and dropped + the range's final interval because its closing fetch had not happened yet. */ + AND interval_start_time_utc >= $2 + AND interval_start_time_utc <= $3 + /* Chunk-exclusion bounds only. */ + AND collection_time >= $2 - interval '1 day' + AND collection_time <= $3 + interval '30 days' + AND interval_start_time_utc IS NOT NULL + ) AS identified + WHERE rn = 1 + + UNION ALL + + /* Arm 2 - rows collected before tier 2. No interval start exists and none can be + reconstructed, so these keep the pre-tier-2 treatment byte for byte. */ + SELECT + collection_time AS point_time, + execution_count, + avg_duration_us + FROM query_store_stats + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND interval_start_time_utc IS NULL + ), + raw AS + ( + SELECT + point_time, + SUM(execution_count * avg_duration_us / 1000.0) AS total_duration_ms, + SUM(execution_count) AS total_executions, + extract(epoch FROM (date_trunc('second', point_time) - date_trunc('second', LAG(point_time) OVER (ORDER BY point_time)))) AS interval_seconds + FROM placed + GROUP BY point_time + ) + SELECT + point_time AS collection_time, + CASE WHEN interval_seconds > 0 THEN total_duration_ms / interval_seconds ELSE 0 END AS duration_ms_per_second, + CASE WHEN interval_seconds > 0 THEN CAST(total_executions AS DOUBLE PRECISION) / interval_seconds ELSE 0 END AS executions_per_second + FROM raw + ORDER BY point_time + """; + + /// Runs . + public static Task> GetProcedureDurationTrendAsync( + NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + => ReadDurationTrendAsync(ProcedureDurationTrendSql, postgres, serverId, startUtc, endUtc, cancellationToken); + + /// Runs . + public static Task> GetQueryStoreDurationTrendAsync( NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) + => ReadDurationTrendAsync(QueryStoreDurationTrendSql, postgres, serverId, startUtc, endUtc, cancellationToken); + + /// + /// Shared reader for the three duration trends. All three project the same three columns - point time, + /// a per-second value, a per-second execution rate - which is the viewer's own arrangement + /// (ReadDurationTrendAsync), kept so the three series cannot drift apart in how they are read. + /// Summed bigint deltas come back as Postgres numeric, so the values Convert tolerantly. + /// + private static async Task> ReadDurationTrendAsync( + string sql, NpgsqlDataSource postgres, int serverId, DateTime startUtc, DateTime endUtc, + CancellationToken cancellationToken) { var items = new List(); - await using var command = postgres.CreateCommand(QueryDurationTrendSql); + await using var command = postgres.CreateCommand(sql); DarlingMcpReadParameters.AddWindow(command, serverId, startUtc, endUtc); await using var reader = await command.ExecuteReaderAsync(cancellationToken); while (await reader.ReadAsync(cancellationToken)) { + var executionsPerSecond = reader.IsDBNull(2) ? 0 : Convert.ToDouble(reader.GetValue(2)); items.Add(new QueryDurationTrendPoint( reader.GetDateTime(0), reader.IsDBNull(1) ? 0 : Convert.ToDouble(reader.GetValue(1)), - reader.IsDBNull(2) ? 0 : (long)Convert.ToDouble(reader.GetValue(2)))); + (long)executionsPerSecond, + executionsPerSecond)); } return items; @@ -494,4 +639,103 @@ private static async Task> ReadQueryHistoryAsync( return items; } + + /* ─────────── "which nothing is this?" probes for the three windowed trends ─────────── */ + + /// + /// Whether this server has EVER recorded a memory sample, ignoring any window. + /// Exists so an empty memory trend can say WHICH kind of nothing it found. "No memory trend data + /// available" is true both of a quiet window and of a server the collector has never touched, and those + /// want opposite responses from the caller — widen the window, versus go find out why collection is not + /// running. Reads v_memory_stats, the same source reads, so it can + /// never report "collected" for rows the trend cannot see. LIMIT 1, so it stops at the first row. + /// + public const string HasAnyMemoryStatSql = """ + SELECT 1 + FROM v_memory_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// Whether this server has EVER recorded a file I/O sample. Reads v_file_io_stats, the + /// same source reads. See . + public const string HasAnyFileIoStatSql = """ + SELECT 1 + FROM v_file_io_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// + /// Whether this server has EVER recorded a query-stats sample. + /// Reads the BASE query_stats table, deliberately, because + /// does: on a V38+ store v_query_stats is the payload-RESOLVING view, not a passthrough, and the + /// duration trend projects no text so it never needs it. Probing the view here would be probing a + /// different relation from the one the read walks — the exact way an existence probe reports the wrong + /// branch in the case it exists to get right. + /// + public const string HasAnyQueryStatSql = """ + SELECT 1 + FROM query_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// Whether this server has EVER recorded a stored-procedure sample. Reads the BASE + /// procedure_stats table for the same reason the query-stats probe above does — it is what + /// reads. See . + public const string HasAnyProcedureStatSql = """ + SELECT 1 + FROM procedure_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// + /// Whether this server has EVER recorded a Query Store sample. Reads the BASE + /// query_store_stats table, the source walks. + /// Worth the most of the five, because zero rows here has a cause the others do not: Query Store + /// can simply be OFF on every database. A server with no Query Store data is not a server with no slow + /// queries, and the read has to say so rather than return a clean-looking empty series. + /// + public const string HasAnyQueryStoreStatSql = """ + SELECT 1 + FROM query_store_stats + WHERE server_id = $1 + LIMIT 1 + """; + + /// Runs . + public static Task HasAnyMemoryStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnySampleAsync(postgres, HasAnyMemoryStatSql, serverId, cancellationToken); + + /// Runs . + public static Task HasAnyFileIoStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnySampleAsync(postgres, HasAnyFileIoStatSql, serverId, cancellationToken); + + /// Runs . + public static Task HasAnyQueryStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnySampleAsync(postgres, HasAnyQueryStatSql, serverId, cancellationToken); + + /// Runs . + public static Task HasAnyProcedureStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnySampleAsync(postgres, HasAnyProcedureStatSql, serverId, cancellationToken); + + /// Runs . + public static Task HasAnyQueryStoreStatAsync( + NpgsqlDataSource postgres, int serverId, CancellationToken cancellationToken = default) + => HasAnySampleAsync(postgres, HasAnyQueryStoreStatSql, serverId, cancellationToken); + + /// All five probes share one shape: a scalar that is null when no row qualifies. + private static async Task HasAnySampleAsync( + NpgsqlDataSource postgres, string sql, int serverId, CancellationToken cancellationToken) + { + await using var command = postgres.CreateCommand(sql); + DarlingMcpReadParameters.AddInt(command, serverId); + return await command.ExecuteScalarAsync(cancellationToken) is not null; + } } diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingWebHostService.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingWebHostService.cs index 533d1e534..48040e0dc 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingWebHostService.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingWebHostService.cs @@ -11,6 +11,7 @@ using System.Net; using System.Runtime.Versioning; using System.Security.Cryptography; +using System.Security.Cryptography.X509Certificates; using System.Text; using System.Threading; using System.Threading.Tasks; @@ -48,7 +49,14 @@ namespace PerformanceMonitor.Darling.Service.Mcp; /// cookie or a valid ?token= (constant-time), which is exchanged for an HMAC-signed HttpOnly /// SameSite=Strict cookie and 302-redirected to strip the token from the URL; out-of-CIDR is 403; no /// cookie/token gets a minimal inline login form. The cookie signing key is a per-process 32-byte RNG value, -/// so a restart invalidates sessions (acceptable). No TLS (the same reverse-proxy story as MCP). +/// so a restart invalidates sessions (acceptable). +/// +/// TLS (#2562): opt-in via web.network.tls and applied to the NETWORK listener only — see +/// for the certificate rules and the Kestrel bind below for why the +/// loopback listeners stay plain HTTP. Without it the exposed listener is plain HTTP and every start warns +/// that the token and its cookie cross the segment in the clear. A certificate that is missing, unreadable, +/// ambiguous or expired fail-closes to loopback-only exactly as an undecryptable token does; it never +/// downgrades to serving the LAN over HTTP. /// /// The dashboard connects to the store as the least-privilege VIEWER role (not owner, not mcp) — a /// read-only pool. Static assets ship from wwwroot (a csproj Content copy) with the content root AND @@ -61,6 +69,14 @@ public sealed class DarlingWebHostService : BackgroundService private readonly WebRuntimeState _state; private WebApplication? _app; private NpgsqlDataSource? _appDataSource; + + /// The TLS certificate the current network listener presents — the leaf AND the intermediates + /// that travel with it (#2562) — held for its lifetime. Disposal is not bookkeeping: on Windows the + /// private key is loaded with MachineKeySet and no PersistKeySet, so disposing is what + /// REMOVES the key material from the machine key store. A rebind that leaked it would accumulate a key + /// per restart. + private DarlingWebTls.LoadedCertificate? _serverCertificate; + private int _runningPort; /// How often the supervisor re-reads the live control-plane state (#1562, mirrors the MCP host). @@ -117,6 +133,9 @@ protected override async Task ExecuteAsync(CancellationToken stoppingToken) lifetime exactly as before (the network exposure block is restart-only by design). */ DarlingConfig? config = null; var lastFailedStartUtc = DateTime.MinValue; + /* #2389: the last control-plane-override report emitted, so a steady disagreement is stated once per + distinct state instead of on every 5s poll tick. */ + string? lastOverrideReport = null; while (!stoppingToken.IsCancellationRequested) { if (config is null && DateTime.UtcNow - lastFailedStartUtc >= FailedStartBackoff) @@ -149,13 +168,28 @@ lifetime exactly as before (the network exposure block is restart-only by design } var published = _state.Read(); - var enabled = published?.Enabled ?? config.Web.Enabled; - var desiredPort = published?.Port ?? config.Web.Port; - switch (DecideWebAction(_app is not null, _runningPort, enabled, desiredPort)) + /* #2389: the store still wins whenever the worker has published (unchanged), but the resolution + now carries WHICH plane supplied each value — the MCP host's twin, sharing one resolver so the + two surfaces cannot drift. */ + var toggle = DarlingHostBinding.ResolveEndpointToggle( + published is null ? null : (published.Enabled, published.Port), config.Web.Enabled, config.Web.Port); + + /* Report the DISAGREEMENT at the point of override, once per distinct state, so a file edit that + the control plane is quietly ignoring says so instead of presenting as a successful start. */ + var overrideReport = DarlingHostBinding.DescribeToggleOverride( + toggle, "web", "Web dashboard", config.Web.Enabled, config.Web.Port); + if (overrideReport is not null && !string.Equals(overrideReport, lastOverrideReport, StringComparison.Ordinal)) + { + _logger.LogWarning("{Report}", overrideReport); + } + + lastOverrideReport = overrideReport; + + switch (DecideWebAction(_app is not null, _runningPort, toggle.Enabled, toggle.Port)) { case WebSupervisorAction.Start when DateTime.UtcNow - lastFailedStartUtc >= FailedStartBackoff: - if (!await TryStartServerAsync(config, desiredPort, stoppingToken)) + if (!await TryStartServerAsync(config, toggle, stoppingToken)) { lastFailedStartUtc = DateTime.UtcNow; } @@ -168,9 +202,9 @@ lifetime exactly as before (the network exposure block is restart-only by design case WebSupervisorAction.Restart: _logger.LogInformation( - "Web dashboard port changed via the control plane ({Old} -> {New}) — rebinding", _runningPort, desiredPort); + "Web dashboard port changed via the control plane ({Old} -> {New}) — rebinding", _runningPort, toggle.Port); await StopServerAsync(stoppingToken); - if (!await TryStartServerAsync(config, desiredPort, stoppingToken)) + if (!await TryStartServerAsync(config, toggle, stoppingToken)) { lastFailedStartUtc = DateTime.UtcNow; } @@ -218,6 +252,9 @@ private async Task StopServerAsync(CancellationToken cancellationToken) _appDataSource = null; } + _serverCertificate?.Dispose(); + _serverCertificate = null; + _runningPort = 0; } @@ -235,6 +272,9 @@ private async Task DisposeFailedStartAsync() try { await _appDataSource.DisposeAsync(); } catch { /* best-effort */ } _appDataSource = null; } + + try { _serverCertificate?.Dispose(); } catch { /* best-effort */ } + _serverCertificate = null; } /// @@ -260,13 +300,18 @@ internal static DarlingHostBinding.BindDecision ResolveWebBind(WebConfig web, bo } /// - /// One start ATTEMPT of the inner web app at : the port comes from the live + /// One start ATTEMPT of the inner web app at 's port: the port comes from the live /// control-plane value, and every bail path returns false so the supervisor retries with backoff instead of /// standing down for the process lifetime. The bind/network/token decisions come from the FILE-loaded config /// (network exposure is deliberately restart-only); returns true when the app is started and listening. + /// #2389: the toggle carries the enable/port PROVENANCE, not just the port, so the start line names + /// the plane each half of the bind came from. /// - private async Task TryStartServerAsync(DarlingConfig config, int effectivePort, CancellationToken stoppingToken) + private async Task TryStartServerAsync( + DarlingConfig config, DarlingHostBinding.EndpointToggle toggle, CancellationToken stoppingToken) { + var effectivePort = toggle.Port; + var web = config.Web; var network = web.Network; @@ -318,6 +363,141 @@ private async Task TryStartServerAsync(DarlingConfig config, int effective } } + /* TLS for the network listener (#2562). Resolved HERE rather than in the pure ResolveBind ladder for + the same reason the token is: loading a certificate reads files and a clock, and the ladder is + kept free of both. It also means a certificate failure degrades exactly the way a token failure + does — Critical, then loopback-only — instead of needing its own BindReason, which the MCP host's + parallel enum would have had to grow a member it can never use. */ + DarlingWebTls.LoadedCertificate? serverCertificate = null; + if (networkMode) + { + var plan = DarlingWebTls.Describe(network!.Tls); + switch (plan.Shape) + { + case DarlingWebTls.TlsShape.NotConfigured: + /* The pre-#2562 behaviour, still the default, and still the only thing a plain-HTTP + reverse proxy in front of the port needs. Warn every start: the token and the cookie + it mints are readable by anything on the segment, and that is easy not to notice + precisely because the dashboard works perfectly. */ + _logger.LogWarning( + "Web dashboard is LAN-exposed WITHOUT TLS — the access token and its session cookie cross the " + + "segment in the clear, and web.network.allowFrom bounds only who can route to the port. " + + "Configure web.network.tls (a PKCS#12 bundle or a PEM pair), or front the port with a " + + "TLS-terminating reverse proxy."); + break; + + case DarlingWebTls.TlsShape.Invalid: + _logger.LogCritical( + "Web dashboard TLS is misconfigured ({Problem}) — refusing to expose; binding loopback-only.", + plan.Problem); + networkMode = false; + break; + + default: + try + { + var loaded = DarlingWebTls.Load(network.Tls!, plan.Shape); + var certificate = loaded.Leaf; + + if (plan.Warning is not null) + { + _logger.LogWarning("Web dashboard TLS: {Warning}", plan.Warning); + } + + /* Lifetime is checked BEFORE the listener is built, not left to the handshake: an + expired certificate takes the dashboard down either way, and this is the only + place the reason reaches an operator's log. + + ToUniversalTime() is not decoration: X509Certificate2.NotBefore/NotAfter return + LOCAL DateTimes, and while the implicit DateTime->DateTimeOffset conversion does + carry the local offset and would compare correctly, it reads as a UTC value to + everyone who follows. Convert where the trap is, not where it detonates. */ + var refusal = DarlingWebTls.LifetimeRefusal( + certificate.NotBefore.ToUniversalTime(), + certificate.NotAfter.ToUniversalTime(), + DateTimeOffset.UtcNow); + if (refusal is not null) + { + loaded.Dispose(); + _logger.LogCritical( + "Web dashboard TLS certificate cannot be used ({Refusal}) — refusing to expose; binding loopback-only.", + refusal); + networkMode = false; + break; + } + + /* Adopted by the field IMMEDIATELY, before any of the bail paths below it (port in + use, store credential not ready), so every one of them releases the key. */ + serverCertificate = loaded; + _serverCertificate = loaded; + + /* The SAN has to name the IP, not a hostname, and that is a consequence of a + control that lives two files away: the anti-DNS-rebind Host allowlist accepts + only `localhost`, a loopback literal, or the configured listen IP, so a LAN + client CANNOT reach this dashboard by DNS name — it is refused 400 before any + route runs. An internal CA asked for "a certificate for darling.corp.local" + issues exactly the certificate that can never match here, and the operator + would learn that from a browser warning rather than from us. + + A warning, not a refusal: the dashboard genuinely works after a click-through, + and taking it down over a name mismatch would be a worse outcome than saying so. + Skipped on a wildcard bind, where there is no single address to match against. + Reuses the store's own SAN reader rather than growing a second one. */ + if (!networkListenIp!.Equals(IPAddress.Any) + && !networkListenIp.Equals(IPAddress.IPv6Any) + && !DarlingManagedPostgres.CertificateSanCoversIp(certificate, networkListenIp)) + { + /* One placeholder per argument, in order. A repeated {Listen} would read + naturally and throw at format time, which the surrounding catch would + report as "web dashboard failed to start". */ + _logger.LogWarning( + "Web dashboard TLS certificate carries no iPAddress SAN for {Listen} — every browser " + + "will report a name mismatch. The anti-DNS-rebind Host allowlist accepts only that " + + "literal IP or loopback in the Host header, so a LAN client has to browse to it by IP " + + "on port {Port} and a DNS-name-only certificate can never match. Reissue the " + + "certificate with an iPAddress SAN.", + networkListenIp, effectivePort); + } + + var expiry = DarlingWebTls.ExpiryWarning( + certificate.NotAfter.ToUniversalTime(), DateTimeOffset.UtcNow); + if (expiry is not null) + { + _logger.LogWarning( + "Web dashboard TLS certificate {Expiry} — subject {Subject}, thumbprint {Thumbprint}. " + + "The dashboard stops serving when it lapses; renew it before then.", + expiry, certificate.Subject, certificate.Thumbprint); + } + else + { + _logger.LogInformation( + "Web dashboard TLS certificate loaded — subject {Subject}, thumbprint {Thumbprint}, expires {NotAfter:u}.", + certificate.Subject, certificate.Thumbprint, certificate.NotAfter.ToUniversalTime()); + } + } + catch (Exception ex) + { + /* Release whatever was already adopted. Adoption happens BEFORE the SAN check and + the expiry logging, and both of those still run inside this try — the SAN reader + re-materializes an extension from raw DER and can throw on a malformed one. A + throw there lands here holding a certificate the listener will never use, and + this degrade RETURNS TRUE (loopback-only started fine), so the method's outer + catch and DisposeFailedStartAsync never see it. Without these two lines the + "every bail path releases the key" claim on _serverCertificate is false. */ + _serverCertificate?.Dispose(); + _serverCertificate = null; + serverCertificate = null; + + _logger.LogCritical( + "Web dashboard TLS certificate could not be loaded ({Message}) — refusing to expose; binding loopback-only.", + ex.Message); + networkMode = false; + } + + break; + } + } + /* The REAL primary bind address (network IP when exposed, else loopback): both the port precheck and the Kestrel bind use it, so the precheck probes the actual address, not always loopback. */ var primaryBind = networkMode ? networkListenIp! : IPAddress.Loopback; @@ -327,6 +507,7 @@ firewall check so a bail here reports nothing about the firewall. */ if (await PortUtilityService.IsTcpPortListeningAsync(effectivePort, primaryBind, stoppingToken)) { _logger.LogError("Port {Port} is already in use — web dashboard not started this attempt; will retry", effectivePort); + await DisposeFailedStartAsync(); return false; } @@ -348,12 +529,14 @@ await CheckWebFirewallAsync( if (!OperatingSystem.IsWindows()) { _logger.LogError("Web dashboard not started: postgres.managed = true requires Windows"); + await DisposeFailedStartAsync(); return false; } storeConnectionString = await WaitForManagedConnectionStringAsync(config.Postgres, stoppingToken); if (storeConnectionString is null) { + await DisposeFailedStartAsync(); return false; } } @@ -375,23 +558,56 @@ in production only. WebRootPath is resolved against ContentRootPath -> AppContex WebRootPath = "wwwroot", }); + var listenerCertificate = serverCertificate; builder.WebHost.ConfigureKestrel(options => { if (networkMode) { /* Bind the specific family, then ALSO both loopback families (unless the listen is itself loopback/wildcard, which would collide) so a local client resolving "localhost" still - reaches the dashboard (loopback stays tokenless). */ - options.Listen(primaryBind, effectivePort); + reaches the dashboard. Loopback is exempt from the CIDR test only — since #1649 it + authenticates like any other remote. */ + options.Listen(primaryBind, effectivePort, listen => + { + if (listenerCertificate is not null) + { + /* ServerCertificateChain, not just ServerCertificate: Kestrel presents ONLY what + it is handed, so an intermediate left out here is an incomplete chain and a + failed handshake on every client that has not independently cached it. Measured + against a real leaf+intermediate PEM before this was wired: the server sent one + certificate. */ + listen.UseHttps(https => + { + https.ServerCertificate = listenerCertificate.Value.Leaf; + if (listenerCertificate.Value.Chain.Count > 0) + { + https.ServerCertificateChain = listenerCertificate.Value.Chain; + } + }); + } + }); + if (DarlingHostBinding.ShouldAddLoopbackListeners(primaryBind)) { + /* The loopback listeners stay PLAIN HTTP even when the network listener is TLS, and + that is a decision rather than an oversight. The certificate names the LAN address + the operator exposes; it almost never also names "localhost", so serving TLS here + would hand the local browser a name-mismatch warning on the one surface that never + leaves the machine. Nothing is lost: loopback traffic is not on the segment this + feature protects. When the listen IS a wildcard, ShouldAddLoopbackListeners is false + and the single listener is TLS for everyone — the operator asked for all interfaces. + + This is also the answer to "redirect or refuse" for the exposed address: one port + cannot speak both schemes, so a plain-HTTP client hitting the TLS listener fails at + the handshake. Adding a second HTTP port to redirect would re-open, on a new port, + the cleartext surface this exists to close. */ options.Listen(IPAddress.Loopback, effectivePort); options.Listen(IPAddress.IPv6Loopback, effectivePort); } } else { - /* The default/degraded loopback-only server — both families. */ + /* The default/degraded loopback-only server — both families, always plain HTTP. */ options.ListenLocalhost(effectivePort); } }); @@ -405,6 +621,11 @@ reaches the dashboard (loopback stays tokenless). */ _app = builder.Build(); + /* #2479 item 5: the gates below used to refuse silently. Rate-limited per (gate, source), + because this port is LAN-exposed on purpose - see DarlingHttpRefusalLog. Created per + started server so a rebind starts with a clean budget. */ + var refusals = new DarlingHttpRefusalLog(); + /* Pipeline order: the Host-allowlist middleware runs FIRST on EVERY request (both modes) as the DNS-rebinding guard, then (network mode only) the auth middleware, then DarlingWebEndpoints.MapAll -> UseDefaultFiles -> UseStaticFiles. WebApplication auto-inserts UseRouting at the head and @@ -423,6 +644,12 @@ reaches the dashboard (loopback stays tokenless). */ { if (!IsAllowedHost(context.Request.Host.Host, networkListenIp)) { + refusals.Report( + _logger, "Web dashboard", DarlingRefusalGate.HostAllowlist, StatusCodes.Status400BadRequest, + context.Connection.RemoteIpAddress, + $"the Host header '{DarlingHttpRefusalLog.Sanitize(context.Request.Host.Host)}' is not an address this endpoint binds" + + " (a loopback name/IP, or web.network.listen when LAN-exposed)", + DateTime.UtcNow); context.Response.StatusCode = StatusCodes.Status400BadRequest; return; } @@ -445,8 +672,8 @@ only the network auth decision (session cookie / ?token= / in-CIDR). */ var remote = context.Connection.RemoteIpAddress; var hasValidCookie = TryValidateSessionCookie( context.Request.Cookies[SessionCookieName], signingKey, DateTimeOffset.UtcNow); - var hasValidToken = DarlingHostBinding.FixedTimeTokenEquals( - context.Request.Query["token"].ToString(), token); + var presentedToken = context.Request.Query["token"].ToString(); + var hasValidToken = DarlingHostBinding.FixedTimeTokenEquals(presentedToken, token); switch (DecideWebAuth(remote, cidr, hasValidCookie, hasValidToken)) { @@ -462,10 +689,34 @@ only the network auth decision (session cookie / ?token= / in-CIDR). */ return; case WebAuthAction.Forbid: + /* The ONLY 403 this host produces: an out-of-CIDR remote, or one whose address + ASP.NET Core could not report (which fails closed). A wrong credential from + inside the CIDR is ShowLogin, not this - see below. */ + refusals.Report( + _logger, "Web dashboard", DarlingRefusalGate.SourceCidr, StatusCodes.Status403Forbidden, + remote, + remote is null + ? "its source address could not be determined, which fails closed" + : $"its address is outside web.network.allowFrom ({cidr})", + DateTime.UtcNow); context.Response.StatusCode = StatusCodes.Status403Forbidden; return; default: /* ShowLogin */ + /* A 200 rather than a 401, so this is not a "rejected request" by status - and + it is exactly the state an operator asks about when a ?token= they pasted did + not work. Logged ONLY when a token was actually presented and did not match: + a first visit with no token is the normal path to the login page and logging + it would make every bookmark a warning. */ + if (!string.IsNullOrEmpty(presentedToken)) + { + refusals.Report( + _logger, "Web dashboard", DarlingRefusalGate.Token, StatusCodes.Status200OK, + remote, + "the presented ?token= does not match web.network.encryptedToken, so the login page was served instead", + DateTime.UtcNow); + } + await WriteLoginPageAsync(context); return; } @@ -476,17 +727,37 @@ only the network auth decision (session cookie / ?token= / in-CIDR). */ _app.UseDefaultFiles(); _app.UseStaticFiles(); + /* #2389: name the authority for each half of what is being started — enabled/port from whichever + plane the supervisor resolved, listen/allowFrom/token always from darling.json. */ + var origin = DarlingHostBinding.DescribeToggleOrigin(toggle); if (networkMode) { + /* #2562: name the SCHEME the exposed listener actually speaks. An operator reading this line is + deciding whether the token they are about to paste crosses the wire in the clear, and the + line used to say "http://" unconditionally because that was the only thing it could be. */ _logger.LogInformation( - "Starting web dashboard on http://{Listen}:{Port} (LAN-exposed to {Cidr} behind a token->cookie gate + in-app CIDR; loopback also bound, tokenless)", - primaryBind, effectivePort, allowedCidr); + "Starting web dashboard on {Scheme}://{Listen}:{Port} (LAN-exposed to {Cidr} behind a token->cookie gate + in-app CIDR; " + + "loopback also bound over plain HTTP, and since #1649 it authenticates too) — " + + "enabled/port from {Origin}; listen/allowFrom/token/tls from darling.json web.network (file-only, restart-only)", + serverCertificate is null ? "http" : "https", primaryBind, effectivePort, allowedCidr, origin); } else { - _logger.LogInformation("Starting web dashboard on http://localhost:{Port} (loopback only)", effectivePort); + _logger.LogInformation( + "Starting web dashboard on http://localhost:{Port} (loopback only) — enabled/port from {Origin}", + effectivePort, origin); } + /* #2479 item 6: the network block is read ONCE and held for the process lifetime by design. + Say so at every start, in BOTH modes - the loopback line above never mentioned the block at + all, and loopback-when-you-expected-LAN is exactly the state being diagnosed. */ + _logger.LogInformation( + "{Report}", + DarlingHostBinding.DescribeNetworkBlockLifetime( + "web", "Web dashboard", config.Web.Network?.IsConfigured ?? false, networkMode, + networkMode ? primaryBind.ToString() : null, + networkMode ? allowedCidr.ToString() : null)); + /* StartAsync, not RunAsync: the supervisor loop owns the wait — the app keeps serving until StopServerAsync (toggle-off, port change, or shutdown). */ await _app.StartAsync(stoppingToken); @@ -495,7 +766,9 @@ only the network auth decision (session cookie / ?token= / in-CIDR). */ } catch (OperationCanceledException) when (stoppingToken.IsCancellationRequested) { - /* Normal shutdown mid-start. */ + /* Normal shutdown mid-start — still release anything already acquired (#2562: the certificate's + private key is held in the machine key store until it is disposed). */ + await DisposeFailedStartAsync(); return false; } catch (Exception ex) @@ -576,15 +849,65 @@ still has to authenticate below. */ } /// - /// Builds a session cookie value {expiryUnix}.{base64url(HMAC-SHA256(key, expiryUnix))}. PURE — the - /// HMAC is over the exact expiry string that is stored, so verifies - /// the same bytes it signs (a tampered expiry fails the HMAC). + /// The longest will mint into a cookie. + /// + /// Generous for an OIDC sub — the identifiers real providers issue are GUIDs, opaque + /// 40-character strings, or an email — and a hard REFUSAL rather than a truncation, deliberately. + /// Truncating an identity is the one failure mode that would be worse than not having one: two subjects + /// sharing a prefix would collapse to the same seat, so the surface would report the wrong person as the + /// author of a change and every downstream authorization decision would be made about somebody else. + /// Failing the sign-in loudly is recoverable; silently merging two identities is not. + /// + internal const int MaxSessionSubjectLength = 256; + + /// + /// Builds a session cookie value. PURE. + /// + /// Two shapes, and the difference is whether the seat has a NAME: + /// {expiryUnix}.{base64url(HMAC)} for the shared-token seat, which has no identity to carry, and + /// {expiryUnix}.{base64url(subject)}.{base64url(HMAC)} once a per-user sign-in has established + /// who is holding it (#2550). + /// + /// The signature covers the whole prefix before the FINAL dot, which is what makes the two + /// shapes one construction rather than two. Signing only the expiry and appending the subject beside it + /// would leave the subject unauthenticated — anyone holding a valid cookie could rewrite it to any other + /// subject and be served as that person, which is a worse position than the shared token this exists to + /// improve on, because it would look like identity while providing none. Extending the signed region + /// instead means expiry and subject are tamper-evident together. + /// + /// The two shapes cannot be confused for each other even though they share a separator: the + /// signature is always the last segment and the signed payload is always everything before it, so + /// re-presenting a 3-segment cookie as a 2-segment one requires the subject to be a valid HMAC over the + /// expiry, and the reverse requires forging an HMAC. Neither is available without the key. + /// + /// The subjectless form is byte-for-byte what this method produced before per-user identity existed, + /// so cookies already in browsers keep validating across the upgrade instead of signing everyone out. /// - internal static string BuildSessionCookieValue(byte[] signingKey, DateTimeOffset expiry) + /// The authenticated principal, or null/empty for the shared-token seat. + /// exceeds + /// — see the note there on why this refuses rather than truncates. + /// + internal static string BuildSessionCookieValue(byte[] signingKey, DateTimeOffset expiry, string? subject = null) { + if (subject is { Length: > MaxSessionSubjectLength }) + { + throw new ArgumentException( + $"Session subject is {subject.Length} characters, which exceeds the {MaxSessionSubjectLength}-character " + + "limit. It is refused rather than truncated because two subjects sharing a prefix would become the " + + "same seat.", + nameof(subject)); + } + var expiryUnix = expiry.ToUnixTimeSeconds().ToString(CultureInfo.InvariantCulture); - var signature = HMACSHA256.HashData(signingKey, Encoding.ASCII.GetBytes(expiryUnix)); - return $"{expiryUnix}.{Base64UrlEncode(signature)}"; + var payload = string.IsNullOrEmpty(subject) + ? expiryUnix + : $"{expiryUnix}.{Base64UrlEncode(Encoding.UTF8.GetBytes(subject))}"; + + /* ASCII, as before: the payload is an integer and base64url by construction, so every byte is in + range, and keeping the encoding means the subjectless cookie is identical to the one this minted + before the subject existed. */ + var signature = HMACSHA256.HashData(signingKey, Encoding.ASCII.GetBytes(payload)); + return $"{payload}.{Base64UrlEncode(signature)}"; } /// @@ -593,20 +916,50 @@ internal static string BuildSessionCookieValue(byte[] signingKey, DateTimeOffset /// The signature compare is constant-time. /// internal static bool TryValidateSessionCookie(string? cookieValue, byte[] signingKey, DateTimeOffset now) + => TryValidateSessionCookie(cookieValue, signingKey, now, out _); + + /// + /// PURE verify, additionally reporting WHO the cookie says is holding it (#2550). + /// + /// is null for the shared-token seat, which genuinely has no identity — + /// distinct from an empty string, which would read as an authenticated principal with a blank name. Every + /// caller that stamps provenance has to keep those apart, because "the shared token did this" and "a + /// signed-in person we failed to name did this" are different facts. + /// + /// The subject is decoded only AFTER the HMAC verifies, so nothing derived from an unauthenticated + /// cookie ever reaches a caller. + /// + internal static bool TryValidateSessionCookie(string? cookieValue, byte[] signingKey, DateTimeOffset now, out string? subject) { + subject = null; + if (string.IsNullOrEmpty(cookieValue)) { return false; } - var dot = cookieValue.IndexOf('.'); - if (dot <= 0 || dot >= cookieValue.Length - 1) + /* LAST dot, not the first: the signature is the final segment and the signed payload is everything + before it, which is the rule that lets the subjectless and subject-bearing shapes share one parse. + For a subjectless cookie the last dot IS the first, so this is the original behavior unchanged. */ + var lastDot = cookieValue.LastIndexOf('.'); + if (lastDot <= 0 || lastDot >= cookieValue.Length - 1) { return false; } - var expiryPart = cookieValue.Substring(0, dot); - var signaturePart = cookieValue.Substring(dot + 1); + var payload = cookieValue.Substring(0, lastDot); + var signaturePart = cookieValue.Substring(lastDot + 1); + + var firstDot = payload.IndexOf('.'); + var expiryPart = firstDot < 0 ? payload : payload.Substring(0, firstDot); + var subjectPart = firstDot < 0 ? null : payload.Substring(firstDot + 1); + + /* Exactly two or three segments. A fourth would leave a dot inside the subject segment, and accepting + it would mean two different cookies could parse to the same subject. */ + if (subjectPart is not null && (subjectPart.Length == 0 || subjectPart.Contains('.'))) + { + return false; + } if (!long.TryParse(expiryPart, NumberStyles.None, CultureInfo.InvariantCulture, out var expiryUnix)) { @@ -628,8 +981,28 @@ internal static bool TryValidateSessionCookie(string? cookieValue, byte[] signin return false; } - var expected = HMACSHA256.HashData(signingKey, Encoding.ASCII.GetBytes(expiryPart)); - return CryptographicOperations.FixedTimeEquals(presented, expected); + var expected = HMACSHA256.HashData(signingKey, Encoding.ASCII.GetBytes(payload)); + if (!CryptographicOperations.FixedTimeEquals(presented, expected)) + { + return false; + } + + if (subjectPart is not null) + { + try + { + subject = Encoding.UTF8.GetString(Base64UrlDecode(subjectPart)); + } + catch (FormatException) + { + /* Signed by us and still undecodable means we minted it wrong, not that a caller tampered. + Refusing is still right: serving a session whose identity we cannot read would put an + unnamed principal behind the write paths this exists to attribute. */ + return false; + } + } + + return true; } private static void AppendSessionCookie(HttpContext context, byte[] signingKey) @@ -639,9 +1012,13 @@ private static void AppendSessionCookie(HttpContext context, byte[] signingKey) { HttpOnly = true, SameSite = SameSiteMode.Strict, - /* The dashboard endpoint is plain HTTP (no TLS on this surface); a Secure cookie would never be - sent, so it must be false here. The reverse-proxy story is the same as MCP for on-wire secrecy. */ - Secure = false, + /* PER-REQUEST, not a fixed value, because since #2562 this ONE host can serve both schemes at once: + the network listener is TLS when web.network.tls is configured while the loopback listeners + deliberately stay plain HTTP. A hardcoded true would mint a cookie the loopback browser then + refuses to send back (the login would loop forever); a hardcoded false would let a cookie issued + over TLS be replayed on any http:// downgrade. IsHttps is the connection's own answer, so each + cookie is marked for the transport it was actually issued on. */ + Secure = context.Request.IsHttps, Path = "/", MaxAge = SessionLifetime, }); diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfCallToolFilter.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfCallToolFilter.cs new file mode 100644 index 000000000..f48bbd40c --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfCallToolFilter.cs @@ -0,0 +1,44 @@ +using System.Collections.Generic; +using ModelContextProtocol.Protocol; +using ModelContextProtocol.Server; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +// A single call-tool filter that offers each tool's result as GCF instead of JSON when +// DARLING_OUTPUT_FORMAT=gcf. Registered once in DarlingMcpHostService, so it covers every +// tool with no per-tool changes. The re-encode is conservative (never larger, never lossy, +// see GcfOutput); anything it cannot faithfully shrink is returned as the original JSON. +public static class GcfCallToolFilter +{ + // The filter: run the tool, then transform its result. A filter is `next => handler`. + public static McpRequestFilter Instance => + next => async (request, cancellationToken) => Transform(await next(request, cancellationToken)); + + // Replaces a single JSON text-content block with its GCF wire when GCF is enabled and + // the wire is smaller and lossless; otherwise returns the result unchanged. Exposed for + // testing. StructuredContent (if a tool sets it) is left untouched. + public static CallToolResult Transform(CallToolResult result) + { + if (!GcfOutput.Enabled || result.Content == null || result.IsError == true) + return result; + + // Only a lone text block is re-encoded, so an image or other block sent alongside + // it is never dropped. + if (result.Content.Count != 1 || result.Content[0] is not TextContentBlock text) + return result; + + var wire = GcfOutput.TryEncode(text.Text); + if (wire == null) + return result; + + // Return a new result with only the text block replaced, rather than mutating the + // one the tool produced; the other fields are carried over unchanged. + return new CallToolResult + { + Content = new List { new TextContentBlock { Text = wire } }, + StructuredContent = result.StructuredContent, + IsError = result.IsError, + Meta = result.Meta, + }; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfOutput.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfOutput.cs new file mode 100644 index 000000000..fa6e18749 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/GcfOutput.cs @@ -0,0 +1,219 @@ +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Text.Json; +using BlackwellSystems.Gcf; + +namespace PerformanceMonitor.Darling.Service.Mcp; + +// Optional GCF (Graph Compact Format, https://gcformat.com) output for the MCP tool +// results. When DARLING_OUTPUT_FORMAT=gcf, a call-tool filter (GcfCallToolFilter) +// re-encodes each tool's JSON result as a GCF generic wire: the repeated field names of +// the record arrays these tools return (blocking events, alerts, wait stats, config, ...) +// are factored into a single header, cutting the token cost of a record-heavy result by +// roughly a quarter to a half of the server's compact JSON depending on shape (uniform +// numeric records win most; mostly-text results win least). Opt-in, lossless, and never +// larger than the JSON. +public static class GcfOutput +{ + // True when GCF output is requested. Read from the environment on each call so it can + // be toggled per process (or per test) without a restart. + public static bool Enabled => + string.Equals( + Environment.GetEnvironmentVariable("DARLING_OUTPUT_FORMAT")?.Trim(), + "gcf", + StringComparison.OrdinalIgnoreCase + ); + + // Returns a GCF wire for the given JSON, or null to keep the JSON. Null is returned + // whenever the JSON does not parse, contains a number GCF cannot carry exactly (a + // non-integer beyond double precision, e.g. a high-precision decimal), GCF is not + // smaller than the JSON (never-grow guard), or the decoded wire does not equal the input + // (fail-safe), so enabling GCF never grows, drops, or garbles a tool result. + public static string? TryEncode(string json) + { + if (string.IsNullOrEmpty(json)) + return null; + + object? native; + try + { + using var doc = JsonDocument.Parse(json); + native = FromJson(doc.RootElement); + } + catch + { + return null; + } + + string wire; + try + { + wire = Gcf.EncodeGeneric(native); + } + catch + { + return null; + } + + // Never-grow guard: only offer GCF when it is actually smaller than the JSON the + // tool would otherwise return. + if (wire.Length >= json.Length) + return null; + + // Fail-safe: verify the wire against the INPUT, not against itself. Decode the wire + // back to a value and require it to equal the model the tool's JSON parsed to + // (`native`). FromJson has already declined any number the wire could not carry + // exactly, so a match here means the JSON survives the full JSON -> GCF -> value + // round-trip. Object key order may normalize to header order (semantically equal for + // JSON objects), so the key comparison is order-insensitive. + try + { + if (!ValuesEqual(Gcf.DecodeGeneric(wire), native)) + return null; + } + catch + { + return null; + } + + return wire; + } + + // Order-insensitive structural equality over the gcf-dotnet model (OrderedMap / List / + // long / double / string / bool / null), used to confirm a decoded wire equals the + // input model. + private static bool ValuesEqual(object? a, object? b) + { + if (a is null || b is null) + return a is null && b is null; + + if (a is OrderedMap ma && b is OrderedMap mb) + { + if (ma.Count != mb.Count) + return false; + foreach (var key in ma.Keys) + { + if (!mb.TryGetValue(key, out var vb) || !ValuesEqual(ma[key], vb)) + return false; + } + return true; + } + + if (a is List la && b is List lb) + { + if (la.Count != lb.Count) + return false; + for (var i = 0; i < la.Count; i++) + if (!ValuesEqual(la[i], lb[i])) + return false; + return true; + } + + if (a is string sa && b is string sb) + return sa == sb; + if (a is bool ba && b is bool bb) + return ba == bb; + if (IsNumber(a) && IsNumber(b)) + return NumbersEqual(a!, b!); + return false; + } + + private static bool IsNumber(object? v) => v is long || v is double; + + // long/long compare exactly; a long and an integer-valued double (an integer can decode + // as either) compare by value. Precision-lossy numbers never reach here: FromJson + // declined them before the wire was produced; the mixed branch still avoids widening the + // long to double so the guard cannot itself launder a value above 2^53. + private static bool NumbersEqual(object a, object b) + { + if (a is long al && b is long bl) + return al == bl; + if (a is double ad && b is double bd) + return ad.Equals(bd); + + // Mixed long/double. Compare without casting the long to double (that cast rounds + // above 2^53 and would let a lost value compare equal). The two are equal only when + // the double is integral, sits inside the long range, and equals the long exactly. + long lng; + double dbl; + if (a is long la) + { + lng = la; + dbl = (double)b; + } + else + { + lng = (long)b; + dbl = (double)a; + } + return dbl == Math.Floor(dbl) + && dbl >= long.MinValue + && dbl <= long.MaxValue + && (long)dbl == lng; + } + + // Converts a parsed JSON value into the gcf-dotnet native model (OrderedMap / List / + // scalars), preserving object key order. Integers are kept as long rather than double + // so large ids, counts, and durations are never float-rounded. + private static object? FromJson(JsonElement e) + { + switch (e.ValueKind) + { + case JsonValueKind.Object: + var map = new OrderedMap(); + foreach (var p in e.EnumerateObject()) + map.Add(p.Name, FromJson(p.Value)); + return map; + + case JsonValueKind.Array: + var list = new List(); + foreach (var item in e.EnumerateArray()) + list.Add(FromJson(item)); + return list; + + case JsonValueKind.String: + return e.GetString(); + + case JsonValueKind.Number: + if (e.TryGetInt64(out var l)) + return l; + // A non-integer is carried on the wire as an IEEE-754 double (SPEC 2.3.2). + // Keep it only when the double holds the JSON token exactly; otherwise + // decline the whole payload (this throw is caught in TryEncode and the tool + // result stays JSON) rather than emit a wire that has silently dropped + // precision. A token outside the decimal range is inherently double-domain. + var d = e.GetDouble(); + if (!NumberSurvivesAsDouble(e, d)) + throw new NotSupportedException("number not exactly representable as a double"); + return d; + + case JsonValueKind.True: + return true; + + case JsonValueKind.False: + return false; + + default: + return null; // Null / Undefined + } + } + + // Reports whether a double holds the JSON number token exactly, so a non-integer can + // be carried on the wire without silently dropping precision. A non-finite double (a + // token that overflowed to +/-Infinity, e.g. 1e400) never represents a finite token and + // is declined. Otherwise the double is exact when its shortest round-trip form + // reproduces the token, or, for a token in the decimal range, when the decimal it parsed + // to is unchanged. The shortest round-trip comparison is what makes a 16-significant-digit + // value such as 0.5029000043869019 (exactly representable, but 16 digits) survive: a + // (decimal)d cast keeps only 15 digits and would wrongly reject it. + private static bool NumberSurvivesAsDouble(JsonElement e, double d) + { + if (!double.IsFinite(d)) + return false; + if (d.ToString("R", CultureInfo.InvariantCulture) + .Equals(e.GetRawText(), StringComparison.Ordinal)) + return true; + return e.TryGetDecimal(out var exact) && (decimal)d == exact; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj b/Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj index 8a1f4f119..79f792fe5 100644 --- a/Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj +++ b/Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj @@ -5,7 +5,7 @@ disable PerformanceMonitor.Darling.Service PerformanceMonitor.Darling.Service - 3.5.0 + 3.6.0 + @@ -56,11 +58,16 @@ - + + diff --git a/Darling/PerformanceMonitor.Darling.Service/Program.cs b/Darling/PerformanceMonitor.Darling.Service/Program.cs index b46c94ba3..0cc76e5f5 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Program.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Program.cs @@ -93,6 +93,33 @@ network endpoint (darling-network-endpoints D8). It DPAPI-decrypts the network r return await DarlingCliCommands.PrintViewerConnectionAsync(configPath, Console.Out, Console.Error, CancellationToken.None); } +/* CLI verbs: --print-mcp-token / --print-web-token (#2479 item 2) — reprint an endpoint's access token from + darling.json. --configure-network shows each generated token once; losing it used to leave regeneration as + the only path, which invalidates every client already configured against it. The verbs refuse when not + elevated, and both are Windows-only for the same reason --print-viewer-connection is: the token is a DPAPI + blob. The reasoning about what this does and does NOT disclose lives with the implementation. */ +if (args.Length > 0 && DarlingCliCommands.IsPrintMcpTokenVerb(args[0])) +{ + if (!OperatingSystem.IsWindows()) + { + Console.Error.WriteLine("--print-mcp-token requires Windows (DPAPI)."); + return 1; + } + + return DarlingCliCommands.PrintMcpToken(args.Length > 1 ? args[1] : null, Console.Out, Console.Error); +} + +if (args.Length > 0 && DarlingCliCommands.IsPrintWebTokenVerb(args[0])) +{ + if (!OperatingSystem.IsWindows()) + { + Console.Error.WriteLine("--print-web-token requires Windows (DPAPI)."); + return 1; + } + + return DarlingCliCommands.PrintWebToken(args.Length > 1 ? args[1] : null, Console.Out, Console.Error); +} + /* CLI verb: --export-viewer-config (#1953) — write the viewer machine's whole handoff folder (a complete darling.json with "managed": false and the resolved connection string, the store's server.crt beside it, and a README.txt documenting every field) instead of making the operator hand-merge --print-viewer-connection's @@ -185,8 +212,9 @@ taking ownership needs a privilege a virtual service account is not granted — { if (!OperatingSystem.IsWindows()) { - Console.Error.WriteLine("--enable-mcp requires Windows (DPAPI + firewall)."); - return 1; + /* #2626: the refusal names the path that WORKS on this host, not just the platform it is not. */ + return DarlingCliCommands.WriteEndpointVerbPlatformRefusal( + isMcp: true, enable: true, args.Length > 1 ? args[1] : null, Console.Error); } var configPath = args.Length > 1 ? args[1] : null; @@ -197,8 +225,9 @@ taking ownership needs a privilege a virtual service account is not granted — { if (!OperatingSystem.IsWindows()) { - Console.Error.WriteLine("--disable-mcp requires Windows (DPAPI + firewall)."); - return 1; + /* #2626: the refusal names the path that WORKS on this host, not just the platform it is not. */ + return DarlingCliCommands.WriteEndpointVerbPlatformRefusal( + isMcp: true, enable: false, args.Length > 1 ? args[1] : null, Console.Error); } var configPath = args.Length > 1 ? args[1] : null; @@ -209,8 +238,9 @@ taking ownership needs a privilege a virtual service account is not granted — { if (!OperatingSystem.IsWindows()) { - Console.Error.WriteLine("--enable-web requires Windows (DPAPI + firewall)."); - return 1; + /* #2626: the refusal names the path that WORKS on this host, not just the platform it is not. */ + return DarlingCliCommands.WriteEndpointVerbPlatformRefusal( + isMcp: false, enable: true, args.Length > 1 ? args[1] : null, Console.Error); } var configPath = args.Length > 1 ? args[1] : null; @@ -221,8 +251,9 @@ taking ownership needs a privilege a virtual service account is not granted — { if (!OperatingSystem.IsWindows()) { - Console.Error.WriteLine("--disable-web requires Windows (DPAPI + firewall)."); - return 1; + /* #2626: the refusal names the path that WORKS on this host, not just the platform it is not. */ + return DarlingCliCommands.WriteEndpointVerbPlatformRefusal( + isMcp: false, enable: false, args.Length > 1 ? args[1] : null, Console.Error); } var configPath = args.Length > 1 ? args[1] : null; diff --git a/Darling/PerformanceMonitor.Darling.Service/QueryStoreBackfill.cs b/Darling/PerformanceMonitor.Darling.Service/QueryStoreBackfill.cs index c45380df1..453c26c4a 100644 --- a/Darling/PerformanceMonitor.Darling.Service/QueryStoreBackfill.cs +++ b/Darling/PerformanceMonitor.Darling.Service/QueryStoreBackfill.cs @@ -25,7 +25,7 @@ namespace PerformanceMonitor.Darling.Service; /// live path never takes. Phase 1 made the LIVE path hole-free, but two bounded windows still /// discard history by design: first contact takes only the trailing 60 minutes of a ~30-day /// catalog, and post-outage catch-up is clamped to (the -/// #1556 incident fix, tightened to 1h by #2102) as a bounded, logged hole. One mechanism fills +/// #1556 incident fix, tightened by #2102) as a bounded, logged hole. One mechanism fills /// both, and every slice of it windows at most /// at a time (#2102 — the query's cost grows with window width, so an unchunked wide range on a /// big database re-times-out forever instead of draining): @@ -40,7 +40,7 @@ namespace PerformanceMonitor.Darling.Service; /// is the #1960 design constraint: the two paths can never race for the same boundary. /// /// Clamp holes (post-outage). An interior gap is invisible to MIN/MAX, so the runner -/// records it at the moment the 24h clamp fires — (raw watermark, clamped floor), merged wider on +/// records it at the moment the catch-up clamp fires — (raw watermark, clamped floor), merged wider on /// a repeat clamp — under this worker's own collector_state rows /// (; deliberately NOT the definition's StateKeys machinery, so /// query_store itself still declares none). The worker services the hole newest-first and shrinks diff --git a/Darling/PerformanceMonitor.Darling.Service/StoreConfigProvider.cs b/Darling/PerformanceMonitor.Darling.Service/StoreConfigProvider.cs index 1c79789c7..c2980e6cd 100644 --- a/Darling/PerformanceMonitor.Darling.Service/StoreConfigProvider.cs +++ b/Darling/PerformanceMonitor.Darling.Service/StoreConfigProvider.cs @@ -113,9 +113,11 @@ public async Task SeedIfEmptyAsync(DarlingConfig config, CancellationToken cance { /* #2254: the seed is skipped, so any server added to darling.json AFTER the first start is silently ignored — and --test-connection reads the FILE, so it validates those servers - happily while the service never monitors them. Say so once per start instead of leaving the - operator to discover it. */ - await WarnAboutFileOnlyServersAsync(connection, config, cancellationToken); + happily while the service never monitors them. #2552: the same skip makes every per-server + SETTING in the file dead text for a server that IS registered, which was silent for the + same reason and is the more common edit. Say both once per start instead of leaving the + operator to discover them. */ + await WarnAboutFileVersusStoreAsync(connection, config, cancellationToken); } /* LAST — its presence marks the seed complete (the reload gate keys on config_version). */ @@ -133,46 +135,79 @@ operator to discover it. */ } /// - /// #2254: names the servers present in darling.json that the store does not have, once per start. + /// #2254 + #2552: the once-per-start report on where darling.json and the registry disagree. Two causes, + /// reported separately because they have two different remedies. /// - /// The seed runs only while config_monitored_servers is empty, so a server added to the file - /// after the first successful start is a permanent no-op. What made that expensive in the field is that - /// --test-connection reads the FILE and validated the new server as PASS, so the operator had two - /// outputs that were each correct about different things and no way to see the disagreement: config edit, - /// service restart, support round trip. + /// Cause A — a server the store has never had (#2254). The seed runs only while + /// config_monitored_servers is empty, so a server added to the file after the first successful + /// start is a permanent no-op. What made that expensive in the field is that --test-connection + /// reads the FILE and validated the new server as PASS, so the operator had two outputs that were each + /// correct about different things and no way to see the disagreement: config edit, service restart, + /// support round trip. /// - /// Compared on server_id OR name (#2158). It used to be id alone, on the grounds that the id - /// is what the collectors key on — correct while every row's id equalled the hash of its own address, and - /// wrong the moment an edit began PRESERVING a row's identity so a re-addressed server keeps its history. - /// After such an edit the file's derived id matches nothing, and an id-only comparison would report a - /// server that IS monitored as absent, then advise re-adding it — wrong advice, on every start, about the - /// one server the operator had just fixed. The name arm covers that; a genuinely removed server is gone - /// from the store under both keys, so the Viewer-Remove case still reports exactly as before. + /// Cause B — a server the store HAS, whose settings the file disagrees with (#2552). Cause + /// A's warning made this one WORSE rather than better: it teaches the operator "adding a server to the + /// file does not register it", and the natural inference from that is that the file still drives the + /// servers the store already knows about. It does not — for a REGISTERED server every per-server setting + /// in darling.json is dead text. The field report was a PostgreSQL target refusing a self-signed + /// certificate: the operator applied the documented fix ("trustServerCertificate": true), + /// restarted, and got a BYTE-IDENTICAL error, because the store's row still said false and nothing + /// anywhere compared the two. That is the worst loop to leave open — a connection failure is exactly the + /// class of problem an operator fixes by editing config and restarting, and this was the one class of + /// edit that produced an unchanged error with no explanation. + /// + /// Cause B does NOT make the file authoritative and changes none of the ordering: the store still + /// wins, deliberately (#2254). The defect is that the disagreement was INVISIBLE. + /// + /// Both causes are matched on server_id OR name (#2158). It used to be id alone, on the + /// grounds that the id is what the collectors key on — correct while every row's id equalled the hash of + /// its own address, and wrong the moment an edit began PRESERVING a row's identity so a re-addressed + /// server keeps its history. After such an edit the file's derived id matches nothing, and an id-only + /// comparison would report a server that IS monitored as absent, then advise re-adding it — wrong advice, + /// on every start, about the one server the operator had just fixed. The name arm covers that; a genuinely + /// removed server is gone from the store under both keys, so the Viewer-Remove case still reports exactly + /// as before. /// - private async Task WarnAboutFileOnlyServersAsync( + private async Task WarnAboutFileVersusStoreAsync( NpgsqlConnection connection, DarlingConfig config, CancellationToken ct) { + var storeServers = await ReadRegisteredServersForComparisonAsync(connection, ct); + + /* Cause A asks "is this file entry REGISTERED at all", which a control-plane pause does not change — + so these sets are built from every row, disabled included. Filtering them would report a paused + server as never-monitored and advise re-adding it, which is the #2158 failure this method was + already fixed once to avoid. The enabled/disabled split belongs to the DRIFT pass, which asks a + different question; see WarnAboutSettingDrift. */ var storeIds = new HashSet(); var storeNames = new HashSet(StringComparer.OrdinalIgnoreCase); - using (var command = new NpgsqlCommand("SELECT server_id, name FROM config_monitored_servers", connection)) - await using (var reader = await command.ExecuteReaderAsync(ct)) + foreach (var stored in storeServers) { - while (await reader.ReadAsync(ct)) + storeIds.Add(stored.Config.ServerId); + if (!string.IsNullOrEmpty(stored.Config.Name)) { - storeIds.Add(reader.GetInt32(0)); - if (!reader.IsDBNull(1)) - { - storeNames.Add(reader.GetString(1)); - } + storeNames.Add(stored.Config.Name); } } var fileOnly = ServersOnlyInFile(config.Servers, storeIds, storeNames); - if (fileOnly.Count == 0) + if (fileOnly.Count > 0) { - return; + await WarnAboutFileOnlyServersAsync(connection, fileOnly, ct); } + /* Runs unconditionally rather than behind the file-only early return that used to end this method: + the #2552 case is precisely the one where fileOnly is EMPTY — every server in the file is already + registered — so returning there is what made a registered server's settings drift silent. */ + WarnAboutSettingDrift(config.Servers, storeServers); + } + + /// + /// Cause A's two log lines (#2252 / #2258), split out from the caller so the drift pass is not gated on + /// there being anything to say here. + /// + private async Task WarnAboutFileOnlyServersAsync( + NpgsqlConnection connection, IReadOnlyList fileOnly, CancellationToken ct) + { /* #2258: the OBSERVED registry is the tombstone, and it already exists. collect.servers gets a row upserted on every successful connect, and the Viewer's Remove deletes only from config_monitored_servers (the DESIRED config) — so a row surviving there means "this server really @@ -229,6 +264,142 @@ Information and says so plainly rather than advising anything. It is not silent } } + /// + /// Cause B's line (#2552): names every registered server whose darling.json entry disagrees with its + /// registry row, and every field it disagrees about, with both values. + /// + /// The remedy names the Viewer and explicitly rules OUT add_servers. That tool cannot do + /// this: it partitions an already-monitored server out as status:"duplicate" before it validates + /// anything, so pointing an operator at it for a REGISTERED server would send them somewhere that + /// silently does nothing — the exact failure this warning exists to end. The Viewer's own duplicate + /// message ("Edit it from Manage Servers instead") already says where the edit lives. + /// + /// The --test-connection caveat is appended only when a CONNECTION-relevant field drifted. + /// The verb probes the file's connection settings, so it can report PASS for a connection the service + /// will never make — but it does not exercise the display name, the excluded-database list, the cost + /// figure or the delivery override, and a warning must not claim more than it can support. + /// + /// A DISABLED server gets its own line, at Information — the same two-line shape cause A + /// already uses for "never monitored" (a warning) versus "deliberately removed" (information). Raised in + /// review on #2556, and the reviewer's first suggestion — filter the read to is_enabled = TRUE — + /// is the one option that must not be taken: that read also feeds cause A, which asks whether a file entry + /// is REGISTERED, and a paused server is. Filtering there would report it as never-monitored and advise + /// re-adding it. Suppressing it from the drift pass only was the other candidate and is rejected too, more + /// narrowly: #2552 is a defect about SILENCE being expensive, so answering it by adding a new silence is + /// the wrong direction, and the drift is exactly what the operator will walk back into when they + /// re-enable. What was genuinely wrong is the CLAIM — "the registry is what the service uses" is not true + /// of a server nothing is connecting to — so the disabled line drops it, drops the remedy, and drops the + /// --test-connection caveat, saying only what is true of a paused server. + /// + private void WarnAboutSettingDrift( + IReadOnlyList? fileServers, IReadOnlyList storeServers) + { + var drifted = DescribeSettingDrift(fileServers, storeServers); + if (drifted.Count == 0) + { + return; + } + + var live = drifted.Where(d => d.IsEnabled).ToList(); + var paused = drifted.Where(d => !d.IsEnabled).ToList(); + + if (live.Count > 0) + { + var connectionCaveat = live.Any(d => d.Fields.Any(f => f.AffectsConnection)) + ? " Note --test-connection reads darling.json, so it probes the FILE's settings and can report " + + "PASS for a connection the service will never make." + : ""; + + _logger?.LogWarning( + "darling.json disagrees with the registry about {Count} monitored server(s), and the registry is " + + "what the service uses: {Details}. The store is authoritative after the first seed, so editing a " + + "registered server's settings in the file changes nothing and a restart cannot change that — " + + "edit them in the Viewer's Manage Servers window (the MCP add_servers tool cannot: an " + + "already-registered server is skipped as a duplicate).{ConnectionCaveat}", + live.Count, + FormatSettingDrift(live, MaxDriftedServersLogged), + connectionCaveat); + } + + if (paused.Count > 0) + { + _logger?.LogInformation( + "darling.json also disagrees with the registry about {Count} server(s) that are registered but " + + "DISABLED: {Details}. Nothing is connecting to them, so neither value is in force today — but " + + "the registry's is the one that would take effect if they were re-enabled from the Viewer, not " + + "the file's.", + paused.Count, + FormatSettingDrift(paused, MaxDriftedServersLogged)); + } + } + + /// + /// How many drifted servers the log line NAMES. The count in the message is always the true total — this + /// is a display budget, like the alert renderer's incident cap, and nothing derives state from the + /// truncated list. A whole-fleet drift (a regenerated darling.json against a 42-server registry) is the + /// realistic way this line becomes unreadable, and an unreadable warning is a silent one. + /// + private const int MaxDriftedServersLogged = 10; + + /// + /// The registry rows as s, for both halves of the report above. + /// + /// encrypted_password is not in the SELECT list, and that is the credential guarantee. + /// #2552 requires that no credential is ever compared or printed, and the way to guarantee that is + /// structural rather than editorial: the blob is never read, so there is nothing in memory for a later + /// edit to this file to leak into a log line. It must not be compared even if it were read — a file entry + /// legitimately carries a file:/env: reference () or a dev + /// plaintext password while the store row carries a DPAPI blob, and + /// backfills exactly that pairing at read time. That is the SUPPORTED shape, so comparing them would + /// report a working configuration as drift on every start. + /// + /// capture_plans is left unread for the opposite reason: it has no per-server darling.json + /// counterpart to disagree with (capturePlans is a top-level service setting and the seed writes + /// the per-server column NULL), so there is nothing to compare. is_enabled is read but never + /// COMPARED — it has no file counterpart either; it decides which of the two drift lines a server belongs + /// on. Every row is returned, enabled or not, because cause A's question is about registration. + /// + private static async Task> ReadRegisteredServersForComparisonAsync( + NpgsqlConnection connection, CancellationToken ct) + { + var servers = new List(); + using var command = new NpgsqlCommand(@" +SELECT server_id, name, host, database, auth, username, encrypt_mode, trust_server_certificate, + read_only_intent, multi_subnet_failover, excluded_databases, monthly_cost_usd, + alert_delivery_mode_override, engine, port, is_enabled +FROM config_monitored_servers", connection); + await using var reader = await command.ExecuteReaderAsync(ct); + while (await reader.ReadAsync(ct)) + { + servers.Add(new RegisteredServer( + new MonitoredServer + { + /* The row's own primary key, so ServerId resolves to it rather than re-deriving the hash — + the same reason BuildServerFromRow reads it (#2218). */ + StoredServerId = reader.GetInt32(0), + Name = reader.IsDBNull(1) ? "" : reader.GetString(1), + Host = reader.IsDBNull(2) ? "" : reader.GetString(2), + Database = reader.IsDBNull(3) ? null : reader.GetString(3), + Auth = reader.IsDBNull(4) ? "integrated" : reader.GetString(4), + Username = reader.IsDBNull(5) ? null : reader.GetString(5), + EncryptMode = reader.IsDBNull(6) ? "Mandatory" : reader.GetString(6), + TrustServerCertificate = !reader.IsDBNull(7) && reader.GetBoolean(7), + ReadOnlyIntent = !reader.IsDBNull(8) && reader.GetBoolean(8), + MultiSubnetFailover = !reader.IsDBNull(9) && reader.GetBoolean(9), + ExcludedDatabases = ReadTextArray(reader, 10), + MonthlyCostUsd = reader.IsDBNull(11) ? 0m : reader.GetDecimal(11), + AlertDeliveryModeOverride = ParseDeliveryOverride(reader.IsDBNull(12) ? null : reader.GetString(12)), + Engine = reader.IsDBNull(13) ? "sqlserver" : reader.GetString(13), + Port = reader.IsDBNull(14) ? 0 : reader.GetInt32(14), + }, + /* NOT NULL DEFAULT TRUE in the table, so the guard is for a store mid-migration; an unknown + enablement reads as ENABLED, which is the direction that keeps the drift visible. */ + IsEnabled: reader.IsDBNull(15) || reader.GetBoolean(15))); + } + + return servers; + } + /// /// Splits the file-only servers into "never monitored" and "monitored once, since removed" (#2258), using the /// observed registry as the evidence. @@ -280,6 +451,410 @@ internal static (IReadOnlyList NeverRegistered, IReadOnlyList De return (never, removed); } + /// + /// One field on which a darling.json entry and its registry row disagree, ALREADY NORMALIZED — the two + /// values carried here are the ones that were compared, so the log line can never print a pair that + /// differs only in a way the service does not act on. + /// marks the fields the connection string is built from, which + /// is what gates the --test-connection caveat on the warning. + /// + internal readonly record struct SettingDrift( + string Field, string FileValue, string StoreValue, bool AffectsConnection); + + /// + /// One config_monitored_servers row: its settings, plus whether the control plane has it ENABLED. + /// Enablement is not a darling.json concept and is never compared — it only decides which of the two + /// drift lines the server belongs on, because "the registry is what the service uses" is not a true + /// sentence about a server nothing is connecting to. + /// + internal sealed record RegisteredServer(MonitoredServer Config, bool IsEnabled); + + /// One registered server and every field its darling.json entry disagrees with (#2552). + internal sealed record ServerSettingDrift(string Server, bool IsEnabled, IReadOnlyList Fields); + + /// + /// Pairs each darling.json entry with its registry row and reports the fields they disagree on (#2552). + /// Pure, so the whole comparison is testable without a store. + /// + /// Pairing. Same either-or key as : the stored + /// server_id first, then the display name case-folded. One rule governs both arms — the match must + /// be unambiguous in BOTH directions, or the entry is skipped rather than guessed at. Nothing enforces + /// display-name uniqueness, and two file entries can derive one server_id (identical addresses, + /// where the seed's ON CONFLICT DO NOTHING left a single row), so a guess would print two entries' + /// settings as one server's drift and could contradict itself on the same line. That is the same "exactly + /// one match" discipline applies before it copies a bootstrap secret, for + /// the same reason: a wrong pairing here is worse than no pairing. + /// + internal static IReadOnlyList DescribeSettingDrift( + IEnumerable? fileServers, IReadOnlyList? storeServers) + { + var drifted = new List(); + if (fileServers is null || storeServers is null || storeServers.Count == 0) + { + return drifted; + } + + var file = fileServers.Where(s => s is not null).ToList(); + var byId = new Dictionary(); + foreach (var stored in storeServers) + { + /* config_monitored_servers.server_id is the PRIMARY KEY, so a duplicate can only come from a + reader that produced one; last-wins rather than throwing, since this is a diagnostic. */ + byId[stored.Config.ServerId] = stored; + } + + var storeNameCounts = CountNames(storeServers.Select(s => s.Config.DisplayName)); + var fileNameCounts = CountNames(file.Select(s => s.DisplayName)); + var fileIdCounts = new Dictionary(); + foreach (var entry in file) + { + fileIdCounts[entry.ServerId] = fileIdCounts.TryGetValue(entry.ServerId, out var n) ? n + 1 : 1; + } + + foreach (var entry in file) + { + RegisteredServer? match = null; + if (byId.TryGetValue(entry.ServerId, out var byIdMatch) && fileIdCounts[entry.ServerId] == 1) + { + match = byIdMatch; + } + else if (Count(storeNameCounts, entry.DisplayName) == 1 && Count(fileNameCounts, entry.DisplayName) == 1) + { + match = storeServers.First(s => + string.Equals(s.Config.DisplayName, entry.DisplayName, StringComparison.OrdinalIgnoreCase)); + } + + if (match is null) + { + /* Either not registered at all — ServersOnlyInFile reports that, and it is a different + remedy — or an ambiguous name, which is deliberately not guessed at. */ + continue; + } + + var fields = CompareServerSettings(entry, match.Config); + if (fields.Count > 0) + { + drifted.Add(new ServerSettingDrift(match.Config.DisplayName, match.IsEnabled, fields)); + } + } + + return drifted; + } + + private static Dictionary CountNames(IEnumerable names) + { + var counts = new Dictionary(StringComparer.OrdinalIgnoreCase); + foreach (var name in names) + { + var key = name ?? ""; + counts[key] = counts.TryGetValue(key, out var n) ? n + 1 : 1; + } + + return counts; + } + + private static int Count(Dictionary counts, string? name) => + counts.TryGetValue(name ?? "", out var n) ? n : 0; + + /// + /// Field-by-field, through the SAME folds the connect path and the collectors apply — because a + /// comparison that does not normalize warns about differences that do not exist, and an operator who is + /// told twice about a non-difference stops reading the line that matters. Every fold below is a real one + /// somewhere else in the service, not a convenience: + /// + /// Every field name below is the darling.json KEY, spelled exactly. The message is about + /// darling.json, so it names what the operator would edit — and it lets + /// RegisteredServerSettingDriftTests pin the set by REFLECTION over + /// 's [JsonPropertyName] properties, with no mapping table in + /// between. A per-server setting added to the file later is therefore covered here or named by a red + /// test; it cannot go back to being silently dead text, which is the category #2552 belongs to rather + /// than the single field that was reported. + /// + /// + /// name — falls back to the host when blank (), + /// then compared case-SENSITIVELY: the store's spelling is what the Viewer and every alert render, so a + /// re-cased name in the file really is a change that is not taking effect. + /// host — trimmed, case-insensitive, matching DarlingWorker.ServerDefinitionEquals. + /// database — blank folds to the engine's implicit default (master for SQL Server, + /// postgres for PostgreSQL), which is what substitutes. So + /// a NULL column against an explicit "master" in the file is not drift, and the message prints the + /// database that is actually connected to. Compared case-SENSITIVELY on PostgreSQL and case-insensitively + /// on SQL Server, because that is what each engine does with the name: PostgreSQL matches the startup + /// packet's database against pg_database.datname byte for byte, so ReportingDB and + /// reportingdb are two databases there and folding them would MISS a real difference — the silent + /// direction, which is the one #2552 is about. On SQL Server the resolution is collation-dependent and + /// case-insensitive on every default collation, which is also what ServerDefinitionEquals assumes; + /// the stated limit is that a case-SENSITIVE server collation would hide a re-cased catalog name + /// there. + /// auth — trimmed, case-insensitive; config validation already restricts it to + /// integrated/sql. + /// username — only when the STORE row uses SQL auth, because that is the only case where the + /// connection string carries one; a stale username beside integrated auth is inert. Case-sensitive, since + /// a SQL login can be. + /// encryptMode — the connect path's own fail-closed fold (trim, upper, anything unrecognized + /// becomes Mandatory), so "strict" against "Strict" is not drift and a typo against + /// Mandatory is not either. THE hazard #2552 named. It has a SECOND, engine-shaped half that the + /// first cut missed: SQL Server really does have three behaviours (three distinct + /// SqlConnectionEncryptOption values), but the PostgreSQL builder branches on OPTIONAL + /// alone — Strict and Mandatory both land on the same SslMode — so on a PostgreSQL target they are + /// ONE bucket, and reporting them as drift would be reporting a connection difference that does not + /// exist. The comparison collapses them there; the message still prints what each side actually says, so + /// the values shown are never invented. + /// engine — folded through , so "aurora" + /// against "postgres" is one engine and not a disagreement. + /// port — PostgreSQL only (SQL Server carries its port inside the host as host,1433), + /// with 0 folded to the driver default 5432 — so an unset port against an explicit 5432 is not drift. + /// readOnlyIntent / multiSubnetFailover — SQL Server only: ApplicationIntent and + /// MultiSubnetFailover have no PostgreSQL equivalent and the Npgsql builder never sees them. + /// trustServerCertificate — everywhere EXCEPT a PostgreSQL target whose stored mode is + /// Optional, where SslMode.Prefer is chosen without consulting the flag at all, so it is inert. + /// The gate reads the STORE's mode, since the store's is the one in force. + /// excludedDatabases — trimmed, blanks dropped, de-duplicated and ORDERED, because it is + /// consumed as a NOT IN set (DatabaseExclusionFilter) where order and repetition change + /// nothing. Case follows the engine for the same reason database does, and it is live on BOTH: + /// the SQL Server collectors splice it against d.name, and PostgresTargetProvider splices + /// the same filter against pg_database.datname to choose the per-database fan-out — where + /// NOT IN is case-sensitive, so folding case would silently treat two different exclusions as + /// one. + /// monthlyCostUsd — compared numerically, so 100 against 100.00 is not drift. Not + /// connection-relevant, but it does drive the FinOps figures, so it is not cosmetic either. + /// alertDeliveryModeOverride — null means "inherit the global" (#1236) and prints as such. + /// + /// + /// The credential is absent by construction. password and encryptedPassword + /// are the ONLY two darling.json per-server keys this method deliberately does not compare. Neither is + /// read from the store (see ReadRegisteredServersForComparisonAsync) and neither is compared here, + /// so no blob, no file:/env: reference and no plaintext can reach a log line through this + /// path. + /// + /// The engine-gated fields read the STORE's engine, because the store is what the service connects + /// with — if the two disagree about the engine, that disagreement is reported on its own line and the + /// store's answer is the one in force. + /// + internal static IReadOnlyList CompareServerSettings(MonitoredServer file, MonitoredServer store) + { + var drift = new List(); + if (file is null || store is null) + { + return drift; + } + + var storeIsPostgres = store.IsPostgres; + + AddDrift(drift, "name", file.DisplayName, store.DisplayName, StringComparison.Ordinal, false); + AddDrift(drift, "host", Trimmed(file.Host), Trimmed(store.Host), StringComparison.OrdinalIgnoreCase, true); + /* PostgreSQL matches a database name byte for byte; SQL Server folds case on every default + collation. Getting this backwards on PostgreSQL would MISS a real difference, which is the + silent direction. */ + var nameComparison = storeIsPostgres ? StringComparison.Ordinal : StringComparison.OrdinalIgnoreCase; + + AddDrift( + drift, + "database", + EffectiveDatabase(file.Database, storeIsPostgres), + EffectiveDatabase(store.Database, storeIsPostgres), + nameComparison, + true); + AddDrift(drift, "auth", Trimmed(file.Auth), Trimmed(store.Auth), StringComparison.OrdinalIgnoreCase, true); + + if (store.UsesSqlAuth) + { + AddDrift( + drift, + "username", + NoneIfBlank(Trimmed(file.Username)), + NoneIfBlank(Trimmed(store.Username)), + StringComparison.Ordinal, + true); + } + + /* Displayed and compared SEPARATELY, and only here. The two are the same everywhere else, but on a + PostgreSQL target Strict and Mandatory are one connection while remaining two different words in + the file — so the comparison collapses them and the message still prints what each side says + rather than a value neither holds. */ + var fileMode = EffectiveEncryptMode(file.EncryptMode); + var storeMode = EffectiveEncryptMode(store.EncryptMode); + if (!string.Equals( + EncryptModeConnectKey(fileMode, storeIsPostgres), + EncryptModeConnectKey(storeMode, storeIsPostgres), + StringComparison.Ordinal)) + { + drift.Add(new SettingDrift("encryptMode", fileMode, storeMode, true)); + } + + AddDrift(drift, "engine", EngineToken(file), EngineToken(store), StringComparison.Ordinal, true); + + /* Inert on a PostgreSQL target whose stored mode is Optional: that branch picks SslMode.Prefer + without ever reading the flag. Live everywhere else, including SQL Server's Optional, where + SqlClient still validates the certificate if the server negotiates encryption. */ + if (!storeIsPostgres || !string.Equals(storeMode, "Optional", StringComparison.Ordinal)) + { + AddDrift( + drift, + "trustServerCertificate", + Json(file.TrustServerCertificate), + Json(store.TrustServerCertificate), + StringComparison.Ordinal, + true); + } + + if (storeIsPostgres) + { + AddDrift( + drift, + "port", + EffectivePostgresPort(file.Port), + EffectivePostgresPort(store.Port), + StringComparison.Ordinal, + true); + } + else + { + AddDrift(drift, "readOnlyIntent", Json(file.ReadOnlyIntent), Json(store.ReadOnlyIntent), StringComparison.Ordinal, true); + AddDrift( + drift, + "multiSubnetFailover", + Json(file.MultiSubnetFailover), + Json(store.MultiSubnetFailover), + StringComparison.Ordinal, + true); + } + + AddDrift( + drift, + "excludedDatabases", + NormalizeExcludedDatabases(file.ExcludedDatabases, storeIsPostgres), + NormalizeExcludedDatabases(store.ExcludedDatabases, storeIsPostgres), + nameComparison, + false); + + if (file.MonthlyCostUsd != store.MonthlyCostUsd) + { + drift.Add(new SettingDrift( + "monthlyCostUsd", + file.MonthlyCostUsd.ToString(System.Globalization.CultureInfo.InvariantCulture), + store.MonthlyCostUsd.ToString(System.Globalization.CultureInfo.InvariantCulture), + false)); + } + + AddDrift( + drift, + "alertDeliveryModeOverride", + file.AlertDeliveryModeOverride?.ToString() ?? "(inherit)", + store.AlertDeliveryModeOverride?.ToString() ?? "(inherit)", + StringComparison.Ordinal, + false); + + return drift; + } + + private static void AddDrift( + List into, string field, string fileValue, string storeValue, StringComparison comparison, bool affectsConnection) + { + if (!string.Equals(fileValue, storeValue, comparison)) + { + into.Add(new SettingDrift(field, fileValue, storeValue, affectsConnection)); + } + } + + private static string Trimmed(string? value) => value?.Trim() ?? ""; + + private static string NoneIfBlank(string value) => string.IsNullOrEmpty(value) ? "(none)" : value; + + private static string Json(bool value) => value ? "true" : "false"; + + /// The database the connection actually opens: blank means the engine's implicit default. + private static string EffectiveDatabase(string? database, bool isPostgres) => + string.IsNullOrWhiteSpace(database) ? (isPostgres ? "postgres" : "master") : database.Trim(); + + /// + /// The encrypt mode the connection is actually built with — 's + /// fail-closed fold, in canonical casing so the log prints the mode in force rather than what was typed. + /// + private static string EffectiveEncryptMode(string? mode) => mode?.Trim().ToUpperInvariant() switch + { + "STRICT" => "Strict", + "OPTIONAL" => "Optional", + _ => "Mandatory", + }; + + /// + /// How many DISTINCT connections the three modes make on this engine, which is not the same number on + /// both. SQL Server maps them to three SqlConnectionEncryptOption values, so all three are + /// separate. The Npgsql builder branches on OPTIONAL alone — everything else takes the same + /// SslMode arm — so on PostgreSQL Strict and Mandatory are ONE bucket and reporting them as drift + /// would report a difference the connection cannot express. + /// + private static string EncryptModeConnectKey(string canonicalMode, bool isPostgres) => + isPostgres && string.Equals(canonicalMode, "Strict", StringComparison.Ordinal) + ? "Mandatory" + : canonicalMode; + + private static string EngineToken(MonitoredServer server) => + server.IsPostgres ? "postgres" : "sqlserver"; + + /// 0 means "the driver's default", which for Npgsql is 5432 — so the two are one value. + private static string EffectivePostgresPort(int port) => + (port > 0 ? port : 5432).ToString(System.Globalization.CultureInfo.InvariantCulture); + + /// + /// The excluded-database list as the collectors see it: a set. DatabaseExclusionFilter splices it + /// into a NOT IN, so order, repetition and surrounding whitespace change nothing, and warning about + /// any of them would be warning about a difference that does not exist. + /// + /// Case follows the ENGINE, because the NOT IN is evaluated by it: PostgreSQL compares + /// datname byte for byte (the list picks the per-database fan-out in + /// PostgresTargetProvider.BuildDatabaseListPlan), while SQL Server's d.name comparison folds + /// case on every default collation. Folding case on PostgreSQL would treat two genuinely different + /// exclusions as one, which is the silent direction. + /// + private static string NormalizeExcludedDatabases(IEnumerable? names, bool isPostgres) + { + if (names is null) + { + return "(none)"; + } + + var comparer = isPostgres ? StringComparer.Ordinal : StringComparer.OrdinalIgnoreCase; + var set = names + .Where(n => !string.IsNullOrWhiteSpace(n)) + .Select(n => n.Trim()) + .Distinct(comparer) + .OrderBy(n => n, comparer) + .ToList(); + + return set.Count == 0 ? "(none)" : string.Join(", ", set); + } + + /// + /// Renders the drift for the log line: server: field (file=X, store=Y), field (…), servers joined + /// by | . Truncated to with the remainder counted rather than + /// dropped silently. Pure, so the exact operator-facing text is pinned by a test. + /// + internal static string FormatSettingDrift(IReadOnlyList drifted, int maxServers) + { + if (drifted is null || drifted.Count == 0) + { + return ""; + } + + var shown = maxServers > 0 && drifted.Count > maxServers ? maxServers : drifted.Count; + var parts = new List(shown + 1); + for (int i = 0; i < shown; i++) + { + var server = drifted[i]; + var fields = server.Fields.Select(f => $"{f.Field} (file={f.FileValue}, store={f.StoreValue})"); + parts.Add($"{server.Server}: {string.Join(", ", fields)}"); + } + + if (shown < drifted.Count) + { + parts.Add($"and {drifted.Count - shown} more not listed"); + } + + return string.Join(" | ", parts); + } + /// /// The file entries whose server_id is absent from the store, by display name. Pure so the /// comparison is testable without a store — the log line above is the only part that needs one. @@ -335,7 +910,7 @@ private static async Task SeedServiceRowAsync(NpgsqlConnection connection, Darli so the worker's post-seed baseline read reflects the seeded state and triggers no spurious reload. */ using var command = new NpgsqlCommand(@" INSERT INTO config_service (id, paused, capture_plans, query_store_backfill_enabled, query_store_text_budget_mb, max_concurrent_sweeps, plan_xml_compression, mcp_enabled, mcp_port, web_enabled, web_port, plan_content_retention_days, compose_statement_timeout_seconds, config_version, updated_at, updated_by) -VALUES (1, FALSE, $1, $7, $8, $9, $10, $2, $3, $4, $5, $11, 0, $6, 'seed') +VALUES (1, FALSE, $1, $7, $8, $9, $10, $2, $3, $4, $5, $11, $12, 0, $6, 'seed') ON CONFLICT (id) DO NOTHING", connection); command.Parameters.AddWithValue(config.CapturePlans); command.Parameters.AddWithValue(config.Mcp.Enabled); diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/HypotheticalIndexExperiment.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/HypotheticalIndexExperiment.cs new file mode 100644 index 000000000..dc16f8497 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/HypotheticalIndexExperiment.cs @@ -0,0 +1,289 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Data.Common; +using System.Globalization; +using System.Text.Json; +using System.Text.Json.Nodes; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; + +namespace PerformanceMonitor.Darling.Service.Targets; + +/// +/// Plans one statement twice — once as the server would today, once with a hypothetical index visible — +/// and reports whether the planner would switch (#2612). +/// +/// +/// This is the one place the product ACTS on a monitored PostgreSQL server rather than reading it, so the +/// blast radius is worth stating precisely rather than reassuringly. A hypothetical index costs no disk, +/// is visible only inside the session that created it, and is never written anywhere. +/// EXPLAIN without ANALYZE does not execute the statement. The whole experiment runs inside a +/// transaction that is ROLLED BACK, so even the session-local catalog entry does not outlive the call. +/// +/// +/// +/// GENERIC_PLAN is what makes this possible at all. pg_stat_statements stores +/// NORMALIZED text — literals replaced by $1, $2 — which cannot be planned the ordinary way +/// without values nobody has. PostgreSQL 16 added EXPLAIN (GENERIC_PLAN) for exactly this, so the +/// statement is planned as the server would plan it before seeing parameters. On PostgreSQL 15 and older +/// there is no such option and the experiment refuses rather than guessing at values, because inventing a +/// parameter would produce a plan for a query nobody ran. +/// +/// +/// +/// A "no" is a result, not a failure. The planner declining the candidate is the answer that saves +/// someone from building an index — measured on the verification rig, where a candidate on an already +/// well-served predicate left the plan and its cost completely unchanged. The report says so plainly +/// rather than presenting an unchanged cost as an inconclusive run. +/// +/// +public static class HypotheticalIndexExperiment +{ + /// + /// The minimum PostgreSQL major for EXPLAIN (GENERIC_PLAN). Below it the experiment refuses. + /// + public const int MinimumPostgresMajorForGenericPlan = 16; + + /// How long either EXPLAIN may take. Planning is cheap; a planner that is not is a finding of + /// its own, and one this call must not sit inside on a server somebody else is using. + public const int StatementTimeoutSeconds = 15; + + /// The decisive answer. False is a real result. + /// Total estimated cost of the plan the server would use today. + /// Total estimated cost with the candidate visible. Equal to + /// when the planner declined it. + /// What hypopg called the candidate, so the name in the plan can be + /// matched to it. Null when creation itself failed. + public readonly record struct Result( + bool PlannerWouldUseIt, + double CostBefore, + double CostAfter, + string? HypotheticalIndexName, + string? PlanBeforeJson, + string? PlanAfterJson, + string Explanation); + + /// + /// Runs the experiment on an OPEN connection to the monitored server. + /// + /// The caller owns the connection because the caller owns the decision about which server this + /// runs against, and that decision is the one thing about this feature that is not mine to make + /// implicitly. + /// + public static async Task RunAsync( + NpgsqlConnection connection, + string normalizedStatementText, + string createIndexStatement, + int postgresMajorVersion, + CancellationToken cancellationToken = default) + { + ArgumentNullException.ThrowIfNull(connection); + ArgumentException.ThrowIfNullOrWhiteSpace(normalizedStatementText); + ArgumentException.ThrowIfNullOrWhiteSpace(createIndexStatement); + + if (postgresMajorVersion < MinimumPostgresMajorForGenericPlan) + { + return new Result( + false, 0, 0, null, null, null, + $"This server is PostgreSQL {postgresMajorVersion}, and EXPLAIN (GENERIC_PLAN) arrived in " + + $"{MinimumPostgresMajorForGenericPlan}. Stored statement text is normalized — literals " + + "are $1, $2 — so without GENERIC_PLAN there is no way to plan it that does not involve " + + "inventing parameter values, which would produce a plan for a query nobody ran. Refused " + + "rather than guessed."); + } + + /* ROLLED BACK unconditionally: the hypothetical index is session-local, but the session is pooled + and would carry it into the next caller's work. */ + await using var transaction = await connection.BeginTransactionAsync(cancellationToken); + + try + { + await ExecuteAsync(connection, transaction, + $"SET LOCAL statement_timeout = '{StatementTimeoutSeconds}s'", cancellationToken); + + var before = await ExplainAsync(connection, transaction, normalizedStatementText, cancellationToken); + + var indexName = await ScalarTextAsync(connection, transaction, + "SELECT indexname FROM hypopg_create_index($1)", createIndexStatement, cancellationToken); + + var after = await ExplainAsync(connection, transaction, normalizedStatementText, cancellationToken); + + var costBefore = TotalCost(before); + var costAfter = TotalCost(after); + + /* The index NAME appearing anywhere in the second plan is the decisive test, not the cost + falling. Cost can move for reasons that have nothing to do with the candidate, and a cheaper + plan that does not reference it is not evidence for building it. */ + var used = indexName is not null && after is not null + && after.Contains(indexName, StringComparison.Ordinal); + + var saved = costBefore > 0 ? (costBefore - costAfter) / costBefore * 100 : 0; + + /* Formatted once, invariantly, then concatenated. An interpolated-string handler cannot span + a concatenation, and a message this long has to wrap. */ + var beforeText = costBefore.ToString("N2", CultureInfo.InvariantCulture); + var afterText = costAfter.ToString("N2", CultureInfo.InvariantCulture); + var savedText = saved.ToString("N1", CultureInfo.InvariantCulture); + + return new Result( + used, costBefore, costAfter, indexName, before, after, + used + ? $"The planner WOULD use this index: estimated cost falls from {beforeText} to " + + $"{afterText}, a {savedText}% reduction. That is an ESTIMATE from the planner's own " + + "cost model on this server's current statistics, not a measured runtime — no " + + "statement was executed and no index was built." + : $"The planner would NOT use this index. The plan is unchanged at an estimated cost of " + + $"{beforeText}. This is a real answer rather than an inconclusive run: on this " + + "server's current statistics, building it would cost write throughput and disk and " + + "change nothing about this statement."); + } + finally + { + /* Explicit, and not left to disposal: a rollback that is skipped leaves the candidate visible + to whoever gets this pooled session next, and every plan they read after that is wrong in a + way nothing would report. + + Guarded, because a failed EXPLAIN can leave the transaction already aborted and disposed — + and a finally that throws REPLACES the original exception with a meaningless one about + transaction state. Measured: the first run against a real target hid an 08P01 behind an + ObjectDisposedException from this very line. */ + try + { + await transaction.RollbackAsync(CancellationToken.None); + } + catch (Exception) + { + /* Nothing to add and nothing to save: the transaction is already gone, which is the + outcome this block wanted. Swallowed so the real failure reaches the caller. */ + } + + /* AND hypopg_reset(), because the rollback is NOT enough — the assumption that it was is the + one this code originally shipped with, and the verification run disproved it: after two + experiments and two rollbacks, hypopg_list_indexes still returned 2. + + Hypothetical indexes are SESSION-local, not transaction-local. They are held in the + extension's own memory rather than in the catalog, so a transaction never owned them and + rolling one back was never going to remove them. On a pooled connection that means every + plan the next caller reads is planned against phantom indexes — wrong in a way nothing + anywhere would report, which is the worst shape a defect can have in a monitoring tool. + + Its own try: reset failing must not replace a real failure either, and a connection too + broken to run it is a connection the pool will discard. */ + try + { + await using var reset = new NpgsqlCommand("SELECT hypopg_reset()", connection); + await reset.ExecuteNonQueryAsync(CancellationToken.None); + } + catch (Exception) + { + /* Same reasoning as above. */ + } + } + } + + private static async Task ExecuteAsync( + NpgsqlConnection connection, DbTransaction transaction, string sql, CancellationToken cancellationToken) + { + await using var command = new NpgsqlCommand(sql, connection, (NpgsqlTransaction)transaction); + await command.ExecuteNonQueryAsync(cancellationToken); + } + + private static async Task ScalarTextAsync( + NpgsqlConnection connection, DbTransaction transaction, string sql, string argument, CancellationToken cancellationToken) + { + await using var command = new NpgsqlCommand(sql, connection, (NpgsqlTransaction)transaction); + command.Parameters.AddWithValue(argument); + + var value = await command.ExecuteScalarAsync(cancellationToken); + return value as string; + } + + /// + /// The statement text is BOUND, never interpolated — and getting there took two failed designs worth + /// recording, because both look correct. + /// + /// + /// The obvious one is EXPLAIN (GENERIC_PLAN, FORMAT JSON) {text} on an ordinary command. It + /// fails: normalized text carries $1, PostgreSQL's extended protocol parses that as a required + /// parameter, and the bind step supplies none — 08P01: bind message supplies 0 parameters, but + /// prepared statement "" requires 1. Adding a NULL parameter makes the error go away and produces + /// a WRONG answer: amount > NULL is provably NULL, so the planner returns a degenerate + /// Result node instead of the generic plan. The whole point of GENERIC_PLAN is to plan + /// without values, and binding one defeats it silently. + /// + /// + /// + /// So the statement travels as a VALUE into a transaction-local GUC, a DO block runs the EXPLAIN + /// through EXECUTE where $1 is just text, and the plan comes back through a second GUC. + /// The bind step never sees a placeholder because the SQL never contains one. That the statement text + /// is a bound parameter for its whole journey is not incidental — it is the property that makes this + /// safe, and the first design did not have it. + /// + /// + private const string ExplainThroughGucSql = """ + DO $pm$ + DECLARE line text; acc text := ''; + BEGIN + FOR line IN EXECUTE 'EXPLAIN (GENERIC_PLAN, FORMAT JSON) ' || current_setting('pm.stmt') LOOP + acc := acc || line; + END LOOP; + PERFORM set_config('pm.plan', acc, true); + END + $pm$; + """; + + private static async Task ExplainAsync( + NpgsqlConnection connection, DbTransaction transaction, string statementText, CancellationToken cancellationToken) + { + /* is_local = true on both: the settings die with the transaction that is rolled back below, so + nothing survives into the next caller of this pooled session. */ + await using (var stage = new NpgsqlCommand("SELECT set_config('pm.stmt', $1, true)", connection, (NpgsqlTransaction)transaction)) + { + stage.Parameters.AddWithValue(statementText); + await stage.ExecuteNonQueryAsync(cancellationToken); + } + + await using (var explain = new NpgsqlCommand(ExplainThroughGucSql, connection, (NpgsqlTransaction)transaction)) + { + await explain.ExecuteNonQueryAsync(cancellationToken); + } + + await using var read = new NpgsqlCommand("SELECT current_setting('pm.plan', true)", connection, (NpgsqlTransaction)transaction); + return (await read.ExecuteScalarAsync(cancellationToken))?.ToString(); + } + + /// + /// The root node's Total Cost. Zero when the plan cannot be parsed, which the caller reads + /// alongside PlannerWouldUseIt — a zero cost with a false verdict is the shape of a plan that + /// could not be read, and neither number is quoted on its own. + /// + public static double TotalCost(string? planJson) + { + if (string.IsNullOrWhiteSpace(planJson)) + { + return 0; + } + + try + { + return JsonNode.Parse(planJson) is JsonArray { Count: > 0 } array + && array[0] is JsonObject root + && root["Plan"] is JsonObject plan + && plan["Total Cost"] is JsonValue cost + ? cost.GetValue() + : 0; + } + catch (JsonException) + { + return 0; + } + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/RdsEndpoint.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsEndpoint.cs new file mode 100644 index 000000000..070b65e8f --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsEndpoint.cs @@ -0,0 +1,104 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Text.RegularExpressions; + +namespace PerformanceMonitor.Darling.Service.Targets; + +/// +/// What an RDS or Aurora endpoint hostname tells us, for the log-download route (#2538). +/// +/// Why this has to be parsed at all. A target is a connection string. The RDS API needs a +/// DBInstanceIdentifier and a region, and neither appears in a connection string — but both are +/// encoded in the endpoint hostname AWS hands out. Asking an operator to configure them separately would +/// mean two sources of truth for the same server, and the one they typed would eventually disagree with the +/// one they connect to. +/// +/// The cluster/instance distinction is the part that matters. +/// DownloadDBLogFilePortion takes an INSTANCE identifier. An Aurora CLUSTER endpoint +/// (…​.cluster-xxxx.…) names a cluster, whose writer is whatever instance currently holds that role — +/// so a cluster endpoint cannot be used directly and has to be resolved through the API first. Reading a +/// cluster endpoint as an instance id yields DBInstanceNotFound against a perfectly healthy cluster, +/// which is a confusing way to fail. +/// +/// Reader endpoints are called out separately (.cluster-ro-). They round-robin across +/// replicas, so the instance behind one is not stable between calls — and plans captured from a reader are +/// a different workload from the writer's, not a substitute for it. +/// +public static class RdsEndpoint +{ + /// The instance or cluster id, depending on . + /// From the hostname, so it always matches the endpoint actually connected to. + public readonly record struct Parsed(string Identifier, string Region, RdsEndpointKind Kind); + + /* Instance: name.hash.region.rds.amazonaws.com + Cluster: name.cluster-hash.region.rds.amazonaws.com + Reader: name.cluster-ro-hash.region.rds.amazonaws.com + Custom: name.cluster-custom-hash.region.rds.amazonaws.com + + The suffix is matched loosely on purpose: commercial AWS, China (.com.cn) and GovCloud all differ + there, and pinning the exact tail would silently refuse to parse a perfectly valid GovCloud endpoint. + What must be exact is the SECOND label, because that is what separates a cluster from an instance. */ + private static readonly Regex s_endpoint = new( + @"^(?[a-z0-9][a-z0-9-]*)\.(?cluster-ro-[a-z0-9]+|cluster-custom-[a-z0-9]+|cluster-[a-z0-9]+|[a-z0-9]+)\.(?[a-z]{2}(?:-[a-z]+)+-\d)\.rds\.", + RegexOptions.Compiled | RegexOptions.IgnoreCase); + + /// + /// Null when the host is not an RDS endpoint at all — a self-hosted server, a proxy, or an IP. That is + /// an ordinary answer rather than an error: the caller falls back to the pg_read_file route. + /// + public static Parsed? TryParse(string? host) + { + if (string.IsNullOrWhiteSpace(host)) + { + return null; + } + + var match = s_endpoint.Match(host.Trim()); + + if (!match.Success) + { + return null; + } + + var second = match.Groups["second"].Value; + + var kind = second.StartsWith("cluster-ro-", StringComparison.OrdinalIgnoreCase) + ? RdsEndpointKind.ClusterReader + : second.StartsWith("cluster-custom-", StringComparison.OrdinalIgnoreCase) + ? RdsEndpointKind.ClusterCustom + : second.StartsWith("cluster-", StringComparison.OrdinalIgnoreCase) + ? RdsEndpointKind.ClusterWriter + : RdsEndpointKind.Instance; + + return new Parsed( + Identifier: match.Groups["name"].Value, + Region: match.Groups["region"].Value.ToLowerInvariant(), + Kind: kind); + } +} + +/// +/// Which shape of endpoint was found. The distinction decides whether the identifier can be used with +/// DownloadDBLogFilePortion directly or has to be resolved to a writer instance first. +/// +public enum RdsEndpointKind +{ + /// A DB instance. Its identifier IS the log-download identifier. + Instance, + + /// An Aurora cluster writer endpoint. Names a cluster; the writer must be resolved. + ClusterWriter, + + /// An Aurora reader endpoint. Round-robins across replicas, so no stable instance behind it. + ClusterReader, + + /// A custom endpoint over a chosen subset of instances. Same resolution problem as a reader. + ClusterCustom, +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogSource.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogSource.cs new file mode 100644 index 000000000..823939525 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogSource.cs @@ -0,0 +1,172 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Amazon; +using Amazon.RDS; +using Amazon.RDS.Model; + +namespace PerformanceMonitor.Darling.Service.Targets; + +/// +/// Fetches PostgreSQL server-log text from the RDS API, for targets with no filesystem to read (#2538). +/// +/// Why this exists at all. auto_explain writes plans to the server log and nowhere else. +/// On a self-hosted server the collector reads that log with pg_read_file; on Aurora and RDS there is +/// no filesystem, pg_read_server_files is not grantable, and the log is only reachable through +/// DownloadDBLogFilePortion. Same text, different transport — which is exactly why the parsing and +/// redaction were moved into PgPlanLogParser first. +/// +/// This is the only code in the product that reaches a monitored target other than through a +/// database connection. It holds no credentials of its own: the SDK's default chain finds the EC2 +/// instance profile the service already runs under, which is how the monitoring hosts reach every other AWS +/// API today. Nothing is stored, so nothing can leak from config — and a host with no role simply fails the +/// call and the collector degrades, rather than the product asking anyone to paste keys into a file. +/// +/// The marker is deliberately in memory rather than in the store. RDS returns a position to +/// resume from, and keeping it per-process means a restart re-reads a bounded tail instead of nothing. +/// Re-reading is HARMLESS here and that is not luck: plan rows dedup on (queryid, plan_hash), so an +/// overlapping window produces the same shapes rather than duplicates — the same property the +/// pg_read_file route already relies on. Persisting the marker would buy nothing and add a schema +/// rung that could disagree with reality after a log rotation. +/// +public sealed class RdsLogSource +{ + /// + /// How much of a log file to take on the FIRST read of a target, before any marker exists. Bounded for + /// the same reason the file route reads a tail: #2565 measured 772 MB of log in twenty seconds at + /// capture-everything, and an unbounded first read would pull all of it across the network. + /// + private const int FirstReadLines = 10_000; + + private readonly Dictionary _markers = new(StringComparer.Ordinal); + + private readonly Func _clientFactory; + + public RdsLogSource(Func? clientFactory = null) + => _clientFactory = clientFactory + ?? (region => new AmazonRDSClient(RegionEndpoint.GetBySystemName(region))); + + /// Raw log text, to be handed to PgPlanLogParser.Extract unchanged. + /// RDS had more than one call's worth. The caller decides whether to keep + /// pulling; this type does not loop, so one cycle cannot spend unbounded time on one target. + public readonly record struct LogChunk(string Text, bool MoreAvailable); + + /// + /// The newest PostgreSQL log file's unread portion, or null when this target is not RDS at all. + /// + /// A cluster endpoint is resolved to its WRITER, because DownloadDBLogFilePortion takes an + /// instance identifier and because the writer is where the workload worth capturing runs. A READER + /// endpoint is refused rather than guessed at: it round-robins across replicas, so the instance behind + /// it is not stable between calls and plans captured through one would be attributed to whichever + /// replica answered. + /// + public async Task ReadNewestAsync(string host, CancellationToken cancellationToken = default) + { + var endpoint = RdsEndpoint.TryParse(host); + + if (endpoint is null) + { + return null; + } + + var parsed = endpoint.Value; + + if (parsed.Kind is RdsEndpointKind.ClusterReader or RdsEndpointKind.ClusterCustom) + { + throw new InvalidOperationException( + $"'{host}' is an Aurora {(parsed.Kind == RdsEndpointKind.ClusterReader ? "reader" : "custom")} " + + "endpoint, which does not resolve to a stable instance — it moves between replicas " + + "call to call. Point the target at the cluster writer endpoint or at a specific instance " + + "so captured plans belong to a server that can be named."); + } + + using var client = _clientFactory(parsed.Region); + + var instanceId = parsed.Kind == RdsEndpointKind.ClusterWriter + ? await ResolveWriterAsync(client, parsed.Identifier, cancellationToken) + : parsed.Identifier; + + var newest = await NewestLogFileAsync(client, instanceId, cancellationToken); + + if (newest is null) + { + return new LogChunk(string.Empty, false); + } + + var key = instanceId + "|" + newest; + _markers.TryGetValue(key, out var marker); + + var response = await client.DownloadDBLogFilePortionAsync( + new DownloadDBLogFilePortionRequest + { + DBInstanceIdentifier = instanceId, + LogFileName = newest, + /* "0" means from the start, which on a rotated multi-GB log is not what anyone wants on a + first read. NumberOfLines with no marker asks RDS for the TAIL, matching the file + route's bounded-tail behaviour. */ + Marker = marker, + NumberOfLines = marker is null ? FirstReadLines : 0, + }, + cancellationToken); + + if (!string.IsNullOrEmpty(response.Marker)) + { + /* Keyed by FILE as well as instance, so a log rotation starts a fresh marker instead of + resuming a new file at an old file's offset. */ + _markers[key] = response.Marker; + } + + /* AdditionalDataPending is bool? in the SDK. Treated as false when null: claiming more is + pending when the API did not say so would make a caller loop for data that is not there. */ + return new LogChunk(response.LogFileData ?? string.Empty, response.AdditionalDataPending == true); + } + + private static async Task ResolveWriterAsync( + IAmazonRDS client, string clusterId, CancellationToken cancellationToken) + { + var clusters = await client.DescribeDBClustersAsync( + new DescribeDBClustersRequest { DBClusterIdentifier = clusterId }, cancellationToken); + + var cluster = clusters.DBClusters.FirstOrDefault() + ?? throw new InvalidOperationException($"Aurora cluster '{clusterId}' was not found."); + + var writer = cluster.DBClusterMembers.FirstOrDefault(m => m.IsClusterWriter == true) + ?? throw new InvalidOperationException( + $"Aurora cluster '{clusterId}' reports no writer. That is a real state during a failover, " + + "so this is worth retrying rather than treating as a configuration error."); + + return writer.DBInstanceIdentifier; + } + + /// + /// The newest PostgreSQL log file. Filtered by name because an instance's log list also carries + /// upgrade and other logs, and sorted by last-written rather than by name — the filename embeds a + /// timestamp, but sorting text would order 2026-08-9 after 2026-08-10. + /// + private static async Task NewestLogFileAsync( + IAmazonRDS client, string instanceId, CancellationToken cancellationToken) + { + var files = await client.DescribeDBLogFilesAsync( + new DescribeDBLogFilesRequest + { + DBInstanceIdentifier = instanceId, + FilenameContains = "postgresql", + }, + cancellationToken); + + return files.DescribeDBLogFiles + .OrderByDescending(f => f.LastWritten) + .Select(f => f.LogFileName) + .FirstOrDefault(); + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogUnavailableException.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogUnavailableException.cs new file mode 100644 index 000000000..4ac9bb9c3 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsLogUnavailableException.cs @@ -0,0 +1,80 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; + +namespace PerformanceMonitor.Darling.Service.Targets; + +/// +/// The RDS log could not be READ — as distinct from being read and holding nothing (#2633). +/// +/// +/// This type exists because those two states were the same value. IngestAsync caught every failure, +/// logged a warning and returned 0 rows, and the runner turned 0 into a +/// collection_log row reading SUCCESS with the note "no new auto_explain plans in the RDS log +/// window" — a positive claim that the log was opened and was empty. On the monitoring host the truth was +/// rds:DescribeDBLogFiles denied by IAM: nothing was opened at all. +/// +/// +/// +/// That is a REGRESSION against the route this one replaced. The pg_read_file path answers the same +/// situation with PERMISSIONS and a message naming the grant; the managed path, which is the one a +/// real fleet is on, was the one that went quiet. And the app-log warning is not a substitute: +/// collection_log is where collection health is read, and it was saying this collector was fine. +/// +/// +/// +/// Tolerating the failure stays right — one target's IAM gap must not take the cycle down. What changes is +/// that the cycle now says which kind of nothing it found. +/// +/// +public sealed class RdsLogUnavailableException : Exception +{ + public RdsLogUnavailableException(string message, bool isAuthorizationFailure, Exception innerException) + : base(message, innerException) + => IsAuthorizationFailure = isAuthorizationFailure; + + /// + /// The call was DENIED rather than failing for some other reason. + /// + /// Only this case degrades to PERMISSIONS. Everything else — a throttle, a failover, an + /// endpoint that stopped resolving — stays loud, because the store's own rule is that an unclassified + /// failure must be loud rather than quietly swallowed, and a permanent-sounding status on a transient + /// fault is how a real outage gets read as a configuration choice. + /// + public bool IsAuthorizationFailure { get; } + + /// + /// Whether an exception from the AWS SDK is an authorization refusal. + /// + /// Matched on the MESSAGE as well as the type, because the SDK reports this in more than one + /// shape depending on the call: an AmazonServiceException carrying an + /// AccessDenied/AccessDeniedException error code, and — as measured on the fleet — a + /// plain "User: … is not authorized to perform: rds:DescribeDBLogFiles … because no identity-based + /// policy allows" message. Deliberately not matched on the HTTP status alone: 403 also covers an + /// expired credential, which is a different thing to tell an operator. + /// + public static bool IsAuthorizationRefusal(Exception ex) + { + ArgumentNullException.ThrowIfNull(ex); + + for (var current = ex; current is not null; current = current.InnerException) + { + var message = current.Message ?? string.Empty; + + if (message.Contains("is not authorized to perform", StringComparison.OrdinalIgnoreCase) + || message.Contains("no identity-based policy allows", StringComparison.OrdinalIgnoreCase) + || message.Contains("AccessDenied", StringComparison.OrdinalIgnoreCase)) + { + return true; + } + } + + return false; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/RdsPlanIngestor.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsPlanIngestor.cs new file mode 100644 index 000000000..f5db3073c --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/RdsPlanIngestor.cs @@ -0,0 +1,186 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Threading; +using System.Threading.Tasks; +using Microsoft.Extensions.Logging; +using Npgsql; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Service.Targets; + +/// +/// Stores auto_explain plans fetched from the RDS log API into collect.pg_plan_capture +/// (#2538) — the managed-PostgreSQL half of plan capture. +/// +/// Why this is not a collector. Every collector in this product is +/// BuildQuery → DbDataReader → rows. There is no reader here: the text arrives from an AWS +/// API over HTTPS, with no database connection to the target involved at all. Forcing that through a +/// SQL-shaped framework would mean a fake reader wrapping an HTTP response, which is more machinery and +/// less honesty than a small component that says what it is. +/// +/// What it deliberately does NOT duplicate. Parsing, redaction and hashing come from +/// PgPlanLogParser, shared with the pg_read_file route — the redaction living in one place is +/// the whole reason that type exists. The WRITE goes through PgCollectorRowWriter and +/// PgPlanCaptureCollector's own definition, so the column order, the COPY command and the standard +/// prefix are the collector's rather than a second opinion about them. This adds a source, not a schema. +/// +/// Failure is per target and non-fatal. A missing IAM permission, a cluster mid-failover, or a +/// reader endpoint someone pointed at by mistake are all ordinary states, and none of them should stop the +/// other targets — or the rest of the cycle — from collecting. +/// +public sealed class RdsPlanIngestor +{ + private readonly NpgsqlDataSource _postgres; + private readonly RdsLogSource _logs; + private readonly ILogger? _logger; + + public RdsPlanIngestor(NpgsqlDataSource postgres, RdsLogSource? logs = null, ILogger? logger = null) + { + _postgres = postgres ?? throw new ArgumentNullException(nameof(postgres)); + _logs = logs ?? new RdsLogSource(); + _logger = logger; + } + + /// The target's connection host. A non-RDS host is skipped silently — that target + /// uses the pg_read_file route instead, and there is nothing to report. + /// Rows stored, or zero when this target is not RDS or had nothing new. + public async Task IngestAsync( + int serverId, + string storageName, + string host, + CancellationToken cancellationToken = default) + { + RdsLogSource.LogChunk? chunk; + + try + { + chunk = await _logs.ReadNewestAsync(host, cancellationToken); + } + catch (Exception ex) when (ex is not OperationCanceledException) + { + /* #2633: RETHROWN, not returned as zero rows. The warning below stays — it names the target and + carries the AWS message — but the app log is not where collection health is read. Returning 0 + made the runner write SUCCESS with "no new auto_explain plans in the RDS log window", which + is a claim that the log was opened. On the monitoring host the truth was + rds:DescribeDBLogFiles denied: nothing was opened at all, and the row said the collector was + fine. + + Still tolerated at the cycle level — the runner degrades an authorization refusal to + PERMISSIONS and moves on, exactly as the pg_read_file route already does for its own 42501. + What changes is that the cycle now says WHICH kind of nothing it found. */ + _logger?.LogWarning( + "RDS plan log unavailable for {Server}: {Message} — plan capture is skipped for this target " + + "this cycle; every other collector is unaffected.", + storageName, ex.Message); + + throw new RdsLogUnavailableException( + ex.Message, RdsLogUnavailableException.IsAuthorizationRefusal(ex), ex); + } + + if (chunk is null || string.IsNullOrEmpty(chunk.Value.Text)) + { + return 0; + } + + var plans = PgPlanLogParser.Extract(chunk.Value.Text); + + if (plans.Count == 0) + { + /* A log slab with no plans in it is the ordinary case on a server whose threshold nothing + crossed. Not worth a log line every cycle. */ + return 0; + } + + var rows = new List(plans.Count); + + foreach (var plan in plans) + { + rows.Add(new PgPlanCaptureCollector.Row( + QueryId: plan.QueryId, + PlanHash: plan.PlanHash, + DurationMs: plan.DurationMs, + NodeCount: plan.NodeCount, + TopNodeType: plan.TopNodeType, + PlanJson: plan.PlanJson)); + } + + return await WriteAsync(serverId, storageName, rows, cancellationToken); + } + + /// + /// The same binary COPY the collector runner uses, driven by the collector's own definition — so the + /// column order and COPY command come from one place and cannot drift from the table. + /// + private async Task WriteAsync( + int serverId, + string storageName, + IReadOnlyList rows, + CancellationToken cancellationToken) + { + var definition = PgPlanCaptureCollector.Instance; + + /* Naive UTC, the store's convention for every collector timestamp: the columns are `timestamp` + without a zone, and letting Kind=Utc through makes Npgsql infer timestamptz and shift the value + by the store session's offset. */ + var collectionTime = DateTime.SpecifyKind(DateTime.UtcNow, DateTimeKind.Unspecified); + + await using var connection = await _postgres.OpenConnectionAsync(cancellationToken); + + var writer = new PgCollectorRowWriter(); + var written = 0; + + using (var importer = await connection.BeginBinaryImportAsync( + PgCollectorRowWriter.CopyCommandFor(definition), cancellationToken)) + { + writer.Importer = importer; + + foreach (var row in rows) + { + await importer.StartRowAsync(cancellationToken); + + if (definition.IncludesCollectionId) + { + writer.Value(CollectionIdGenerator.Next()); + } + + writer.Value(collectionTime) + .Value(serverId) + .Value(storageName); + + writer.BeginPayload(); + definition.WritePayload(row, writer, NullContext(serverId, storageName, collectionTime)); + writer.EndPayload(definition.PayloadColumns.Count); + written++; + } + + await importer.CompleteAsync(cancellationToken); + } + + _logger?.LogInformation( + "Stored {Count} auto_explain plan(s) for {Server} from the RDS log API.", written, storageName); + + return written; + } + + /* WritePayload takes a context for the collectors that consult deltas or watermarks. This one reads + none of it - the rows are already fully formed by the parser - so the context exists to satisfy the + signature rather than to carry anything. */ + private static CollectorContext NullContext(int serverId, string storageName, DateTime collectionTime) + => new() + { + ServerId = serverId, + ServerName = storageName, + CollectionTime = collectionTime, + Deltas = new CollectorDeltaCalculator(), + Target = new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql }, + }; +} diff --git a/Darling/PerformanceMonitor.Darling.Service/Targets/SqlServerTargetProvider.cs b/Darling/PerformanceMonitor.Darling.Service/Targets/SqlServerTargetProvider.cs index 9927685c4..dc5f24b51 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Targets/SqlServerTargetProvider.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Targets/SqlServerTargetProvider.cs @@ -99,7 +99,10 @@ caller to drop and re-probe the connection rather than just failing one collecto return yieldsOnLockTimeout ? CollectorTargetFault.LockTimeoutYield : CollectorTargetFault.Unclassified; } - if (sql.Number is 229 or 297 or 300 or 8189 or 916) + /* #2512: the numbers live in SqlServerPermissionErrors now, so this and the two runner catch + filters cannot disagree about what a denial is — they already had (916 was here and in + neither of them). */ + if (SqlServerPermissionErrors.IsPermissionDenied(sql.Number)) { return CollectorTargetFault.Permissions; } diff --git a/Darling/PerformanceMonitor.Darling.Service/darling.sample.json b/Darling/PerformanceMonitor.Darling.Service/darling.sample.json index 88b533d38..aa2d21c47 100644 --- a/Darling/PerformanceMonitor.Darling.Service/darling.sample.json +++ b/Darling/PerformanceMonitor.Darling.Service/darling.sample.json @@ -68,7 +68,8 @@ // secure default (loopback-only, byte-for-byte today's behavior: no pg_hba network rule, no // firewall rule, ssl=off). When present with a non-loopback "listen", the service reconciles it // on every start: it adds the bind IP to listen_addresses, generates a self-signed TLS cert for - // verify-full, writes a marked `hostssl darling scram-sha-256` pg_hba rule + + // verify-full, writes a marked `hostssl darling scram-sha-256` pg_hba rule + // per admitted role + // reloads, and best-effort adds a firewall rule (also add the exact scoped rule yourself; see // Darling/README.md). Reconciliation is symmetric (delete the block to close the box) and // fail-closed: any invalid/incomplete field degrades the store to loopback + a critical log line, @@ -79,11 +80,29 @@ // "listen": "192.168.1.205", // the specific bind IP (preferred; = the cert's iPAddress SAN). // // "0.0.0.0" = all interfaces (then connect by a cert SAN name). // "allowFrom": "192.168.1.0/24", // pg_hba + firewall CIDR (its address family must match "listen"). - // "role": "viewer" // "viewer" (default, read-only remote — the secure default) or + // "role": "viewer" // "viewer" (default, read-only remote — the secure default), // // "admin" (full remote WRITES; the service warns, because admin - // // holds the config-table service-credential pivot). Never the superuser. + // // holds the config-table service-credential pivot), or BOTH as + // // "admin,viewer" — one hostssl line per role, so an admin seat + // // and read-only seats reach the same store, and narrowing + // // allowFrom later tightens every one of them. Never the + // // superuser, never all; one bad element rejects the whole value. // } }, + // SEEDED ONCE, then read-only. This list bootstraps config.config_monitored_servers on the first + // successful start and is never consulted for it again: after that the STORE is authoritative for + // which servers are monitored AND for every per-server setting below. So on a seeded box, editing + // "host" / "database" / "auth" / "username" / "encryptMode" / "trustServerCertificate" / + // "excludedDatabases" / "port" / "engine" / "name" here changes nothing, and neither does adding a + // server — a restart cannot apply either. Add servers with the Viewer's Add Server dialog or the MCP + // add_servers tool; CHANGE an already-registered server from the Viewer's Manage Servers window + // (add_servers skips one that is already monitored as a duplicate). The service names any + // disagreement between this list and the registry once per start, in the log. + // Two exceptions that ARE still read from this file on every start, because they are never written to + // the store: the SQL-auth secret ("encryptedPassword" / "password", including an "env:NAME" or + // "file:/path" reference), and this whole file when the store is unreachable. + // Note --test-connection reads THIS FILE, so on a seeded box it probes the file's settings rather than + // the ones the service will actually connect with. "servers": [ { // Integrated auth (recommended): grant the service account VIEW SERVER STATE etc. @@ -242,10 +261,19 @@ // (loopback-only HTTP, today's behavior). When present with a non-loopback "listen", the web host binds the // LAN interface behind an in-app CIDR check and a token gate: a browser presents the token ONCE via // ?token=, which is exchanged for an HMAC-signed session cookie (then the token is stripped from the URL). - // LOOPBACK stays TOKENLESS even while exposed (the surface is read-only). Fail-closed: any missing - // precondition (token / valid allowFrom / managed mode) keeps it loopback-only. There is deliberately NO - // TLS (a browser over plain HTTP on a trusted LAN); the MITM control is a TLS REVERSE PROXY in front of the - // port, same as MCP -- so this is a trusted-LAN opt-in, NEVER internet-exposed. + // LOOPBACK is exempt from the CIDR check ONLY -- while exposed it presents the token like anyone else + // (#1649: the Custom Views composer is a write path, so a tokenless local process could mutate views). + // A dashboard with NO "network" block registers no auth middleware at all and stays tokenless. + // Fail-closed: any missing precondition (token / valid allowFrom / managed mode) keeps it loopback-only. + // TLS is OPT-IN via the "tls" block (#2562) and applies to the network listener only -- loopback stays + // plain HTTP, because the certificate names the LAN address rather than "localhost". WITHOUT it the token + // and its session cookie cross the segment in the clear and every exposed start warns about exactly that; + // "allowFrom" bounds who can ROUTE to the port, never what an on-path attacker reads off the wire. The + // product consumes a certificate, it does not issue one: no ACME, and no self-signed fallback (that would + // buy encryption without authentication and train you to click through the warning). An internal CA is the + // normal answer here. A missing, unreadable, ambiguous or EXPIRED certificate keeps the dashboard + // loopback-only rather than quietly serving the LAN over HTTP. A TLS reverse proxy in front of the port + // remains a perfectly good alternative -- it is what MCP still relies on. NEVER internet-exposed either way. // "enabled"/"port" here are the SEED; after first start they live in the control plane and the Viewer's // Settings toggles them LIVE (the service starts/stops/rebinds the endpoint within seconds - no restart). // The "network" exposure block below stays file-defined and restart-only; the Viewer toggle still stops an @@ -258,7 +286,21 @@ // "network": { // "listen": "192.168.1.205", // specific bind IP preferred; "0.0.0.0" = all interfaces. // "allowFrom": "192.168.1.0/24", // in-app RemoteIpAddress check + firewall CIDR (loopback always allowed). - // "encryptedToken": "" // REQUIRED to expose; or a plaintext "token" (dev only, warned). + // "encryptedToken": "", // REQUIRED to expose; or a plaintext "token" (dev only, warned). + // + // // OPTIONAL TLS (#2562) -- omit the whole block for plain HTTP. Give ONE form, never both. + // // The certificate must name the address browsers actually use ("listen" above), or every visit + // // is a warning; issue it from your internal CA. Renew it: an EXPIRED certificate takes the + // // dashboard down (by design -- it will not fall back to HTTP), and the service logs a warning + // // for the last 30 days before that happens. + // "tls": { + // "pfxPath": "C:\\ProgramData\\PerformanceMonitorDarling\\certs\\dashboard.pfx", + // "encryptedPfxPassword": "" // or "pfxPassword" (file:/env: reference, or plaintext for dev) + // // omit both if the bundle has no password + // // ... OR a PEM pair instead of the two lines above: + // // "certPath": "C:\\ProgramData\\PerformanceMonitorDarling\\certs\\dashboard.crt", + // // "keyPath": "C:\\ProgramData\\PerformanceMonitorDarling\\certs\\dashboard.key" + // } // } }, @@ -291,12 +333,12 @@ // "thisStoreCovers": "the 42 us-east-1 SQL Server primaries", // "stores": [ // { - // "name": "prod-pos-use2-monitor-01", + // "name": "prod-sql-use2-monitor-01", // "covers": "the readable replicas of those same 42 primaries, in-region from us-east-2", // "matches": ["use2"] // }, // { - // "name": "prod-pos-pg-monitor-01", + // "name": "prod-sql-pg-monitor-01", // "covers": "the Aurora PostgreSQL clusters", // "matches": ["-aurora-", ".cluster-"] // } diff --git a/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/app.css b/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/app.css index c9048ceee..9ff24d5a3 100644 --- a/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/app.css +++ b/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/app.css @@ -645,6 +645,11 @@ table.data td.sev-Healthy { color: var(--ok); } .page-head .server-title { display: flex; align-items: center; gap: 0.5rem; } .page-head .server-title h2 { margin: 0; } .page-head .server-band { display: inline-flex; } +.page-head .server-engine { display: inline-flex; } +/* The engine badge names what the server RUNS, and is the answer to "why does this one have six tabs and that + one twelve" (#2530). Deliberately neutral rather than band-coloured: an engine is not a health state, and a + card whose engine_kind is NULL renders no badge at all rather than one asserting a default. */ +.badge.engine { background: var(--lighter); color: var(--muted); font-weight: 500; } /* ─────────────────────────── sort control ─────────────────────────── */ @@ -671,6 +676,48 @@ table.data td.sev-Healthy { color: var(--ok); } } .search-input:focus { outline: none; border-color: var(--accent); } +/* ─────────────────────────── needs-attention filter (#2437) ─────────────────────────── */ + +/* The header toggle, styled as the group-by-tag toggle beside it — it is the same kind of thing, a view + control over the same cards, so it should not look like a new concept. */ +.attention-control { display: inline-flex; align-items: center; gap: 0.35rem; font-size: 0.78rem; color: var(--muted); cursor: pointer; } +.attention-control input { cursor: pointer; } + +/* The "+N more need attention" line and the notice's way back out. Underlined, accented and pointer-cursored + because a line that navigates has to look like one: the defect being fixed was an inert muted div, which is + the shape of a caption, and every reader treated it as one. The focus ring matches the page's other + activatable divs (util.el's onActivate makes them role=button + tabbable). */ +.attention-link { + display: inline-block; + margin: 0.4rem 0 0.2rem; + color: var(--accent); + font-size: 0.85rem; + text-decoration: underline; + cursor: pointer; +} +.attention-link:hover { filter: brightness(1.15); } +.attention-link:focus-visible { outline: 2px solid var(--accent); outline-offset: 2px; } + +/* The active state, on the grid it shrank. It carries its own colour: the amber arm is a count of servers + wanting attention, the green arm is an all-clear, and painting an all-clear amber would be a colour + contradicting its own text. */ +.attention-note { + display: flex; + align-items: center; + gap: 0.75rem; + flex-wrap: wrap; + margin: 0 0 0.75rem; + padding: 0.5rem 0.75rem; + border-radius: var(--radius); + font-size: 0.85rem; +} +.attention-note.warn { background: rgba(255, 213, 79, 0.10); border: 1px solid rgba(255, 213, 79, 0.35); color: var(--warn); } +.attention-note.ok { background: rgba(129, 199, 132, 0.10); border: 1px solid rgba(129, 199, 132, 0.35); color: var(--ok); } +/* Neither: the search matched nothing, so the filter judged nothing. Green would read as an all-clear the + data does not support, amber would claim a problem it never looked for. */ +.attention-note.none { background: var(--bg-dark); border: 1px solid var(--border); color: var(--muted); } +.attention-note .attention-link { margin: 0; } + /* ─────────────────────────── status bar footer ─────────────────────────── */ #statusbar { @@ -865,3 +912,58 @@ pre.code { color: var(--dim); margin-left: 0.3rem; } + +/* ─────────────────────────── server page: sub-tabs + range ─────────────────────────── */ + +/* The web port of ViewerServerTab.xaml's per-server TabControl. Real links (the tab id rides in the + hash), so a tab is bookmarkable and middle-clickable; the strip scrolls horizontally on a narrow window + rather than wrapping into a block that pushes the panels off-screen. */ +.subtabs { + display: flex; + gap: 0.15rem; + overflow-x: auto; + border-bottom: 1px solid var(--border); + margin-bottom: 1rem; + scrollbar-width: thin; +} +.subtab { + flex: 0 0 auto; + padding: 0.45rem 0.8rem; + font-size: 0.85rem; + color: var(--muted); + text-decoration: none; + border-bottom: 2px solid transparent; + white-space: nowrap; +} +.subtab:hover { color: var(--fg); background: rgba(127, 127, 127, 0.08); } +.subtab.active { color: var(--fg); font-weight: 600; border-bottom-color: var(--accent); } +.subtab:focus-visible { outline: 2px solid var(--muted); outline-offset: -2px; } + +/* The tab's honesty note sits between the tab strip and the panels — it describes what this tab does NOT do, + so it belongs above the panels it qualifies rather than buried under them. */ +.subtabs + .strip.notice { margin-bottom: 1rem; } + +/* The page-level time-range preset picker (the web twin of the viewer's toolbar window presets). */ +.range-control { display: inline-flex; align-items: center; gap: 0.4rem; font-size: 0.78rem; color: var(--muted); } +.range-select-inline { + background: var(--card); + color: var(--fg); + border: 1px solid var(--border); + border-radius: 4px; + padding: 0.2rem 0.4rem; + font-size: 0.78rem; +} + +/* The in-panel picker row (wait-type / perfmon-counter composites) — the single-select web twin of the desktop + viewer's checkbox pickers. Sits between a panel's table and its chart. */ +.picker-row { display: flex; align-items: center; gap: 0.5rem; margin: 0.85rem 0 0.4rem; } + +/* The server header's "why": the pre-banded reason sentence (when the fleet ranked this server) over the same + metric chips the fleet card carries. A band word with no way to ask why is #2422 on a new surface. */ +.server-why { margin: -0.4rem 0 1rem; } +.server-reason { font-size: 0.85rem; margin-bottom: 0.5rem; } +.server-reason.band-Critical { color: var(--err); } +.server-reason.band-Warning { color: var(--warn); } +.server-reason.band-Offline { color: var(--err); } +.server-reason.band-Healthy { color: var(--ok); } +.server-why .metric-bands { max-width: 900px; } diff --git a/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/editor.css b/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/editor.css index d53a1d104..b91e449cc 100644 --- a/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/editor.css +++ b/Darling/PerformanceMonitor.Darling.Service/wwwroot/css/editor.css @@ -530,3 +530,57 @@ min-height: 190px; overflow-x: auto; } + +/* ── starter dashboard templates (#2476) ── */ + +/* The template menu now carries two groups (dashboards, notebooks), so each needs a heading, and the dashboard + items are + + private readonly ViewerAppSettingsStore _appSettingsStore = new(); + /// + /// The per-user settings files that were on disk at startup and could not be read AT ALL, each with why + /// (#2434). Collected in the constructor, where all three stores are loaded, and reported once from + /// — a dialog raised from the constructor has no window behind it. Empty on + /// every ordinary run INCLUDING a first run: an absent file is not a problem, and saying so would make + /// a warning out of the most normal thing the viewer ever does. + /// + private readonly List _unreadableSettingsFiles = new(); + + /// + /// The individual SETTINGS that could not be read, per file, when the rest of that file loaded normally + /// (#2456). Kept apart from because the two need different + /// sentences: one says every setting in the file reverted, the other says these did and nothing else + /// did. Collapsing them would put an overstatement in front of whichever case was not being described, + /// and the whole complaint here was a dialog that could not say which settings were lost. + /// + private readonly List _unreadableSettingsValues = new(); + /// /// The system-tray icon + minimize-to-tray behavior (ported from Lite's SystemTrayService, adapted for /// the headless viewer). Null until initializes it once the store connects. @@ -126,6 +144,24 @@ the constructor and refreshed when the Settings window closes (mirroring ViewerE /// at the fleet bind site so tiles re-sort deterministically every refresh. private ServerOverviewSortMode _overviewSortMode; + /// + /// The FULL, sorted Overview card set — the authority behind OverviewItemsControl.ItemsSource since + /// #2424, which now shows a possibly-filtered projection of it. Everything that used to read the bound list + /// back to answer "what cards are there" reads this instead, for the same reason + /// separates All from Visible: the moment the bound list is a SUBSET, a consumer that reads it + /// back is silently answering a different question. + /// + private List _overviewCards = new(); + + /// + /// True while the Overview grid is filtered to servers that need attention (#2424) — the destination the + /// "+N more need attention" line finally has. Deliberately NOT persisted to ViewerAppSettings the way + /// is: a sort is a preference, this is a triage action tied to a moment, and + /// an Overview that opens with 52 of 57 servers already hidden is a support ticket even with the toggle in + /// plain sight. + /// + private bool _overviewAttentionOnly; + /// /// The non-secret configuration block shown on every connection/config failure (#1954): which /// darling.json won and what was parsed from it. Built at startup before the load is attempted (so a @@ -145,12 +181,22 @@ public MainWindow() { InitializeComponent(); _preferences = _preferencesStore.Load(); + NoteUnreadableSettingsFile(_preferencesStore.FilePath, _preferencesStore.LastLoadState, _preferencesStore.LastLoadProblem, + _preferencesStore.LastLoadUnreadableMembers); + /* The registry loads in its own field initializer, which has already run by the time the + constructor body does, so its state is available here — and it is the one of the three worth + surfacing most: a viewer that opens with no servers because it could not read viewer-servers.json + looks exactly like a viewer nobody has configured yet. */ + NoteUnreadableSettingsFile(_serverStore.FilePath, _serverStore.LastLoadState, _serverStore.LastLoadProblem, + _serverStore.LastLoadUnreadableMembers); /* Seed the persisted app settings the viewer honors at runtime BEFORE any tab/chart renders: the CSV export separator (grid exports) and the Server/Local/UTC time-display mode (every timestamp render routes through ViewerTimeHelper, which reads CurrentDisplayMode). UiTimeContext is deliberately left at its identity default — Darling charts pre-convert their X through ForDisplay, so wiring it would double-convert on hover/crosshair. */ var appSettings = _appSettingsStore.Load(); + NoteUnreadableSettingsFile(_appSettingsStore.FilePath, _appSettingsStore.LastLoadState, _appSettingsStore.LastLoadProblem, + _appSettingsStore.LastLoadUnreadableMembers); ViewerExportSettings.Apply(appSettings); /* Seed the new-mute-rule default expiration preference so every MuteRuleEditDialog opens on it, and the dismiss/mute action-logging opt-in the Alerts tab honors. */ @@ -178,10 +224,99 @@ and the single fleet-refresh cadence both shell timers run at. */ Closed += OnClosed; } + /// + /// Records a settings file that is present and could not be read, for the one startup report. Absent + /// and readable both record nothing, which is the whole point of the three-way split. + /// + /// Since #2456 the state carries a fourth possibility inside Unreadable: the file was read + /// after dropping members the deserializer could not understand, and names + /// them. That goes on the other list, because it is a different claim — those settings reverted and the + /// rest of the file did not. + /// + private void NoteUnreadableSettingsFile( + string filePath, + SettingsFileState state, + string? problem, + IReadOnlyList members) + { + if (state != SettingsFileState.Unreadable) + { + return; + } + + var fileName = System.IO.Path.GetFileName(filePath); + + if (members.Count > 0) + { + _unreadableSettingsValues.Add($"{fileName} — {string.Join(", ", members)}"); + return; + } + + _unreadableSettingsFiles.Add($"{fileName} — {problem}"); + } + + /// + /// One dialog, once, when a settings file the viewer has just fallen back to defaults for is still + /// sitting on disk unreadable (#2434). + /// + /// The fallback used to report itself through Debug.WriteLine, which the compiler removes + /// from a Release build — so in the viewer anyone actually runs, the operator's settings quietly became + /// defaults and there was nowhere whatsoever to find out why. The log carries it now, but a person + /// looking at an app that has forgotten its configuration should not have to go and find a log to learn + /// that it did. + /// + /// Raised from Loaded so it has a window behind it, and before the store connect below on + /// purpose: it costs nothing when there is nothing to say, and the connection-failure overlay would + /// otherwise bury the one message that explains why the settings changed. + /// + /// Two paragraphs rather than one list, because there are two different facts and #2456 is the + /// issue about telling them apart. A file that could not be read at all costs every setting in it; a + /// file that was read after dropping named members costs only those. One sentence covering both would + /// have to overstate one of them, and an overstated capability is worse than an admitted gap — which + /// is also why the second paragraph says the settings are NAMED rather than implying the list is + /// exhaustive of everything wrong with the file. + /// + private void ReportUnreadableSettingsFiles() + { + if (_unreadableSettingsFiles.Count == 0 && _unreadableSettingsValues.Count == 0) + { + return; + } + + var message = new System.Text.StringBuilder(); + + if (_unreadableSettingsFiles.Count > 0) + { + message + .Append("The viewer could not read these settings files, so the settings they hold are at ") + .Append("their defaults for this session:\n\n") + .AppendJoin('\n', _unreadableSettingsFiles) + .Append("\n\n"); + } + + if (_unreadableSettingsValues.Count > 0) + { + message + .Append("These individual settings could not be read and are at their defaults for this ") + .Append("session. Everything else in their file loaded normally:\n\n") + .AppendJoin('\n', _unreadableSettingsValues) + .Append("\n\n"); + } + + message + .Append("The files have been left exactly as they are. Fix the JSON and restart to get those ") + .Append("settings back — or save from the Settings window, which copies an unreadable file ") + .Append("aside as '.unreadable-' before replacing it."); + + MessageBox.Show(message.ToString(), "Settings", MessageBoxButton.OK, MessageBoxImage.Warning); + } + private async void OnLoaded(object sender, RoutedEventArgs e) { Loaded -= OnLoaded; + ReportUnreadableSettingsFiles(); + /* #1954: say WHERE the config came from before trying to read it, so a missing or unparseable darling.json still names the file the viewer looked at and the rule that picked it. Resolving here (rather than inside TryLoad) also means the path we report is provably the path we load — @@ -241,6 +376,9 @@ as the operator wrote it AND as the anchor resolves it (#1970) — the same two var storeVersion = await _dataService.GetStoreSchemaVersionAsync(); if (storeVersion is int version && version < ViewerDataService.RequiredStoreSchemaVersion) { + /* The store answered, so this is not Unreachable — but the seat probe never ran, so the + field stays at "Seat: --" rather than claiming a verdict nothing measured (#2479). */ + ApplySeatState(ViewerSeatState.Unknown); ShowConnectionFailure( $"The Darling store is at schema v{version}, but this viewer needs v{ViewerDataService.RequiredStoreSchemaVersion}. " + "Update or restart the Darling service so it migrates the store, then reopen the viewer."); @@ -252,9 +390,18 @@ as the operator wrote it AND as the anchor resolves it (#1970) — the same two probe fails safe to read-only; the write surfaces gate on ViewerDataService.IsReadOnly and the write paths translate a live 42501 into a friendly message as a backstop. */ await _dataService.DetectReadOnlyAsync(); + + /* #2479 items 3+4: publish the verdict ONCE, globally, instead of letting it be discovered a + dialog at a time. Same probe, same fail-safe — this only makes the answer visible before a + write is attempted, which is the whole of #2400. */ + ApplySeatState(_dataService.IsReadOnly ? ViewerSeatState.ReadOnly : ViewerSeatState.ReadWrite); } catch (ViewerStoreUnreachableException ex) { + /* NOT ReadOnly. #2117 built this arm precisely so an unreachable store stops being misread as + a read-only seat, and a status field that collapsed them would undo it in the one place an + operator now looks first — sending them to fix a role they never needed to touch. */ + ApplySeatState(ViewerSeatState.Unreachable); ViewerLogger.Error("App", "Darling store unreachable", ex); ShowConnectionFailure(ex.Message); return; @@ -263,6 +410,7 @@ the write paths translate a live 42501 into a friendly message as a backstop. */ { /* Reachable but the first connection failed for another reason (e.g. authentication, or the configured database does not exist) — show it rather than dead-ending on a blank window. */ + ApplySeatState(ViewerSeatState.Unreachable); ViewerLogger.Error("App", "Darling store connection failed", ex); ShowConnectionFailure($"Couldn't connect to the Darling store: {ex.Message}"); return; @@ -749,6 +897,44 @@ private void UpdateServerCountText() => ? $"Servers: {_fleet.VisibleServerCount} of {_fleet.TotalCount}" : $"Servers: {_fleet.TotalCount}"; + /// + /// The status bar's "Seat:" field (#2479, items 3 and 4) — text, colour and the tooltip that explains + /// WHY, in one place so no caller can set half of it. + /// + /// The colours are the existing status-bar vocabulary, not a new one: + /// ErrorBrush for a store that cannot be reached (the same severity CollectorHealthText + /// gives an erroring collector), WarningBrush for a read-only seat — visible, because a tester + /// who does not notice it is back to discovering the seat one refusal at a time, but not alarming, + /// because on a remote seat it is the correct and expected default — and the muted foreground for + /// read-write and for not-yet-probed, which are the unremarkable states. + /// + /// only ORDERS the two sentences in the + /// tooltip by which default is likelier to have decided this seat. It cannot claim which one did: a + /// non-loopback store host does not prove a remote seat (#2279), so the tooltip names both. + /// + private void ApplySeatState(ViewerSeatState state) + { + SeatText.Text = ViewerSeatIndicator.Text(state); + SeatText.ToolTip = ViewerSeatIndicator.ToolTip(state, _dataService?.StoreIsOnThisMachine ?? true); + + /* TryFindResource rather than the hard FindResource the sibling status fields use, and the reason is + the call site rather than doubt about the keys: all three brushes exist in all three shipped + themes, but ApplySeatState is called from INSIDE the catch that handles an unreachable store. A + ResourceReferenceKeyNotFoundException thrown from a catch block is unhandled, so a missing key in + some future theme would turn "the service is down" into "the viewer crashed" - the one moment the + viewer most needs to stay up and explain itself. Falling back to the muted foreground costs a + colour and nothing else. */ + SeatText.Foreground = + (state switch + { + ViewerSeatState.Unreachable => TryFindResource("ErrorBrush"), + ViewerSeatState.ReadOnly => TryFindResource("WarningBrush"), + _ => TryFindResource("ForegroundMutedBrush"), + } as System.Windows.Media.Brush) + ?? TryFindResource("ForegroundMutedBrush") as System.Windows.Media.Brush + ?? SeatText.Foreground; + } + private async Task LoadServersAsync(bool preserveSelection = false) { if (_dataService is null) @@ -1031,6 +1217,25 @@ private void OnApplyTimeRangeToAllRequested(ViewerServerTab source, int index, D // ── Time-display mode (Server / Local / UTC) ───────────────────────────────────── + /// + /// Says that a setting the user just changed did not reach disk (#2434). + /// + /// The two handlers below persist on an ordinary click — picking a time-display mode, re-sorting + /// the Overview — so a failure there has no Save button to report through, and until now had nowhere to + /// go at all: the store's Save could throw, neither handler caught it, and the whole-object write had + /// already replaced whatever it could not read. Save answers with a bool instead of throwing now, and + /// this is what does something with the answer. Which file and what was wrong with it is in the log; + /// what the user needs from a dialog is that the click did not stick. + /// + private void WarnSettingNotSaved(string what, string filePath) + { + MessageBox.Show( + $"Your {what} could not be saved to '{System.IO.Path.GetFileName(filePath)}'. It applies for " + + "the rest of this session but will not survive a restart. The viewer log (sidebar: View Log) " + + "says why.", + "Settings", MessageBoxButton.OK, MessageBoxImage.Warning); + } + /// /// A server tab's Server/Local/UTC picker changed (the raising tab already set the global mode via /// and reloaded itself). Persist the new mode and sync @@ -1041,7 +1246,11 @@ private void OnDisplayModeChanged(TimeDisplayMode mode) { var settings = _appSettingsStore.Load(); settings.TimeDisplayMode = mode.ToString(); - _appSettingsStore.Save(settings); + if (!_appSettingsStore.Save(settings)) + { + WarnSettingNotSaved("time display mode", _appSettingsStore.FilePath); + } + SyncDisplayModeToOpenTabs(mode); } @@ -1085,14 +1294,17 @@ private void OverviewSortSelector_SelectionChanged(object sender, SelectionChang var settings = _appSettingsStore.Load(); settings.OverviewSortMode = ServerOverviewSort.ToToken(_overviewSortMode); - _appSettingsStore.Save(settings); - - if (OverviewItemsControl.ItemsSource is IEnumerable current) + if (!_appSettingsStore.Save(settings)) { - OverviewItemsControl.ItemsSource = ServerOverviewSort.Order( - current.ToList(), _overviewSortMode, - s => s.CpuPercentForAlert, s => s.DisplayName, s => s.ServerId); + WarnSettingNotSaved("Overview sort order", _appSettingsStore.FilePath); } + + /* Sorts the FULL card set, never the bound list: with the needs-attention filter on, the bound list is + a subset, and sorting that would quietly discard every card the filter is hiding. */ + _overviewCards = ServerOverviewSort.Order( + _overviewCards, _overviewSortMode, + s => s.CpuPercentForAlert, s => s.DisplayName, s => s.ServerId); + ApplyOverviewCardFilter(); } /// Attaches each server's tags as coloured pills to its Overview summary, joining the loaded tag @@ -1123,17 +1335,19 @@ private void StampTagPills(IReadOnlyList summaries) } } - /// Re-stamps tag pills onto the cards the Overview is currently showing (after a tag colour / - /// assignment change), reassigning ItemsSource so the change renders without a full per-server reload. + /// Re-stamps tag pills onto the Overview's cards (after a tag colour / assignment change) and + /// re-projects so the change renders without a full per-server reload. Stamps the FULL card set rather than + /// the bound list, so a card the needs-attention filter is hiding does not come back missing its pills. /// A no-op unless the Overview is populated. private void RestampOverviewTagPills() { - if (OverviewItemsControl.ItemsSource is IEnumerable items) + if (_overviewCards.Count == 0) { - var list = items.ToList(); - StampTagPills(list); - OverviewItemsControl.ItemsSource = list; + return; } + + StampTagPills(_overviewCards); + ApplyOverviewCardFilter(); } /// @@ -1152,16 +1366,14 @@ private async Task LoadOverviewAsync() var servers = _fleet.All; if (servers.Count == 0) { - OverviewItemsControl.ItemsSource = null; - FleetRollupContainer.Visibility = Visibility.Collapsed; + ClearOverviewCards(); return; } var list = servers.ToList(); if (list.Count == 0) { - OverviewItemsControl.ItemsSource = null; - FleetRollupContainer.Visibility = Visibility.Collapsed; + ClearOverviewCards(); return; } @@ -1190,7 +1402,11 @@ private async Task LoadOverviewAsync() built, _overviewSortMode, s => s.CpuPercentForAlert, s => s.DisplayName, s => s.ServerId); - OverviewItemsControl.ItemsSource = cards; + + /* The full set is the authority; the grid shows whatever the needs-attention filter leaves of it. A + refresh must not silently drop the filter, so this re-applies it rather than binding `cards` directly. */ + _overviewCards = cards; + ApplyOverviewCardFilter(); /* Fleet-wide NOC roll-up above the cards — the cross-server totals SQL aggregate (the read only Darling's central store can serve) plus the band counts / worst-N ranking reduced from the @@ -1224,11 +1440,85 @@ private void ApplyFleetRollup(FleetRollup rollup) FleetRollupContainer.Visibility = rollup.TotalServers > 0 ? Visibility.Visible : Visibility.Collapsed; FleetWorstServersList.Visibility = rollup.HasProblems ? Visibility.Visible : Visibility.Collapsed; - FleetAdditionalProblemsText.Visibility = + FleetAdditionalProblems.Visibility = rollup.AdditionalProblemCount > 0 ? Visibility.Visible : Visibility.Collapsed; FleetAllHealthyText.Visibility = rollup.HasProblems ? Visibility.Collapsed : Visibility.Visible; } + /// Empties the Overview: no cards, no roll-up, and no stale authority left behind for the filter + /// or the tag re-stamp to project from. + private void ClearOverviewCards() + { + _overviewCards = new List(); + OverviewItemsControl.ItemsSource = null; + FleetRollupContainer.Visibility = Visibility.Collapsed; + ApplyOverviewAttentionCount(shown: 0); + } + + /// + /// Projects onto the grid through the needs-attention filter (#2424) — the + /// ONE place ItemsSource is set for a populated Overview, so a refresh, a re-sort and a tag re-stamp cannot + /// each have their own opinion about whether the filter is still on. + /// + /// The predicate is , the same banding the roll-up counted + /// the "+N more" with, so the grid the link lands on holds exactly the servers the link was counting. + /// + private void ApplyOverviewCardFilter() + { + var shown = _overviewAttentionOnly + ? FleetRollup.NeedsAttention(_overviewCards) + : _overviewCards; + + OverviewItemsControl.ItemsSource = shown; + ApplyOverviewAttentionCount(shown.Count); + } + + /// + /// Shows what the filter did, beside the toggle that did it. Only rendered while the filter is on: a grid + /// showing every server needs no arithmetic, and a filtered one must never be mistakable for it. + /// + /// The colour follows the sentence. This line has two of them — a count of servers wanting attention, + /// and an all-clear when the filter matched none — and painting the all-clear amber would be a colour + /// contradicting its own text, which is the family of defect this whole change is about. Set through + /// SetResourceReference (the alert badge's idiom) so it still tracks a theme change. + /// + private void ApplyOverviewAttentionCount(int shown) + { + if (_overviewAttentionOnly) + { + OverviewAttentionCountText.Text = FleetRollup.AttentionFilterCountText(shown, _overviewCards.Count); + OverviewAttentionCountText.SetResourceReference( + ForegroundProperty, shown > 0 ? "WarningBrush" : "SuccessBrush"); + OverviewAttentionCountText.Visibility = Visibility.Visible; + return; + } + + OverviewAttentionCountText.Text = string.Empty; + OverviewAttentionCountText.Visibility = Visibility.Collapsed; + } + + /// The needs-attention toggle: re-projects the grid from the unchanged card set, so turning it off + /// restores the full fleet without a store round-trip. Deliberately carries no IsLoaded guard, unlike the sort + /// selector — that one needs it because seeding the ComboBox raises SelectionChanged during construction, and + /// nothing seeds this box. A guard here would only be able to drop a real transition and leave a ticked box + /// over an unfiltered grid, which is the same lie as the one this fixes, in the other direction. + private void OverviewAttentionOnlyCheck_Changed(object sender, RoutedEventArgs e) + { + _overviewAttentionOnly = OverviewAttentionOnlyCheck.IsChecked == true; + ApplyOverviewCardFilter(); + } + + /// + /// "+N more need attention" — the servers the ranking's cap could not show. Clicking it turns the filter on + /// rather than doing the filtering itself, so the state lands in the toggle where the reader can see it is on + /// and can turn it off; the alternative, a grid that silently shrank, is the worse version of this defect. + /// + private void FleetAdditionalProblems_Click(object sender, MouseButtonEventArgs e) + { + OverviewAttentionOnlyCheck.IsChecked = true; + e.Handled = true; + } + /// Double-clicking an Overview card opens (or focuses) that server's tab (Lite's rule). private void OverviewCard_MouseLeftButtonDown(object sender, MouseButtonEventArgs e) { @@ -1522,7 +1812,10 @@ mean you cannot edit the schedule of a server it happens to be hiding. */ if (settings.ShowDialog() == true && settings.Result is not null) { _preferences = settings.Result; - _preferencesStore.Save(_preferences); + if (!_preferencesStore.Save(_preferences)) + { + WarnSettingNotSaved("default time range and auto-refresh", _preferencesStore.FilePath); + } } /* Re-seed the runtime settings the viewer honors from whatever the window persisted (it self-saves): diff --git a/Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj b/Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj index 149a8fb65..cdb6fbed3 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj +++ b/Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj @@ -15,7 +15,7 @@ hooks go unhandled. It is a harmless no-op for the co-located (non-Velopack) viewer that ships inside the service zip. --> PerformanceMonitor.Darling.Viewer.Program - 3.5.0 + 3.6.0 + diff --git a/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml b/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml index 274a14f28..2c412fe80 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml +++ b/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml @@ -438,6 +438,19 @@ + + + + + + + + + diff --git a/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml.cs b/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml.cs index ae4fe0b4f..a02f802f7 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/SettingsWindow.xaml.cs @@ -748,6 +748,10 @@ private void SeedAlertControlsFrom(AlertSettingsRow r) AlertPvsCheckBox.IsChecked = r.PvsEnabled; AlertPvsThresholdPercentBox.Text = r.PvsThresholdPercent.ToString(CultureInfo.InvariantCulture); AlertPvsFloorGbBox.Text = r.PvsFloorGb.ToString(CultureInfo.InvariantCulture); + AlertFileGrowthCheckBox.IsChecked = r.FileGrowthEnabled; + AlertFileGrowthRiseMbBox.Text = r.FileGrowthRiseMb.ToString(CultureInfo.InvariantCulture); + AlertFileGrowthVolumePercentBox.Text = r.FileGrowthVolumePercent.ToString(CultureInfo.InvariantCulture); + AlertFileGrowthLookbackMinutesBox.Text = r.FileGrowthLookbackMinutes.ToString(CultureInfo.InvariantCulture); AlertLongRunningJobCheckBox.IsChecked = r.LongRunningJobEnabled; AlertLongRunningJobMultiplierBox.Text = r.LongRunningJobMultiplier.ToString(CultureInfo.InvariantCulture); AlertFailedJobCheckBox.IsChecked = r.FailedJobEnabled; @@ -802,6 +806,7 @@ private AlertSettingsRow BuildAlertRowFromControls(List errors) TempDbSpaceEnabled = AlertTempDbSpaceCheckBox.IsChecked == true, LowDiskEnabled = AlertLowDiskCheckBox.IsChecked == true, PvsEnabled = AlertPvsCheckBox.IsChecked == true, + FileGrowthEnabled = AlertFileGrowthCheckBox.IsChecked == true, LongRunningJobEnabled = AlertLongRunningJobCheckBox.IsChecked == true, FailedJobEnabled = AlertFailedJobCheckBox.IsChecked == true, DatabaseStateEnabled = AlertDatabaseStateCheckBox.IsChecked == true, @@ -855,6 +860,13 @@ make the gate impossible to turn back off once enabled. */ row.PvsThresholdPercent = pvsPct; if (int.TryParse(AlertPvsFloorGbBox.Text, out var pvsFloor) && pvsFloor >= 0) row.PvsFloorGb = pvsFloor; + /* #2391: validated to the same ranges DarlingAlertSettings clamps. */ + if (int.TryParse(AlertFileGrowthRiseMbBox.Text, out var growthRise) && growthRise >= 0) + row.FileGrowthRiseMb = growthRise; + if (int.TryParse(AlertFileGrowthVolumePercentBox.Text, out var growthPct) && growthPct is >= 0 and <= 100) + row.FileGrowthVolumePercent = growthPct; + if (int.TryParse(AlertFileGrowthLookbackMinutesBox.Text, out var growthLookback) && growthLookback is >= 5 and <= 1440) + row.FileGrowthLookbackMinutes = growthLookback; if (int.TryParse(AlertLongRunningJobMultiplierBox.Text, out var jobMult) && jobMult is >= 2 and <= 20) row.LongRunningJobMultiplier = jobMult; if (int.TryParse(AlertFailedJobLookbackBox.Text, out var failedJobLookback) && failedJobLookback is >= 1 and <= 1440) @@ -932,6 +944,9 @@ private void RestoreAlertDefaultsButton_Click(object sender, RoutedEventArgs e) AnalysisNotifyCooldownBox.Text = "360"; AlertPvsThresholdPercentBox.Text = "40"; AlertPvsFloorGbBox.Text = "1"; + AlertFileGrowthRiseMbBox.Text = "10240"; + AlertFileGrowthVolumePercentBox.Text = "60"; + AlertFileGrowthLookbackMinutesBox.Text = "60"; AlertLongRunningJobMultiplierBox.Text = "3"; AlertFailedJobLookbackBox.Text = "60"; AlertCooldownBox.Text = "5"; @@ -984,6 +999,8 @@ private void UpdateAlertPreviewText() parts.Add($"disk free < {AlertLowDiskThresholdPercentBox.Text}% or {AlertLowDiskThresholdGbBox.Text}GB"); if (AlertPvsCheckBox.IsChecked == true) parts.Add($"PVS >= {AlertPvsThresholdPercentBox.Text}% of database"); + if (AlertFileGrowthCheckBox.IsChecked == true) + parts.Add($"file growth > {AlertFileGrowthRiseMbBox.Text}MB/{AlertFileGrowthLookbackMinutesBox.Text}m or volume > {AlertFileGrowthVolumePercentBox.Text}%"); if (AlertLongRunningJobCheckBox.IsChecked == true) parts.Add($"jobs > {AlertLongRunningJobMultiplierBox.Text}x avg"); if (AlertFailedJobCheckBox.IsChecked == true) @@ -1038,6 +1055,10 @@ private void UpdateAlertControlStates() AlertCollectionStaleMinutesBox.IsEnabled = enabled; AlertCollectionFailureThresholdBox.IsEnabled = enabled; AlertStoreJobCadenceWarnPercentBox.IsEnabled = enabled; + AlertFileGrowthCheckBox.IsEnabled = enabled; + AlertFileGrowthRiseMbBox.IsEnabled = enabled; + AlertFileGrowthVolumePercentBox.IsEnabled = enabled; + AlertFileGrowthLookbackMinutesBox.IsEnabled = enabled; AlertLongRunningJobCheckBox.IsEnabled = enabled; AlertLongRunningJobMultiplierBox.IsEnabled = enabled; AlertFailedJobCheckBox.IsEnabled = enabled; @@ -1485,7 +1506,22 @@ to fix the URL. NOT blocked — a plaintext POST to a trusted LAN listener is le SaveCsvSeparator(); SaveTimeDisplayMode(); SaveColorTheme(); - _appSettingsStore.Save(_appSettings); + + /* #2434: a whole-object replace that did not happen must not pass for one that did. Said here, + at the point it happens, rather than folded into the validation list below — that list's + sentence is about values the window rejected, which is a different thing from a file it could + not write. The operator config further down goes to the Darling store and reports separately, + so a failure here does not stop it. */ + if (!_appSettingsStore.Save(_appSettings)) + { + MessageBox.Show( + "The viewer's own settings could not be written to " + + $"'{System.IO.Path.GetFileName(_appSettingsStore.FilePath)}', so the viewer-local " + + "preferences on this page (theme, CSV separator, timestamp display, tray options) will be " + + "back to their previous values on the next launch. The viewer log says why.", + "Settings", MessageBoxButton.OK, MessageBoxImage.Warning); + } + Result = BuildViewerPreferences(); if (errors.Count > 0) diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerActualPlanFlow.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerActualPlanFlow.cs index c4c2fef00..eb45dacab 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerActualPlanFlow.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerActualPlanFlow.cs @@ -179,7 +179,8 @@ private static void ShowOutcome(Window owner, ActualPlanResult result) private static void ShowReadOnly(Window owner) => MessageBox.Show(owner, "Capturing an actual plan asks the service to re-execute the query, which it does by running a " + - "command — a read-only viewer seat can't enqueue commands. Reconnect with a read-write profile to " + - "capture actual plans.", + "command — a read-only viewer seat can't enqueue commands. The command is queued in the MONITORING " + + "STORE, not on the monitored server. Reconnect with a read-write store profile to capture actual " + + "plans.", "Read-Only Viewer", MessageBoxButton.OK, MessageBoxImage.Information); } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerAppSettings.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerAppSettings.cs index 7a28fac20..ef5073bf0 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerAppSettings.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerAppSettings.cs @@ -8,7 +8,6 @@ using System; using System.Collections.Generic; -using System.Diagnostics; using System.Globalization; using System.IO; using System.Text.Json; @@ -262,6 +261,9 @@ public sealed class ViewerAppSettingsStore { private static readonly JsonSerializerOptions s_jsonOptions = new() { WriteIndented = true }; + /// The source every diagnostic from this store is filed under. + private const string LogSource = "ViewerAppSettingsStore"; + private readonly string _filePath; /// Override the on-disk location (tests pass a temp file); null uses . @@ -273,6 +275,28 @@ public ViewerAppSettingsStore(string? filePath = null) /// The resolved settings file path (surfaced mainly so tests and diagnostics can name it). public string FilePath => _filePath; + /// + /// What the last on this instance found: for + /// a first run, for an ordinary one, and + /// when defaults were substituted for a file that is still + /// on disk. reads it to raise ONE startup dialog for the last case; every + /// other caller can ignore it, because has already logged. + /// + public SettingsFileState LastLoadState { get; private set; } = SettingsFileState.Absent; + + /// Why the last could not read the file — the line and position of a parse + /// error where there is one — or null when there was nothing wrong with it. + public string? LastLoadProblem { get; private set; } + + /// + /// The settings the last could not read, by NAME, and empty when there were none + /// (#2456). Non-empty means the rest of the file loaded normally and only these reverted to their + /// defaults — the distinction alone cannot make, and the reason the + /// startup dialog can now say which settings were lost instead of only where the parse stopped. + /// + public IReadOnlyList LastLoadUnreadableMembers { get; private set; } = + Array.Empty(); + /// %APPDATA%\PerformanceMonitorDarling\viewer-settings.json. public static string DefaultFilePath() { @@ -283,41 +307,30 @@ public static string DefaultFilePath() } /// - /// Reads the settings, returning defaults when the file does not exist yet or cannot be read/parsed — - /// the viewer never blocks on a first run or a corrupt file. Loaded values are normalized into range. + /// Reads the settings, returning defaults when the file does not exist yet or cannot be read/parsed. + /// Loaded values are normalized into range. A file that is present and unreadable is reported to the + /// log and left on — it is never treated as a first run (#2434). /// public ViewerAppSettings Load() { - try - { - if (!File.Exists(_filePath)) - { - return new ViewerAppSettings(); - } - - var json = File.ReadAllText(_filePath); - var settings = JsonSerializer.Deserialize(json, s_jsonOptions); - return (settings ?? new ViewerAppSettings()).Normalize(); - } - catch (Exception ex) - { - Debug.WriteLine($"ViewerAppSettingsStore: failed to load '{_filePath}', using defaults: {ex.Message}"); - return new ViewerAppSettings(); - } + var read = ViewerSettingsFile.Load(_filePath, LogSource, s_jsonOptions); + LastLoadState = read.State; + LastLoadProblem = read.Problem; + LastLoadUnreadableMembers = read.UnreadableMembers ?? Array.Empty(); + return read.Value!.Normalize(); } - /// Writes the settings as indented JSON, creating the app-data directory on first save. - public void Save(ViewerAppSettings settings) + /// + /// Writes the settings as indented JSON, creating the app-data directory on first save, and returns + /// whether the write happened. False means nothing reached disk — either the write itself failed, or + /// the existing file could not be read AND could not be copied aside, in which case it is deliberately + /// left exactly as it is rather than replaced with defaults (#2434). + /// + public bool Save(ViewerAppSettings settings) { ArgumentNullException.ThrowIfNull(settings); - var directory = Path.GetDirectoryName(_filePath); - if (!string.IsNullOrEmpty(directory)) - { - Directory.CreateDirectory(directory); - } - - File.WriteAllText(_filePath, JsonSerializer.Serialize(settings, s_jsonOptions)); + return ViewerSettingsFile.Save(_filePath, settings, LogSource, s_jsonOptions); } } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertSettings.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertSettings.cs index 2e69f1ff6..6b744fa77 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertSettings.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.AlertSettings.cs @@ -60,7 +60,10 @@ notify toggle (V20) are appended so the existing ordinals stay pinned. */ "pvs_enabled, pvs_threshold_percent, pvs_floor_gb, database_state_enabled, " + "self_disk_free_warn_percent, collection_stale_minutes, collection_failure_threshold, " + "disk_critical_free_percent, disk_critical_free_gb, analysis_notify_cooldown_minutes, " + - "store_job_cadence_warn_percent"; + "store_job_cadence_warn_percent, " + + /* #2391: V79 (#2349). APPENDED — this list drives the SELECT ordinals AND the upsert + parameter positions, so inserting anywhere but the end re-maps both at once. */ + "file_growth_enabled, file_growth_rise_mb, file_growth_volume_percent, file_growth_lookback_minutes"; /// The single global alert-settings row (id=1), for the Settings window prefill + the migrate-in /// defaults check. Column order matches . @@ -74,7 +77,7 @@ notify toggle (V20) are appended so the existing ordinals stay pinned. */ INSERT INTO config_alert_settings (id, " + AlertSettingsColumns + @", modified_at) VALUES (1, $1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13, $14, $15, $16, $17, $18, $19, $20, $21, $22, $23, $24, $25, $26, $27, $28, $29, $30, $31, $32, $33, $34, $35, $36, $37, $38, $39, $40, $41, $42, $43, - $44, $45, $46, $47, $48, $49, $50, $51, $52, $53, $54, + $44, $45, $46, $47, $48, $49, $50, $51, $52, $53, $54, $55, $56, $57, $58, (now() AT TIME ZONE 'UTC')) ON CONFLICT (id) DO UPDATE SET enabled = EXCLUDED.enabled, @@ -131,6 +134,10 @@ ON CONFLICT (id) DO UPDATE SET disk_critical_free_gb = EXCLUDED.disk_critical_free_gb, analysis_notify_cooldown_minutes = EXCLUDED.analysis_notify_cooldown_minutes, store_job_cadence_warn_percent = EXCLUDED.store_job_cadence_warn_percent, + file_growth_enabled = EXCLUDED.file_growth_enabled, + file_growth_rise_mb = EXCLUDED.file_growth_rise_mb, + file_growth_volume_percent = EXCLUDED.file_growth_volume_percent, + file_growth_lookback_minutes = EXCLUDED.file_growth_lookback_minutes, modified_at = (now() AT TIME ZONE 'UTC')"; /// The two cpu_mode values the service honors (it compares case-insensitively against @@ -214,6 +221,10 @@ private static void BindAlertSettings(NpgsqlCommand command, AlertSettingsRow r) command.Parameters.Add(new NpgsqlParameter { TypedValue = r.DiskCriticalFreeGb }); // $52 (#2107, V55) command.Parameters.Add(new NpgsqlParameter { TypedValue = r.AnalysisNotifyCooldownMinutes }); // $53 (#2107, V55) command.Parameters.Add(new NpgsqlParameter { TypedValue = r.StoreJobCadenceWarnPercent }); // $54 (#2136, V57) + command.Parameters.Add(new NpgsqlParameter { TypedValue = r.FileGrowthEnabled }); // $55 (#2349/#2391, V79) + command.Parameters.Add(new NpgsqlParameter { TypedValue = r.FileGrowthRiseMb }); // $56 (#2349/#2391, V79) + command.Parameters.Add(new NpgsqlParameter { TypedValue = r.FileGrowthVolumePercent }); // $57 (#2349/#2391, V79) + command.Parameters.Add(new NpgsqlParameter { TypedValue = r.FileGrowthLookbackMinutes }); // $58 (#2349/#2391, V79) } private static AlertSettingsRow ReadAlertSettingsRow(NpgsqlDataReader reader) => new() @@ -279,6 +290,10 @@ private static void BindAlertSettings(NpgsqlCommand command, AlertSettingsRow r) DiskCriticalFreeGb = reader.GetInt32(51), AnalysisNotifyCooldownMinutes = reader.GetInt32(52), StoreJobCadenceWarnPercent = reader.GetInt32(53), + FileGrowthEnabled = reader.GetBoolean(54), + FileGrowthRiseMb = reader.GetInt32(55), + FileGrowthVolumePercent = reader.GetInt32(56), + FileGrowthLookbackMinutes = reader.GetInt32(57), }; /// Maps the Settings window's CPU-mode combo tag ("Total"/"SqlOnly") to the store value. @@ -339,6 +354,13 @@ public sealed class AlertSettingsRow public int AnalysisNotifyCooldownMinutes { get; set; } = 360; public int StoreJobCadenceWarnPercent { get; set; } = 25; + /* #2391: defaults mirror the V79 column defaults, so a viewer prefilling against a store that has + not seeded the row shows what the store would have given it. Ships OFF, per #2349. */ + public bool FileGrowthEnabled { get; set; } + public int FileGrowthRiseMb { get; set; } = 10240; + public int FileGrowthVolumePercent { get; set; } = 60; + public int FileGrowthLookbackMinutes { get; set; } = 60; + public bool CpuEnabled { get; set; } = true; public int CpuThresholdPercent { get; set; } = 80; diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Blocking.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Blocking.cs index e5d9aa067..d85868bb5 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Blocking.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Blocking.cs @@ -367,8 +367,10 @@ internal async Task> GetBlockingPairRowsAsync( // Always-on DMV blocking snapshot: merge in the fallback rows so the viewer works even when the // blocked-process-report XE captured nothing (threshold unset / AWS RDS). Same connection. + /* #2443: the viewer passes its own window token — this read serves a person waiting at a + grid, not an analysis pass, so there is no budget or service stop for it to abandon under. */ await PgBlockingPairRowQuery.AppendDmvSnapshotRowsAsync( - connection.CreateCommand, rows, serverId, startUtc, endUtc); + connection.CreateCommand, rows, serverId, startUtc, endUtc, cancellationToken); return rows; } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.BlockingStats.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.BlockingStats.cs index 28181db6d..0b614a5a5 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.BlockingStats.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.BlockingStats.cs @@ -28,21 +28,6 @@ namespace PerformanceMonitor.Darling.Viewer; public sealed record BlockingDurationStatsPoint( DateTime Time, int EventCount, long TotalDurationMs, long MaxDurationMs, double AvgDurationMs); -/// -/// One per-minute bucket of deadlock SEVERITY (the Blocking Stats sub-tab's deadlock companion to -/// ): the count of deadlock VICTIM processes in the bucket plus the -/// total / max / avg deadlock wait_time_ms over EVERY process in the bucket's graphs. The on-the-fly -/// Darling equivalent of the Dashboard's pre-aggregated collect.blocking_deadlock_stats -/// (victim_count / total_deadlock_wait_time_ms) — computed by parsing the -/// deadlock_graph_xml the store already keeps, so (like the blocking-duration aggregate) it needs no -/// collector and no schema change. Matches the Dashboard analyzer's exact semantics -/// (26_blocking_deadlock_analyzer.sql: victim_count = SUM(CASE WHEN … VICTIM), -/// total_deadlock_wait_time_ms = SUM(every process's wait_time)) via the shared -/// 's per-process / -/// . -/// -public sealed record DeadlockSeverityStatsPoint( - DateTime Time, int VictimCount, long TotalWaitMs, long MaxWaitMs, double AvgWaitMs); public sealed partial class ViewerDataService { @@ -178,72 +163,13 @@ public async Task> GetDeadlockSeverityStatsAsyn } /// - /// Parses each deadlock graph via the shared and rolls the processes up - /// into per-minute severity buckets (bucketed on deadlock_time truncated to the minute, matching the - /// count trend's DATE_TRUNC('minute', deadlock_time)). Per bucket: victim_count = SUM of - /// , and total / max / avg deadlock wait over EVERY process's - /// (avg is process-weighted = total ÷ process count) — the - /// Dashboard collect.blocking_deadlock_stats semantics. A graph with no parseable process (empty / - /// malformed XML) or a null deadlock_time (unplaceable on the time axis) contributes nothing. - /// Pure + off-WPF so Darling.Tests pin it without a live store; the ORDER BY is re-applied here because the - /// dictionary rollup does not preserve the read order. + /// Delegates to in Common. + /// The arithmetic moved there so the headless service's Blocking Stats endpoint (#2484) could + /// reach it. Two copies of "what counts as a victim" is how two surfaces end up quietly disagreeing + /// about the same deadlock. This entry point stays because the pins that cover the aggregation are + /// written against it. /// internal static List AggregateDeadlockSeverity( - IReadOnlyList<(DateTime? DeadlockTime, string? Xml)> graphs) - { - var buckets = new Dictionary(); - - foreach (var (deadlockTime, xml) in graphs) - { - if (deadlockTime is not { } dt) - continue; - - var model = DeadlockGraphParser.Parse(xml); - if (model.IsEmpty) - continue; - - var bucket = TruncateToMinute(dt); - if (!buckets.TryGetValue(bucket, out var acc)) - { - acc = new DeadlockSeverityAccumulator(); - buckets[bucket] = acc; - } - - foreach (var process in model.Processes) - { - if (process.IsVictim) - acc.VictimCount++; - acc.TotalWaitMs += process.WaitTimeMs; - if (process.WaitTimeMs > acc.MaxWaitMs) - acc.MaxWaitMs = process.WaitTimeMs; - acc.ProcessCount++; - } - } - - return buckets - .OrderBy(kvp => kvp.Key) - .Select(kvp => new DeadlockSeverityStatsPoint( - kvp.Key, - kvp.Value.VictimCount, - kvp.Value.TotalWaitMs, - kvp.Value.MaxWaitMs, - kvp.Value.ProcessCount > 0 ? (double)kvp.Value.TotalWaitMs / kvp.Value.ProcessCount : 0.0)) - .ToList(); - } - - /// Mutable per-bucket accumulator for (a reference type so - /// the TryGetValue handle mutates the stored instance in place — one alloc per populated minute, - /// which is negligible at deadlock volumes). - private sealed class DeadlockSeverityAccumulator - { - public int VictimCount; - public long TotalWaitMs; - public long MaxWaitMs; - public int ProcessCount; - } - - /// Truncates to the minute to match Postgres DATE_TRUNC('minute', …) (the count trend's - /// bucketing), preserving the value's (naive UTC from the store). - private static DateTime TruncateToMinute(DateTime value) => - new(value.Ticks - (value.Ticks % TimeSpan.TicksPerMinute), value.Kind); + IReadOnlyList<(DateTime? DeadlockTime, string? Xml)> graphs) => + DeadlockSeverityAggregator.Aggregate(graphs); } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Fleet.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Fleet.cs index d81bb0c60..2ce4d56f1 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Fleet.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Fleet.cs @@ -187,7 +187,17 @@ private static SolidColorBrush MakeBrush(string hex) /// public sealed class FleetRollup { - /// The default depth of the worst-first ranking (the "Needs attention" list caps at this). + /// + /// The default depth of the worst-first ranking (the "Needs attention" list caps at this). + /// + /// Deliberately small, and deliberately unchanged by #2424 even though a 57-server fleet can leave + /// 52 servers behind it. The ranking renders inside the fleet roll-up panel, which is docked to the top + /// of the Overview and does NOT scroll — only the card grid beneath it does. A list that grew with the + /// fleet would therefore push the cards it is pointing at off the screen, and it would stop being a + /// triage shortlist ("look at these first") and become a second, worse copy of the grid without the + /// metrics. The overflow is answered by giving it somewhere to go instead — the needs-attention card + /// filter, over the grid that scrolls and carries the six metric rows. + /// public const int DefaultWorstCount = 5; /// An empty fleet (no servers registered) — every count zero, no ranking. @@ -211,7 +221,12 @@ public sealed class FleetRollup /// The worst-first problem servers (band != Healthy), capped at the requested depth. public IReadOnlyList WorstServers { get; init; } = Array.Empty(); - /// Problem servers beyond the capped ranking — surfaced as a "+N more" affordance. + /// + /// Problem servers beyond the capped ranking. Surfaced as the "+N more need attention" affordance, + /// which is a LINK into the needs-attention card filter rather than a dead count (#2424): the ranking + /// deliberately stays short (see ), so this number is the whole reason + /// the filter exists. + /// public int AdditionalProblemCount { get; init; } /// Any server needs attention (band != Healthy) — drives the ranking list vs the all-clear line. @@ -321,7 +336,14 @@ public static FleetRollup Build(IReadOnlyList summaries, Flee /// Warning too; otherwise Healthy. No new thresholds are introduced here. /// public static FleetHealthBand ClassifyBand(ServerSummaryItem s) => - ServerHealthClassifier.ClassifyBand(s.IsOnline, s.AwaitingFirstCollection, s.HasCollectorErrors, s.OverallMetricSeverity); + ServerHealthClassifier.ClassifyBand( + s.IsOnline, + /* Via the card's discriminant, not the raw flag. The shared classifier honours an awaiting marker + whatever IsOnline says, so an online card carrying a stray marker banded Warning while the card + said "Online" and had nothing to report — a third reading of the same pair. See ServerCollectionStatus. */ + s.CardStatus == ServerCollectionStatus.AwaitingFirstCollection, + s.HasCollectorErrors, + s.OverallMetricSeverity); /// /// The worst-first ordering score — the SHARED over @@ -339,14 +361,19 @@ public static long FleetHealthScore(ServerSummaryItem s) => /// public static string BuildReason(ServerSummaryItem s) { - if (s.IsOnline == false) + /* Keyed on the card's own status discriminant rather than on the flags behind it. Reading + AwaitingFirstCollection independently of IsOnline is what let an online card claim it was awaiting + its first collection — see ServerCollectionStatus. */ + if (s.CardStatus == ServerCollectionStatus.Offline) { return "Offline — no recent collection"; } - if (s.AwaitingFirstCollection) + if (s.CardStatus == ServerCollectionStatus.AwaitingFirstCollection) { - return "Awaiting first collection"; + /* The word itself, not a copy of it — this line held the fourth spelling of the phrase, in the + fourth file, which is exactly the shape the #2473 pin now forbids. */ + return s.CardStatus.Word(); } var parts = new List(); @@ -380,6 +407,126 @@ public static string BuildReason(ServerSummaryItem s) parts.Add("collection stale"); } - return parts.Count > 0 ? string.Join(", ", parts) : "Needs attention"; + return parts.Count > 0 ? string.Join(", ", parts) : UnspecifiedReason; + } + + /// + /// What answers when it can name nothing — a card banded away from Healthy by a + /// severity whose display the reason does not cover. It reads fine in the ranking, where every row is a + /// problem server, and reads as an unexplained demand on a card, so the tooltip degrades to the band label + /// rather than repeating it. Named so the two sides cannot drift apart on the spelling. + /// + public const string UnspecifiedReason = "Needs attention"; + + /// The line every card tooltip ends on. The sidebar alert badge's tooltip has the same shape — + /// the breakdown, then how to act on it — so the Overview card is not the surface that explains itself + /// least; this one names the gesture the card actually supports. + private const string CardTooltipAction = "Double-click the card to open this server's tab"; + + /// + /// The Overview card's status tooltip (#2422): the band the card's border is painted from, WHY it is in + /// that band, and what to do next. Reported as "the card says Warning and will not say why" — the reader + /// had to scan six metric rows hunting for the amber one, once per card, on a 57-server fleet. + /// + /// It is 's output verbatim, the same sentence the Needs Attention ranking + /// already shows for that exact server. Re-deriving it here instead would forfeit the one property + /// BuildReason exists for: it is built from the card's OWN metric displays, so it cannot disagree with the + /// six rows the reader is looking at while they are looking at them. + /// + public static string BuildStatusTooltip(ServerSummaryItem s) + { + ArgumentNullException.ThrowIfNull(s); + + return Headline(s) + "\n" + CardTooltipAction; + } + + /// + /// The tooltip's first line: what this card's state IS, in the words the card's own status line uses, + /// followed by the reason when there is one to give. + /// + /// It switches on — the SAME discriminant + /// renders — rather than re-reading the flags underneath it. + /// A tooltip that hangs off a word and then contradicts it is the defect this change exists to remove, and + /// two independent readings of the same two flags is precisely how it comes back. + /// + private static string Headline(ServerSummaryItem s) + { + var band = ClassifyBand(s); + + return s.CardStatus switch + { + /* Offline and never-reached come back from BuildReason as whole sentences that already name the + state — the same sentence the status word shows — so a band label in front would only say + "Offline" twice. */ + ServerCollectionStatus.Offline or ServerCollectionStatus.AwaitingFirstCollection => BuildReason(s), + + /* "Unknown" is the one status word with no band behind it: ClassifyBand goes straight to the + metrics and, on a clean card, answers Healthy. The word wins, and the metrics are appended when + they have something to add — not knowing whether a server is reporting is no reason to withhold + the CPU number that WAS collected. */ + ServerCollectionStatus.Unknown => WithReason(UnknownStatus, "; ", s), + + /* Online and stale: the band is the headline. A healthy card gets an all-clear rather than + BuildReason's "Needs attention" fallback, which is written for a ranking that only ever holds + problem servers and on a grid showing EVERY server would say the opposite of the truth. */ + _ => band == FleetHealthBand.Healthy + ? "Healthy — every metric on this card is inside its threshold" + : WithReason(ServerHealthClassifier.BandLabel(band), " — ", s), + }; + } + + /// + /// A headline plus what the card can actually name — or the headline alone when it can name nothing. + /// + /// Every arm that appends a reason goes through here, because a card CAN sit outside Healthy with + /// nothing to say: a Blocking band raised by a long max-wait while the reason's own gate wants a non-zero + /// event count, for one. Appending unguarded produces "Warning — Needs attention", which tells the reader + /// exactly what they already knew and is how the ranking-only fallback reaches a card at all. Two arms had + /// their own copy of the append and only one of them was guarded, which is the same lesson as + /// one level down. + /// + private static string WithReason(string headline, string separator, ServerSummaryItem s) + { + var reason = BuildReason(s); + + return reason == UnspecifiedReason ? headline : headline + separator + reason; + } + + /// The words behind StatusDisplay's "Unknown" — a card whose freshness was never classified. + private const string UnknownStatus = "Unknown — no collection status for this server"; + + /// + /// The problem servers among the Overview's cards — band != Healthy — kept in the order the caller handed + /// them in, so the grid's chosen sort survives the filter. This is the SAME predicate + /// uses to decide who is in the ranking and who counts toward , which + /// is what makes "+52 more need attention" land on exactly 52 cards rather than on a second opinion. + /// + public static List NeedsAttention(IEnumerable summaries) + { + ArgumentNullException.ThrowIfNull(summaries); + + return summaries.Where(s => ClassifyBand(s) != FleetHealthBand.Healthy).ToList(); + } + + /// + /// The count shown beside the filter toggle while it is on. A filtered grid that looks like an unfiltered + /// one is a worse defect than the one the filter fixes, so the active state carries its own arithmetic — + /// including the all-clear case, which is otherwise an empty grid with nothing saying why. + /// + public static string AttentionFilterCountText(int shown, int total) + { + if (shown > 0) + { + return $"showing {shown} of {total}"; + } + + /* Nothing left to show. On a populated fleet that is the honest all-clear; with no cards at all it + must not report that zero servers are healthy, which is the sort of line that gets screenshotted. */ + return total switch + { + <= 0 => "no servers to filter", + 1 => "the 1 server monitored is healthy", + _ => $"all {total} servers are healthy", + }; } } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.MonitoredServers.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.MonitoredServers.cs index 51ebbd46a..aef759ef5 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.MonitoredServers.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.MonitoredServers.cs @@ -13,6 +13,7 @@ using System.Threading.Tasks; using Npgsql; using NpgsqlTypes; +using PerformanceMonitor.Collectors; using PerformanceMonitor.Common; using PerformanceMonitor.Notifications; @@ -176,6 +177,17 @@ AND database IS NOT DISTINCT FROM $2 /// viewer-added server appear immediately (before its first collection) and a removed one disappear at /// once, instead of waiting for the service's next reconcile. is_enabled and monthly_cost_usd /// come from config so a viewer toggle/edit is reflected without a round-trip through the service. + /// + /// The engine discriminator comes from the OBSERVED side (#2530). engine_kind and + /// sql_engine_edition are read from collect.servers, not from + /// config_monitored_servers.engine, which is the DESIRED configuration and cannot carry + /// Aurora-ness at all - that is probed from aurora_version at connect. A server the operator + /// added but the service has not connected to yet therefore has NO kind, which is the honest answer: + /// it gets the SQL Server tab set by default, exactly as it did before this column existed. + /// + /// This query, not ServersSql, is what the sidebar uses on any seeded store - i.e. every + /// real deployment - so the discriminator has to be on BOTH or the viewer would have kept rendering + /// SQL Server tabs at every PostgreSQL target while a unit test over the other query passed. /// public const string ManagedServersSql = @" SELECT @@ -184,7 +196,9 @@ AND database IS NOT DISTINCT FROM $2 COALESCE(s.display_name, c.name) AS display_name, c.is_enabled, s.sql_major_version, - c.monthly_cost_usd + c.monthly_cost_usd, + s.engine_kind, + COALESCE(s.sql_engine_edition, 0) AS sql_engine_edition FROM config_monitored_servers c LEFT JOIN servers s ON s.server_id = c.server_id ORDER BY COALESCE(s.display_name, c.name)"; @@ -215,7 +229,9 @@ public async Task> GetManagedServersAsync(CancellationToken reader.IsDBNull(2) ? serverName : reader.GetString(2), !reader.IsDBNull(3) && reader.GetBoolean(3), reader.IsDBNull(4) ? null : reader.GetInt32(4), - reader.IsDBNull(5) ? 0m : Convert.ToDecimal(reader.GetValue(5)))); + reader.IsDBNull(5) ? 0m : Convert.ToDecimal(reader.GetValue(5)), + reader.IsDBNull(6) ? null : reader.GetString(6), + reader.IsDBNull(7) ? CollectorEngineCapability.UnknownEngineEdition : reader.GetInt32(7))); } return servers; diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Overview.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Overview.cs index 5836ca77a..ec2f8a1e2 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Overview.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Overview.cs @@ -344,6 +344,16 @@ number alongside (Lite's headline). */ instantUtc.HasValue ? Math.Max(0, (int)(nowUtc - instantUtc.Value).TotalMinutes) : null; } +/* The card's status discriminant used to be declared here, as ServerCardStatus. It is now + PerformanceMonitor.Common's ServerCollectionStatus, because three OTHER places derived the same ladder and + two of them are in the headless service, which cannot reference WPF (#2473). The argument for having one + discriminant at all is unchanged and is written down on the enum; what changed is how far "one" reaches. + + The rename is not cosmetic. Lite has its own ServerCardStatus meaning a CONNECTION check, and #2457 kept + the two axes apart on purpose; giving the shared type the collection name (beside Common's existing + ServerConnectionStatus) means the two can no longer be confused for each other by a reader or by a + using directive. */ + /// /// One Overview server card's view-model — copied from Lite's ServerSummaryItem /// (Lite/Services/LocalDataService.Overview.cs) and enriched toward the Dashboard's @@ -559,24 +569,46 @@ public string ThreadsDisplay ? ViewerTimeHelper.ForDisplay(LastCollectionTime.Value).ToString("HH:mm:ss") : "Never"; - /* Connection status — verbatim from Lite; in the viewer the inputs come from ApplyFreshness. + /* Collection status. The (IsOnline, HasCollectorErrors, AwaitingFirstCollection) triple is resolved by + ServerCollectionStatusRules.Classify and nowhere else in the viewer — the sidebar row's dot carried its + own four-state copy until #2473, which is how a never-collected server got a grey "Unknown" dot beside + this card's amber "Awaiting first collection". Everything downstream — the word, the colour, the + ranking's reason, the tooltip, the sidebar dot — renders the resulting ServerCollectionStatus, so no + two of them can land on different answers for the same server. See ServerCollectionStatus for the + contradictions that motivated collapsing it. + The null arm distinguishes "not reached yet" (bootstrap) from a legacy unknown. */ - public string StatusDisplay => IsOnline switch - { - true when HasCollectorErrors => "Warning", - true => "Online", - false => "Offline", - _ => AwaitingFirstCollection ? "Awaiting first collection" : "Unknown" - }; + public ServerCollectionStatus CardStatus => + ServerCollectionStatusRules.Classify(IsOnline, HasCollectorErrors, AwaitingFirstCollection); + + public string StatusDisplay => CardStatus.Word(); - public SolidColorBrush StatusBrush => MakeBrush(IsOnline switch + /* The palette stays here rather than moving to the rules class: these are the viewer's dark-theme hexes, + and the sidebar dot paints the same states from the THEME dictionaries instead (a DynamicResource, so + it follows the light / cool-breeze themes the cards do not). The states agree; only the colour source + differs, and ViewerSidebarDotRendersTheCardStatusTests pins that every state the card paints has a + trigger on the dot. */ + public SolidColorBrush StatusBrush => MakeBrush(CardStatus switch { - true when HasCollectorErrors => "#FFD54F", // amber — stale collection - true => "#81C784", - false => "#E57373", - _ => AwaitingFirstCollection ? "#FFD54F" : "#888888" // amber — queued, not dead + ServerCollectionStatus.Stale => "#FFD54F", // amber — stale collection + ServerCollectionStatus.Online => "#81C784", + ServerCollectionStatus.Offline => "#E57373", + ServerCollectionStatus.AwaitingFirstCollection => "#FFD54F", // amber — queued, not dead + _ => "#888888", }); + /// + /// What the status word MEANS on this card, for the tooltip the status line carries (#2422). A colour and + /// a one-word band were the whole answer the Overview gave, and the reporter's question — "what is it that + /// this text warns me about?" — is one the card could already answer: + /// builds the sentence out of THIS card's own metric displays, and until now only the Needs Attention + /// ranking got to see it. + /// + /// Delegated rather than reimplemented on purpose: two independent derivations of "why is this amber" + /// would eventually disagree, and the one place they would disagree is a card the reader is staring at. + /// + public string StatusTooltip => FleetRollup.BuildStatusTooltip(this); + public bool IsOffline => IsOnline == false; // ── Per-metric severity bands (delegated to the SHARED ServerHealthClassifier — one place for the @@ -663,25 +695,22 @@ public static ServerFreshness ClassifyFreshness(DateTime? lastCollectionUtc, Dat ServerHealthClassifier.ClassifyFreshness(lastCollectionUtc, nowUtc); /// - /// Maps the freshness band onto Lite's card inputs, taking the live-ping's place: Fresh → Online, - /// Stale → the amber Warning state, Offline → the red Offline overlay, NeverCollected → the amber - /// "Awaiting first collection" state (IsOnline stays null: the truth is "unknown, not reached yet", - /// not "was up and died"). + /// Maps the freshness band onto the card's three status flags, taking the live-ping's place: Fresh → + /// Online, Stale → the amber Warning state, Offline → the red Offline overlay, NeverCollected → the amber + /// "Awaiting first collection" state (IsOnline stays null: the truth is "unknown, not reached yet", not + /// "was up and died"). + /// + /// The mapping itself is , shared with the sidebar + /// row and the service's fleet reader. It was written out here in longhand, and the sidebar's longhand + /// copy set two of the three flags and dropped AwaitingFirstCollection — an omission that is + /// invisible in a block of assignments and impossible when the three arrive together (#2473). /// public void ApplyFreshness(DateTime nowUtc) { - var freshness = ClassifyFreshness(LastCollectionTime, nowUtc); - if (freshness == ServerFreshness.NeverCollected) - { - IsOnline = null; - HasCollectorErrors = false; - AwaitingFirstCollection = true; - return; - } - - AwaitingFirstCollection = false; - IsOnline = freshness != ServerFreshness.Offline; - HasCollectorErrors = freshness == ServerFreshness.Stale; + var flags = ServerCollectionStatusRules.FlagsFor(ClassifyFreshness(LastCollectionTime, nowUtc)); + IsOnline = flags.IsOnline; + HasCollectorErrors = flags.HasCollectorErrors; + AwaitingFirstCollection = flags.AwaitingFirstCollection; } private static SolidColorBrush SeverityBrush(HealthSeverity severity) => severity switch diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Postgres.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Postgres.cs new file mode 100644 index 000000000..ef24e69dc --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.Postgres.cs @@ -0,0 +1,421 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Npgsql; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// The viewer's PostgreSQL reader layer (#2530) — what the six PostgreSQL inner tabs +/// () read, and the projections that make stored rows renderable. +/// +/// The queries are not here, deliberately. Every PostgreSQL read runs the SAME SQL the MCP +/// surface runs, from DarlingPg*Reader in PerformanceMonitor.Darling.Storage — moved there +/// from the service's MCP folder by this change so both front ends can reach them. Copying them into the +/// viewer would have meant a second copy of, among others, a 200-line recursive blocking walk whose revisit +/// guard, root attribution and truncation flag were each a separate review finding; two copies of that +/// diverge, and the copy that diverges is never the one being read. Every other seam in this repo that both +/// SKUs answer follows the same rule (CollectorEngineCapability's message, DescribeEngineKind, +/// the alert gates), and the reason is always this one. +/// +/// What IS here: the per-collector health read the Overview tab needs — which has no MCP twin +/// because the MCP answers it per-read rather than per-server — and the display projections. Those exist +/// because the stored rows carry sentinels and units a grid must not render raw: -1 means "not +/// measured" rather than "zero", pg_stat_io's write side is genuinely absent on Aurora rather than +/// zero, and every timestamp in the store is naive UTC that has to pass through +/// like every other timestamp the viewer shows. +/// +public sealed partial class ViewerDataService +{ + /// + /// Per-collector collection facts for one server over the window — the Overview tab's grid. + /// + /// Scoped by an explicit collector-name array rather than by a LIKE 'pg\_%' pattern: the + /// caller passes the names says are PostgreSQL collectors, so the set + /// tracks the catalog instead of a naming convention a future collector could break. Cast to + /// text[] explicitly — Npgsql's inference through = ANY($4) is a runtime failure, not a + /// compile one. + /// + /// A collector this engine can never run has NO row here at all, by design: dispatch filters it + /// out before it runs, so it writes no collection_log entry (a fake SUCCESS/0-rows would be + /// ~2,880 rows a day per server of noise). That absence is exactly why + /// composes this against the catalog rather than rendering it + /// directly — otherwise the one collector an operator most needs explained is the one row missing. + /// + /// $1 server_id, $2/$3 window (naive UTC), $4 collector names. + /// + public const string PostgresCollectorHealthSql = """ + SELECT + collector_name AS collector_name, + MAX(collection_time) AS last_run_at, + MAX(collection_time) FILTER (WHERE status = 'SUCCESS') AS last_success_at, + CAST(COUNT(*) AS bigint) AS runs, + CAST(COUNT(*) FILTER (WHERE status <> 'SUCCESS') AS bigint) AS failed_runs, + CAST(COALESCE(SUM(rows_collected), 0) AS bigint) AS rows_collected, + (ARRAY_AGG(status ORDER BY collection_time DESC))[1] AS last_status, + /* The newest NON-EMPTY message, not the newest message. A collector that failed an hour ago and + has succeeded quietly since would otherwise show a blank explanation for a non-zero failure + count, which reads as the grid having nothing to say about it. */ + (ARRAY_AGG(error_message ORDER BY (error_message IS NULL), collection_time DESC))[1] + AS last_message + FROM v_collection_log + WHERE server_id = $1 + AND collection_time >= $2 + AND collection_time <= $3 + AND collector_name = ANY($4::text[]) + GROUP BY collector_name + """; + + /// One collector's raw collection_log facts over the window, before the catalog join. + public sealed record PostgresCollectorLogFacts( + string CollectorName, + DateTime? LastRunAt, + DateTime? LastSuccessAt, + long Runs, + long FailedRuns, + long RowsCollected, + string? LastStatus, + string? LastMessage); + + /// Runs for the PostgreSQL collectors the catalog ships. + public async Task> GetPostgresCollectorLogFactsAsync( + int serverId, DateTime startUtc, DateTime endUtc, IReadOnlyList collectorNames, + CancellationToken cancellationToken = default) + { + var facts = new List(); + if (collectorNames.Count == 0) + { + return facts; + } + + await using var command = _dataSource.CreateCommand(PostgresCollectorHealthSql); + command.Parameters.Add(new NpgsqlParameter { TypedValue = serverId }); + /* Kind-Unspecified at the bind, per the store's naive-UTC discipline: a Kind=Utc DateTime makes + Npgsql infer timestamptz, and PostgreSQL then zone-shifts these naive columns to compare them, so + east of UTC the window slides off the data and the read silently returns nothing. */ + command.Parameters.Add(new NpgsqlParameter { TypedValue = DateTime.SpecifyKind(startUtc, DateTimeKind.Unspecified) }); + command.Parameters.Add(new NpgsqlParameter { TypedValue = DateTime.SpecifyKind(endUtc, DateTimeKind.Unspecified) }); + command.Parameters.Add(new NpgsqlParameter { TypedValue = collectorNames.ToArray() }); + + await using var reader = await command.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) + { + facts.Add(new PostgresCollectorLogFacts( + reader.GetString(0), + reader.IsDBNull(1) ? null : reader.GetDateTime(1), + reader.IsDBNull(2) ? null : reader.GetDateTime(2), + reader.IsDBNull(3) ? 0 : reader.GetInt64(3), + reader.IsDBNull(4) ? 0 : reader.GetInt64(4), + reader.IsDBNull(5) ? 0 : reader.GetInt64(5), + reader.IsDBNull(6) ? null : reader.GetString(6), + reader.IsDBNull(7) ? null : reader.GetString(7))); + } + + return facts; + } + + /// One row of the Overview tab's grid: a PostgreSQL collector and what it did for this server. + public sealed class PostgresCollectorHealthRow + { + public string Collector { get; init; } = ""; + public string StoreTable { get; init; } = ""; + public string Status { get; init; } = ""; + public string LastRun { get; init; } = ""; + public long Runs { get; init; } + public long FailedRuns { get; init; } + public long RowsCollected { get; init; } + public string Explanation { get; init; } = ""; + + /// Drives the grid's row highlight: true for a collector this engine can never run, which + /// is a fact to read rather than a fault to chase. + public bool IsPermanentGap { get; init; } + + /// Drives the grid's row highlight for a collector that IS supposed to run and is not. + public bool IsFault { get; init; } + } + + /// + /// Composes the Overview grid from the catalog and the log — pure, so it is unit-tested without a store + /// or a window. + /// + /// Catalog-driven, not log-driven, and that is the whole design. Rows come from the nine + /// PostgreSQL collectors ships; the log only fills them in. A collector + /// with no log row is therefore VISIBLE, with the reason: on stock PostgreSQL the two Aurora-only + /// collectors read aurora_stat_system_waits() and + /// aurora_stat_statements(), which core PostgreSQL has in no version, and + /// says so in the one sentence the MCP + /// surface and the web dashboard also print. The defect #2530 is about is unexplained emptiness; a + /// grid that just dropped those two would have reproduced it one layer down. + /// + /// The three states are deliberately distinct, because they need three different responses: a + /// permanent engine gap is nothing to do, a collector that should be running and has no rows in the + /// window is something to chase, and a collector that ran and failed carries its own error text. + /// + public static List BuildPostgresCollectorHealth( + DarlingServer server, + IReadOnlyList postgresCollectors, + IReadOnlyList logFacts) + { + ArgumentNullException.ThrowIfNull(server); + ArgumentNullException.ThrowIfNull(postgresCollectors); + ArgumentNullException.ThrowIfNull(logFacts); + + var byName = logFacts.ToDictionary(f => f.CollectorName, StringComparer.Ordinal); + var rows = new List(); + + foreach (var definition in postgresCollectors) + { + var gap = CollectorEngineCapability.NotCollectedMessage( + server.ServerName, server.EngineEdition, server.EngineKind, definition.Name); + + byName.TryGetValue(definition.Name, out var facts); + + string status; + var isFault = false; + if (gap is not null) + { + status = "Not collected on this engine"; + } + else if (facts is null) + { + status = "No runs in this window"; + isFault = true; + } + else if (facts.FailedRuns > 0) + { + status = string.Equals(facts.LastStatus, "SUCCESS", StringComparison.Ordinal) + ? "Recovered" + : facts.LastStatus ?? "Failing"; + isFault = !string.Equals(facts.LastStatus, "SUCCESS", StringComparison.Ordinal); + } + else + { + status = "Collecting"; + } + + rows.Add(new PostgresCollectorHealthRow + { + Collector = definition.Name, + StoreTable = definition.TargetTable, + Status = status, + LastRun = facts?.LastRunAt is { } at + ? ViewerTimeHelper.ForDisplay(at).ToString("yyyy-MM-dd HH:mm", CultureInfo.CurrentCulture) + : "", + Runs = facts?.Runs ?? 0, + FailedRuns = facts?.FailedRuns ?? 0, + RowsCollected = facts?.RowsCollected ?? 0, + /* The engine sentence wins over the last error when both exist: a collector that cannot run + here has no error worth showing, and a stale one from before a migration would be the more + confusing of the two. */ + Explanation = gap ?? facts?.LastMessage ?? "", + IsPermanentGap = gap is not null, + IsFault = isFault, + }); + } + + return rows; + } + + // ───────────────────────────────────────────────────────────────────────────────────────────── + // The nine shared reads, as the viewer calls them. Thin on purpose: the SQL, the parameter binding + // and the ordinal mapping all live in PerformanceMonitor.Darling.Storage, shared byte-for-byte with + // the MCP surface, so these add a window and nothing else. + // ───────────────────────────────────────────────────────────────────────────────────────────── + + /// Vacuum tab, panel 1 — what is holding the xmin horizon back, by cause. + public Task> GetPgXminHorizonAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgXminReader.GetPgXminHorizonAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Vacuum tab, panel 2 — per-table autovacuum backlog, ranked by ratio to the table's own threshold. + public Task> GetPgAutovacuumAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 100, CancellationToken cancellationToken = default) => + DarlingPgAutovacuumReader.GetPgAutovacuumAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Vacuum tab, panel 3 — per-database transaction-ID and multixact freeze headroom. + public Task> GetPgWraparoundAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgWraparoundReader.GetPgWraparoundAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Waits tab — Aurora's cumulative wait counters, differenced over the window. + public Task> GetPgWaitStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgWaitReader.GetPgWaitStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// I/O tab — pg_stat_io per backend type / object / context, differenced over the window. + public Task> GetPgIoAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 100, CancellationToken cancellationToken = default) => + DarlingPgIoReader.GetPgIoAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Replication tab — slot WAL retention and the xmin each slot pins. + public Task> GetPgSlotsAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgSlotReader.GetPgSlotsAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Activity tab, panel 1 — the sampling DENOMINATOR: how many captures ran, and how many saw blocking. + public Task GetPgBlockingCaptureCountsAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgBlockingReader.GetPgBlockingCaptureCountsAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Activity tab, panel 2 — blocking chains by root. + public Task> GetPgBlockingChainsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgBlockingReader.GetPgBlockingChainsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Activity tab, panel 3 — lock cycles, which have no root and so appear on no chain. + public Task> GetPgBlockingCyclesAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgBlockingReader.GetPgBlockingCyclesAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Activity tab, panel 4 — top statement shapes by total execution time. + public Task> GetPgTopQueriesAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgStatementReader.GetPgTopQueriesAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Activity tab, panel 5 — per-database counters: temp-file spills, cache hits, deadlocks, commit split. + public Task> GetPgDatabaseStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgDatabaseReader.GetPgDatabaseStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Storage tab, panel 1 - the per-table bloat ESTIMATE with its measured sizes and the trust + /// signals that decide whether the estimate may be rendered at all. + public Task> GetPgTableBloatAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgTableBloatReader.GetPgTableBloatAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Storage tab, panel 2 - per-index scan counts with the catalog facts that decide whether an + /// unscanned index can actually go. + public Task> GetPgIndexUsageAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgIndexUsageReader.GetPgIndexUsageAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Vacuum tab, panel 4 - the sessions holding a transaction open, and which of them actually + /// pins the xmin horizon the panels above it measure. + public Task> GetPgSessionStatesAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgSessionStatesReader.GetPgSessionStatesAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Vacuum tab, panel 5 - whether this target can capture execution plans at all, and if not, + /// which step is missing. Latest state per facet rather than the history: every facet is a + /// parameter-group setting, so the window holds the same rows repeated. + public Task> GetPgPlanCaptureReadinessAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgPlanCaptureReadinessReader.GetPgPlanCaptureReadinessAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Waits tab, write-side panel - checkpoints, background writer and WAL over the window (#2544). + /// Returns a SINGLE row or null: the three source views are cluster-wide singletons, and what is reported + /// is the CHANGE across the window rather than the cumulative levels, so there is one answer rather than a + /// series. Null means fewer than two samples, which is a real state on a freshly added server and is not + /// the same as a quiet one. + public Task GetPgWriteStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, CancellationToken cancellationToken = default) => + DarlingPgWriteStatsReader.GetPgWriteStatsAsync(_dataSource, serverId, startUtc, endUtc, cancellationToken); + + /// Overview tab - the server's own configuration from pg_settings (#2658), non-default first. + /// Anchored on the newest snapshot rather than the toolbar window, deliberately: a configuration is the + /// state NOW, and an hours filter would return nothing for a server whose hourly collector last ran just + /// outside it - which reads as "this server has no configuration" rather than "ask again". + public Task> GetPgServerConfigAsync( + int serverId, int limit = 500, CancellationToken cancellationToken = default) => + DarlingPgServerConfigReader.GetCurrentConfigAsync(_dataSource, serverId, limit, cancellationToken); + + /// Activity tab - PostgreSQL deadlocks reported in the window (#2661), one row per distinct + /// report. Windowed on when the deadlock HAPPENED rather than when it was collected: the collector + /// re-reads an overlapping log tail, so a report is found minutes later and found again for as long as + /// it stays in the window, and filtering on collection time would place it wrongly and move it. + public Task> GetPgDeadlocksAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 100, + CancellationToken cancellationToken = default) => + DarlingPgDeadlockReader.GetDeadlocksAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Overview tab - which extensions this target has, could have, or cannot have (#2545). Latest + /// state per extension rather than the history: installing one is a rare deliberate act, so the window + /// holds the same answer repeated daily. Monitoring-relevant extensions sort first, and within them the + /// ACTIONABLE state (available but not installed) sorts above the rest. + public Task> GetPgExtensionAvailabilityAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 200, CancellationToken cancellationToken = default) => + DarlingPgExtensionAvailabilityReader.GetPgExtensionAvailabilityAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Activity tab - lock state by mode, type and relation over the window (#2544). Every row + /// carries the capture denominator, because these are SAMPLES: three ungranted rows means something + /// different in 60 captures than in 4. Ungranted sorts first, then by the worst wait anyone served. + public Task> GetPgLockStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgLockStatsReader.GetPgLockStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// + /// Wait events attributed to the query that waited (#2603). Deltas and the samples-to-milliseconds + /// inference both live in the reader, beside the numbers they derive from. + /// + public Task> GetPgWaitSamplingAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit, CancellationToken cancellationToken = default) => + DarlingPgWaitSamplingReader.GetPgWaitSamplingAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// + /// The kernel's own CPU and disk per query (#2603). Deltas and reset detection live in the reader. + /// + public Task> GetPgKernelStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit, CancellationToken cancellationToken = default) => + DarlingPgKernelStatsReader.GetPgKernelStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// + /// Which columns are filtered on and how badly the planner estimated them (#2603). Newest per + /// predicate rather than differenced — see the reader for why a rate is the wrong shape here. + /// + public Task> GetPgPredicateStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit, CancellationToken cancellationToken = default) => + DarlingPgPredicateStatsReader.GetPgPredicateStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// + /// Plans captured by auto_explain (#2566), grouped by shape. The JSON is redacted at collection — there + /// is no un-redacted copy in the store for this read to expose. + /// + public Task> GetPgPlanCaptureAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit, CancellationToken cancellationToken = default) => + DarlingPgPlanCaptureReader.GetPgPlanCaptureAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Storage tab - per-column planner statistics (#2543), latest per column and ranked by + /// suspicion rather than alphabetically: heavy skew first (the parameter-sensitivity signal), then low + /// correlation (why an index scan was rejected). Zero rows has TWO causes - no qualifying table, or a + /// monitoring role without SELECT, since pg_stats filters on has_column_privilege - and the caller must + /// not report either as healthy statistics. + public Task> GetPgColumnStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 100, CancellationToken cancellationToken = default) => + DarlingPgColumnStatsReader.GetPgColumnStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Replication tab - connected standbys and how far behind each got (#2544). Returns the latest + /// sample AND the window's worst, because a replica that drifts hundreds of MB behind and recovers reads + /// as healthy in any single sample - and it is the one most likely to be useless when somebody needs to + /// fail over to it. Ranked by worst REPLAY bytes, never by the time lag, which understates a stall. + public Task> GetPgReplicationStatsAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgReplicationStatsReader.GetPgReplicationStatsAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// I/O tab - what is resident in shared buffers (#2544). Latest snapshot only, because + /// residency is a level rather than a counter and averaging it across a day answers nothing. A NULL + /// relation name means another database's relation or a shared catalog, not a missing name. + public Task> GetPgBufferUsageAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgBufferUsageReader.GetPgBufferUsageAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); + + /// Storage tab - MEASURED b-tree index bloat (#2561), latest per index. Ranked by estimated + /// reclaimable BYTES rather than by density, because a small index at 40% density is worth nothing next + /// to a large one at 70%. Indexes too large to measure sort to the TOP: unknown is not zero, and they + /// are the likeliest big win. + public Task> GetPgIndexBloatAsync( + int serverId, DateTime startUtc, DateTime endUtc, int limit = 50, CancellationToken cancellationToken = default) => + DarlingPgIndexBloatReader.GetPgIndexBloatAsync(_dataSource, serverId, startUtc, endUtc, limit, cancellationToken); +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs index 48763ddb5..0c1461355 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor. @@ -14,6 +14,7 @@ using System.Threading; using System.Threading.Tasks; using Npgsql; +using PerformanceMonitor.Collectors; using PerformanceMonitor.Common; using PerformanceMonitor.Darling.Storage; @@ -23,21 +24,36 @@ namespace PerformanceMonitor.Darling.Viewer; /// One row of the servers table, as the viewer's server list shows it. Was a positional record; it is now /// a class because the ported Lite server-row chrome needs mutable, change-notifying runtime state on each /// row — the favorite star (, matched from the viewer's registry) and the -/// collection-freshness status dot ( / → -/// ) update in place on the refresh timers without resetting the list's selection. +/// collection-freshness status dot ( / / +/// → → ) update in +/// place on the refresh timers without resetting the list's selection. /// The Postgres-sourced fields stay immutable (get-only); only the sidebar overlay state is settable. /// Equality is now reference-based, which every consumer already relies on (they key on /// or hold the list's own instances). /// public sealed class DarlingServer : INotifyPropertyChanged { + /// + /// servers.engine_kind as the store holds it (#2530) - one of + /// , or null for a row no connect has stamped since the V82 rung + /// landed. Defaulted so the two reader call sites are the only places that have to know the column + /// exists, and so the test fakes that construct a SQL Server row keep compiling unchanged. + /// + /// + /// servers.sql_engine_edition. Carried beside the kind because + /// takes BOTH axes and asks kind first: a + /// PostgreSQL target's edition is 0, which the edition axis correctly reads as "no claim". Passing 0 + /// here for a server whose edition has not been read is exactly that same silence. + /// public DarlingServer( int serverId, string serverName, string displayName, bool isEnabled, int? sqlMajorVersion, - decimal monthlyCostUsd = 0) + decimal monthlyCostUsd = 0, + string? engineKind = null, + int engineEdition = CollectorEngineCapability.UnknownEngineEdition) { ServerId = serverId; ServerName = serverName; @@ -45,6 +61,8 @@ public DarlingServer( IsEnabled = isEnabled; SqlMajorVersion = sqlMajorVersion; MonthlyCostUsd = monthlyCostUsd; + EngineKind = engineKind; + EngineEdition = engineEdition; } public int ServerId { get; } @@ -54,6 +72,35 @@ public DarlingServer( public int? SqlMajorVersion { get; } public decimal MonthlyCostUsd { get; } + /// The raw servers.engine_kind token, or null when the store makes no claim. Kept raw + /// rather than decoded to an enum so an unrecognised token written by a NEWER service survives the trip + /// to , which shows the operator the literal string to search for. + public string? EngineKind { get; } + + /// servers.sql_engine_edition; + /// when the probe has not run or the target has no such property (every PostgreSQL target). + public int EngineEdition { get; } + + /// + /// True only when the store SAYS this target is PostgreSQL. Absence is not evidence for either engine, + /// so a null or unrecognised token is false and the server keeps the SQL Server surface - the + /// pre-#2530 behaviour, which is the only safe default for the servers that have not reconnected since + /// the engine-kind rung landed. + /// + public bool IsPostgres => MonitoredEngineKind.IsPostgres(EngineKind); + + /// + /// How the engine reads in the per-server header, or null when the store makes no claim - in which case + /// the header shows NO engine label rather than "SQL Server", because the tabs such a server gets are a + /// default rather than a finding. An unrecognised token renders as the raw token: the describer's + /// "an unrecognised engine" is worded to sit mid-sentence in the capability messages and reads wrong as + /// a label, and the literal string is what an operator would grep their store for. + /// + public string? EngineDescription => + !MonitoredEngineKind.IsKnown(EngineKind) + ? (string.IsNullOrWhiteSpace(EngineKind) ? null : EngineKind.Trim()) + : MonitoredEngineKind.DescribeEngineKind(EngineKind); + /// "SQL Server 2022"-style label for the server list; empty when the version is unknown. public string VersionLabel => ViewerDataService.SqlVersionLabel(SqlMajorVersion); @@ -87,7 +134,7 @@ public bool? IsOnline { _isOnline = value; OnPropertyChanged(nameof(IsOnline)); - OnPropertyChanged(nameof(DotStatus)); + RaiseDotChanged(); } } } @@ -104,34 +151,103 @@ public bool HasCollectorErrors { _hasCollectorErrors = value; OnPropertyChanged(nameof(HasCollectorErrors)); - OnPropertyChanged(nameof(DotStatus)); + RaiseDotChanged(); } } } + private bool _awaitingFirstCollection; + /// - /// Sidebar status-dot vocabulary — "Online"/"Warning"/"Offline"/"Unknown", identical to Lite's - /// ServerConnection.DotStatus so the ported server-row DataTriggers colour the Ellipse the same way. + /// True when no collection has EVER landed for this server: the service hasn't reached it yet — a + /// registered-but-queued server during bootstrap, not a dead one. The row had no such flag until #2473, + /// which is precisely why its dot went grey "Unknown" while the Overview card one panel over said amber + /// "Awaiting first collection" about the same server, off the same freshness call. /// - public string DotStatus => IsOnline switch + public bool AwaitingFirstCollection { - true => HasCollectorErrors ? "Warning" : "Online", - false => "Offline", - _ => "Unknown" - }; + get => _awaitingFirstCollection; + set + { + if (_awaitingFirstCollection != value) + { + _awaitingFirstCollection = value; + OnPropertyChanged(nameof(AwaitingFirstCollection)); + RaiseDotChanged(); + } + } + } + + /// + /// The sidebar row's status, as a VALUE — the same the Overview card + /// renders (#2473). This row used to derive its own copy of the status ladder from the same flags, on a + /// different type, on a different surface, and with only four of the five states: the sidebar and the card + /// could therefore say different things about one server, and for a never-collected server they did. Both + /// now read , which is the collapse #2429 argued for — + /// with one discriminant there is no flag combination left for the renderings to disagree about. + /// + public ServerCollectionStatus CardStatus => + ServerCollectionStatusRules.Classify(IsOnline, HasCollectorErrors, AwaitingFirstCollection); + + /// + /// Sidebar status-dot vocabulary — the SAME words the Overview card's StatusDisplay shows, because + /// both render . They are also the DataTrigger values MainWindow.xaml keys + /// the Ellipse fill off, so a state with no trigger falls through to the muted grey default rather than + /// failing anything — which is how "Awaiting first collection" was silently grey. The trigger set is + /// pinned against the enum in ViewerSidebarDotRendersTheCardStatusTests. + /// + public string DotStatus => CardStatus.Word(); + + /// + /// What the dot means, for the ToolTip the sidebar Ellipse carries — the answer to #2422 one surface over, + /// and the thing that makes an amber dot legible at all. The card's amber covers two states (a stale + /// collection and a never-collected server) and the card disambiguates them with a WORD; a dot has no room + /// for one, so the tooltip does that job instead. The first line is + /// , word-for-word what the Overview card says for the + /// same state. + /// + /// The second line names the axis. In Darling every one of these words is a COLLECTION answer — + /// there is no live ping to a monitored server — which is the opposite of Lite, where the identically + /// coloured dot reports a connection check. A reader who uses both should not have to infer which. The + /// third names the gesture THIS surface supports: ServerList_MouseDoubleClick opens the tab, and a + /// single click only selects, so naming one would be naming a no-op. + /// + public string DotTooltip => string.Join( + "\n", + CardStatus.Headline(), + "Darling has no live ping: this is how old the newest collection is, not a connection check.", + "Double-click the row to open this server's tab"); + + /// Raises the change notifications for everything derived from the three status flags. One + /// helper rather than three call sites per setter: was added after + /// , and a derived member that a setter forgets to announce is a dot that stops + /// updating in place — silent, and only visible on a list refresh. + private void RaiseDotChanged() + { + OnPropertyChanged(nameof(CardStatus)); + OnPropertyChanged(nameof(DotStatus)); + OnPropertyChanged(nameof(DotTooltip)); + } /// /// Sets the dot from the same collection-freshness classification the Overview cards use - /// (): Fresh → Online, Stale → the amber Warning, - /// Offline → red, NeverCollected → the grey Unknown dot (the service hasn't reached the server yet — - /// during a fleet bootstrap that is "queued", not "dead"). Both instants are UTC (the store is naive - /// UTC; nowUtc is ). + /// (), through the same + /// mapping: Fresh → Online, Stale → the amber Warning, + /// Offline → red, NeverCollected → the amber "Awaiting first collection" (the service hasn't reached the + /// server yet — during a fleet bootstrap that is "queued", not "dead"). Both instants are UTC (the store + /// is naive UTC; nowUtc is ). + /// + /// This method used to set two flags out of three by hand and drop the awaiting marker on the floor. + /// Nothing about a block of assignments makes a missing one visible, which is why the flags now arrive as + /// a single value (#2473). /// public void ApplyFreshness(DateTime? lastCollectionUtc, DateTime nowUtc) { - var freshness = ServerSummaryItem.ClassifyFreshness(lastCollectionUtc, nowUtc); - IsOnline = freshness == ServerFreshness.NeverCollected ? null : freshness != ServerFreshness.Offline; - HasCollectorErrors = freshness == ServerFreshness.Stale; + var flags = ServerCollectionStatusRules.FlagsFor( + ServerSummaryItem.ClassifyFreshness(lastCollectionUtc, nowUtc)); + IsOnline = flags.IsOnline; + HasCollectorErrors = flags.HasCollectorErrors; + AwaitingFirstCollection = flags.AwaitingFirstCollection; } // ── Per-server alert "needs attention" badge state (from the polled alert history, ack-aware) ── @@ -233,8 +349,16 @@ public void SetSilenced(bool silenced) /// public sealed partial class ViewerDataService : IAsyncDisposable { + /// + /// The observed server registry. engine_kind (#2530) and sql_engine_edition ride along + /// because the per-server tab set and every PostgreSQL panel's own explanation are derived from them - + /// see . Both are read here AND in + /// ViewerDataService.MonitoredServers.cs's ManagedServersSql: that one is what the sidebar + /// actually uses on a seeded store, so a discriminator added to only this query would have left every + /// real deployment on the SQL Server tab set. + /// public const string ServersSql = - "SELECT server_id, server_name, display_name, is_enabled, sql_major_version, COALESCE(monthly_cost_usd, 0) FROM servers ORDER BY display_name"; + "SELECT server_id, server_name, display_name, is_enabled, sql_major_version, COALESCE(monthly_cost_usd, 0), engine_kind, COALESCE(sql_engine_edition, 0) FROM servers ORDER BY display_name"; /// /// The authoritative read-only probe (V8 security hardening): does the connected role hold INSERT @@ -528,7 +652,35 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb') EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'query_store_health'), EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'query_store_text' AND column_name = 'query_hash'), EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'config_service' AND column_name = 'compose_statement_timeout_seconds'), - EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'config_alert_settings' AND column_name = 'file_growth_enabled')"; + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'config_alert_settings' AND column_name = 'file_growth_enabled'), + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'collection_log' AND column_name = 'slowest_item_ms'), + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'tempdb_stats' AND column_name = 'max_size_mb'), + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'servers' AND column_name = 'engine_kind'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_database_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_index_usage_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_table_bloat_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_session_states'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_plan_capture_readiness'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_write_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_extension_availability'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_lock_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_column_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_replication_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_buffer_usage'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_index_bloat'), + /* V95 probes a COLUMN, not a table. The three tables it touches all already exist at V94, so table + existence cannot separate the rungs and information_schema.columns is the only sentinel that can. */ + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'pg_column_stats' + AND column_name = 'database_name'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_wait_sampling'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_kernel_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_predicate_stats'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_plan_capture'), + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'servers' AND column_name = 'postgres_major_version'), + EXISTS (SELECT 1 FROM information_schema.columns WHERE table_name = 'pg_io_stats' AND column_name = 'read_bytes'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_server_config'), + EXISTS (SELECT 1 FROM information_schema.tables WHERE table_name = 'pg_deadlocks'), + EXISTS (SELECT 1 FROM pg_indexes WHERE indexname = 'idx_pg_deadlocks_identity')"; /// The store schema version this viewer build requires — the highest migration it knows /// (). The connect-time gate blocks a store below this. @@ -549,7 +701,7 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb') await using var reader = await command.ExecuteReaderAsync(cancellationToken); if (await reader.ReadAsync(cancellationToken)) { - return MapProbedSchemaVersion(reader.GetBoolean(0), reader.GetBoolean(1), reader.GetBoolean(2), reader.GetBoolean(3), reader.GetBoolean(4), reader.GetBoolean(5), reader.GetBoolean(6), reader.GetBoolean(7), reader.GetBoolean(8), reader.GetBoolean(9), reader.GetBoolean(10), reader.GetBoolean(11), reader.GetBoolean(12), reader.GetBoolean(13), reader.GetBoolean(14), reader.GetBoolean(15), reader.GetBoolean(16), reader.GetBoolean(17), reader.GetBoolean(18), reader.GetBoolean(19), reader.GetBoolean(20), reader.GetBoolean(21), reader.GetBoolean(22), reader.GetBoolean(23), reader.GetBoolean(24), reader.GetBoolean(25), reader.GetBoolean(26), reader.GetBoolean(27), reader.GetBoolean(28), reader.GetBoolean(29), reader.GetBoolean(30), reader.GetBoolean(31), reader.GetBoolean(32), reader.GetBoolean(33), reader.GetBoolean(34), reader.GetBoolean(35), reader.GetBoolean(36), reader.GetBoolean(37), reader.GetBoolean(38), reader.GetBoolean(39), reader.GetBoolean(40), reader.GetBoolean(41), reader.GetBoolean(42), reader.GetBoolean(43), reader.GetBoolean(44), reader.GetBoolean(45), reader.GetBoolean(46), reader.GetBoolean(47), reader.GetBoolean(48), reader.GetBoolean(49), reader.GetBoolean(50), reader.GetBoolean(51), reader.GetBoolean(52), reader.GetBoolean(53), reader.GetBoolean(54)); + return MapProbedSchemaVersion(reader.GetBoolean(0), reader.GetBoolean(1), reader.GetBoolean(2), reader.GetBoolean(3), reader.GetBoolean(4), reader.GetBoolean(5), reader.GetBoolean(6), reader.GetBoolean(7), reader.GetBoolean(8), reader.GetBoolean(9), reader.GetBoolean(10), reader.GetBoolean(11), reader.GetBoolean(12), reader.GetBoolean(13), reader.GetBoolean(14), reader.GetBoolean(15), reader.GetBoolean(16), reader.GetBoolean(17), reader.GetBoolean(18), reader.GetBoolean(19), reader.GetBoolean(20), reader.GetBoolean(21), reader.GetBoolean(22), reader.GetBoolean(23), reader.GetBoolean(24), reader.GetBoolean(25), reader.GetBoolean(26), reader.GetBoolean(27), reader.GetBoolean(28), reader.GetBoolean(29), reader.GetBoolean(30), reader.GetBoolean(31), reader.GetBoolean(32), reader.GetBoolean(33), reader.GetBoolean(34), reader.GetBoolean(35), reader.GetBoolean(36), reader.GetBoolean(37), reader.GetBoolean(38), reader.GetBoolean(39), reader.GetBoolean(40), reader.GetBoolean(41), reader.GetBoolean(42), reader.GetBoolean(43), reader.GetBoolean(44), reader.GetBoolean(45), reader.GetBoolean(46), reader.GetBoolean(47), reader.GetBoolean(48), reader.GetBoolean(49), reader.GetBoolean(50), reader.GetBoolean(51), reader.GetBoolean(52), reader.GetBoolean(53), reader.GetBoolean(54), reader.GetBoolean(55), reader.GetBoolean(56), reader.GetBoolean(57), reader.GetBoolean(58), reader.GetBoolean(59), reader.GetBoolean(60), reader.GetBoolean(61), reader.GetBoolean(62), reader.GetBoolean(63), reader.GetBoolean(64), reader.GetBoolean(65), reader.GetBoolean(66), reader.GetBoolean(67), reader.GetBoolean(68), reader.GetBoolean(69), reader.GetBoolean(70), reader.GetBoolean(71), reader.GetBoolean(72), reader.GetBoolean(73), reader.GetBoolean(74), reader.GetBoolean(75), reader.GetBoolean(76), reader.GetBoolean(77), reader.GetBoolean(78), reader.GetBoolean(79)); } return null; @@ -574,7 +726,7 @@ OR NOT EXISTS (SELECT 1 FROM pg_extension WHERE extname = 'timescaledb') /// is unit-tested without a live store; any schema bump past the newest arm trips the pinning test that keeps /// this in step with . /// - internal static int MapProbedSchemaVersion(bool hasConfigControlPlane, bool hasAlertDeliveryOverride, bool hasAnalysisState, bool hasAlertTuningKnobs, bool hasDefaultTraceEvents, bool hasIndexObjectStatsLatestIndex, bool hasCollectionLogHypertableOrPlainPg, bool hasJobHistory, bool hasAgentStatus, bool hasGenericWebhook, bool hasDeadlocksDatabaseName, bool hasQueryStoreReplicaRole, bool hasLongQueryCompletions, bool hasWebDashboardConfig, bool hasCustomViews, bool hasServerTags, bool hasConnectionRefireKnobs = false, bool hasAgCollectors = false, bool hasAgAlertKnobs = false, bool hasAgLatencyColumns = false, bool hasAgDisconnectRefire = false, bool hasPayloadDimensions = false, bool hasDimFloorIndexes = false, bool hasBlockingWaitThreshold = false, bool hasQueryStoreIntervalIdentity = false, bool hasPagerDutyWebhook = false, bool hasPagerDutyProxy = false, bool hasCollectorState = false, bool hasPlanCorrection = false, bool hasPvsStats = false, bool hasPvsPressureKnobs = false, bool hasDatabaseStateAlert = false, bool hasServerTagColour = false, bool hasQueryStatsHostObject = false, bool hasFindingDrillDown = false, bool hasStoreMetrics = false, bool hasPlanDimGzip = false, bool hasSelfAlertKnobs = false, bool hasJobMetricsColumns = false, bool hasJobCadenceKnob = false, bool hasBackfillSwitch = false, bool hasCollectorMemoryKnobs = false, bool hasDatabaseStateEdgeMemory = false, bool hasIncidentOccurrences = false, bool hasPlanXmlCompressionKnob = false, bool hasMonitoredServerEngine = false, bool hasPgBlockingEdges = false, bool hasQueryStorePlanMap = false, bool hasPgStatementText = false, bool hasQueryStoreText = false, bool hasPlanContentRetentionKnob = false, bool hasQueryStoreHealth = false, bool hasQueryStoreTextHash = false, bool hasComposeTimeoutKnob = false, bool hasFileGrowthAlert = false) + internal static int MapProbedSchemaVersion(bool hasConfigControlPlane, bool hasAlertDeliveryOverride, bool hasAnalysisState, bool hasAlertTuningKnobs, bool hasDefaultTraceEvents, bool hasIndexObjectStatsLatestIndex, bool hasCollectionLogHypertableOrPlainPg, bool hasJobHistory, bool hasAgentStatus, bool hasGenericWebhook, bool hasDeadlocksDatabaseName, bool hasQueryStoreReplicaRole, bool hasLongQueryCompletions, bool hasWebDashboardConfig, bool hasCustomViews, bool hasServerTags, bool hasConnectionRefireKnobs = false, bool hasAgCollectors = false, bool hasAgAlertKnobs = false, bool hasAgLatencyColumns = false, bool hasAgDisconnectRefire = false, bool hasPayloadDimensions = false, bool hasDimFloorIndexes = false, bool hasBlockingWaitThreshold = false, bool hasQueryStoreIntervalIdentity = false, bool hasPagerDutyWebhook = false, bool hasPagerDutyProxy = false, bool hasCollectorState = false, bool hasPlanCorrection = false, bool hasPvsStats = false, bool hasPvsPressureKnobs = false, bool hasDatabaseStateAlert = false, bool hasServerTagColour = false, bool hasQueryStatsHostObject = false, bool hasFindingDrillDown = false, bool hasStoreMetrics = false, bool hasPlanDimGzip = false, bool hasSelfAlertKnobs = false, bool hasJobMetricsColumns = false, bool hasJobCadenceKnob = false, bool hasBackfillSwitch = false, bool hasCollectorMemoryKnobs = false, bool hasDatabaseStateEdgeMemory = false, bool hasIncidentOccurrences = false, bool hasPlanXmlCompressionKnob = false, bool hasMonitoredServerEngine = false, bool hasPgBlockingEdges = false, bool hasQueryStorePlanMap = false, bool hasPgStatementText = false, bool hasQueryStoreText = false, bool hasPlanContentRetentionKnob = false, bool hasQueryStoreHealth = false, bool hasQueryStoreTextHash = false, bool hasComposeTimeoutKnob = false, bool hasFileGrowthAlert = false, bool hasCollectionLogFanoutRollup = false, bool hasTempDbMaxSize = false, bool hasServerEngineKind = false, bool hasPgDatabaseStats = false, bool hasPgIndexUsageStats = false, bool hasPgTableBloatStats = false, bool hasPgSessionStates = false, bool hasPgPlanCaptureReadiness = false, bool hasPgWriteStats = false, bool hasPgExtensionAvailability = false, bool hasPgLockStats = false, bool hasPgColumnStats = false, bool hasPgReplicationStats = false, bool hasPgBufferUsage = false, bool hasPgIndexBloat = false, bool hasPgPerDatabaseAttribution = false, bool hasPgWaitSampling = false, bool hasPgKernelStats = false, bool hasPgPredicateStats = false, bool hasPgPlanCapture = false, bool hasPgMajorVersion = false, bool hasPg18IoBytes = false, bool hasPgServerConfig = false, bool hasPgDeadlocks = false, bool hasPgDeadlockIdentity = false) { /* V71 (the PostgreSQL blocking-edges rung): a table-existence sentinel and now the newest-first arm. A collector table would ordinarily get no arm at all — see the V63-V69 note below — but the TOP @@ -628,6 +780,234 @@ prose mention would exempt it. */ /* #2349: newest first. The arm below STAYS — a store migrated to exactly 78 must map to 78 rather than falling through. The column is named only in the probe line, not in this prose, per the V71 finding: the coverage ratchet strips information_schema lines but cannot strip a comment. */ + /* #2472: newest first. The arm below STAYS — a store migrated to exactly 79 must map to 79 rather + than falling through. The column is named only in the probe line, not in this prose, per the V71 + finding: the coverage ratchet strips information_schema lines but cannot strip a comment. + + This rung's gate earns its place rather than being bookkeeping. The service writes the fan-out + rollup on every productive per-database run and Collection Health reads it; a viewer on a V79 + store would find the column missing and the read would throw rather than degrade, so the banner + has to fire before the tab does. */ + /* #2515: newest first. The arm below STAYS — a store migrated to exactly 80 must map to 80 rather + than falling through. The column is named only in the probe line, not in this prose, per the V71 + finding: the coverage ratchet strips information_schema lines but cannot strip a comment. + + The reason to gate is the standing invariant rather than a read that would throw — no viewer query + names this column; the SERVICE's alert adapter is what reads it. RequiredStoreSchemaVersion is + StorageVersion.SchemaVersion, so a fully-migrated store has to map to EXACTLY this or the banner + reports a mismatch on a store that is current. What the banner buys on the way is worth having: a + store still at 80 has no ceiling recorded anywhere in its history, so every tempdb percentage in it + is the old distance-to-next-autogrow measurement, and the operator should know that before reading + one. */ + /* #2540: newest first, and the arms below STAY — a store migrated to exactly 85 must map to 85 + rather than falling through. The table is named only in the probe line, not in this prose, per + the V71 finding: the coverage ratchet strips information_schema lines but cannot strip a comment. + + A collector-table rung would ordinarily get no arm at all — see the V63-V69 note below — and this + one is here for the standing invariant: it is the TOP rung, RequiredStoreSchemaVersion is + StorageVersion.SchemaVersion, and a fully-migrated store must map to EXACTLY that or the version + banner reports a mismatch against a store that is current. Nothing in the viewer reads the table; + the read that does is on the MCP surface. + + The three comment blocks that used to stack here were moved down onto the arms they describe + (#2530 → V82, #2539 → V83, #2542 → V85). They had drifted upward as each new rung inserted its + own block above them, which is how a comment ends up explaining an arm two rungs away from the + one it was written for. Keep the block with its arm. */ + /* V104 (#2661): the deadlock lookup index. An INDEX sentinel rather than a table one, because + V103 created the table and V104 only indexes it — a store stopped between the two has the table + and not the index, which is a real interrupted-upgrade state and exactly what these arms exist to + distinguish. The TOP rung, so it must map exactly or the connect-time gate refuses a store that + is perfectly current. */ + if (hasPgDeadlockIdentity) + { + return 104; + } + + /* V103 (#2661): collect.pg_deadlocks itself. */ + if (hasPgDeadlocks) + { + return 103; + } + + /* V102 (#2658): the pg_settings snapshot. Table-existence sentinel. Was the top rung until V103 + landed and keeps its own arm, because a store stopped between them is a real state during an + interrupted upgrade. */ + if (hasPgServerConfig) + { + return 102; + } + + /* V101 (#2655): PostgreSQL 18's measured I/O byte columns. Column sentinel. Was the top rung until + V102 landed and keeps its own arm, because a store stopped between the two is a real state + during an interrupted upgrade. */ + if (hasPg18IoBytes) + { + return 101; + } + + /* V100 (#2653): the PostgreSQL major on the registry. A COLUMN sentinel, like V82's next door. Was + the top rung until V101 landed and keeps its own arm, because a store stopped between the two is + a real state during an interrupted upgrade. */ + if (hasPgMajorVersion) + { + return 100; + } + + /* V99 (#2566): pg_plan_capture. Was the top rung until V100 landed and keeps its own arm, because a + store stopped between the two is a real state during an interrupted upgrade. */ + if (hasPgPlanCapture) + { + return 99; + } + + /* V98 (#2603): pg_predicate_stats. Was the top rung until V99 landed and keeps its own arm, because + a store stopped between the two is a real state during an interrupted upgrade. */ + if (hasPgPredicateStats) + { + return 98; + } + + /* V97 (#2603): pg_kernel_stats. Was the top rung until V98 landed and keeps its own arm, because a + store stopped between the two is a real state during an interrupted upgrade. */ + if (hasPgKernelStats) + { + return 97; + } + + /* V96 (#2603): pg_wait_sampling. Was the top rung until V97 landed and keeps its own arm, because a + store stopped between the two is a real state during an interrupted upgrade. */ + if (hasPgWaitSampling) + { + return 96; + } + + /* V95 (#2599): database_name on the three per-database PostgreSQL tables. Was the top rung until + V96 landed and keeps its own arm, because a store stopped between the two is a real state during + an interrupted upgrade. */ + if (hasPgPerDatabaseAttribution) + { + return 95; + } + + /* V94 (#2561): b-tree index bloat, measured. Was the top rung until V95 landed and keeps its own + arm, because a store stopped between the two is a real state during an interrupted upgrade. */ + if (hasPgIndexBloat) + { + return 94; + } + + /* V93 (#2544, buffers): what is resident in shared buffers. Was the top rung until V94 landed and + keeps its own arm, because a store stopped between the two is a real state during an interrupted + migration. Formerly the TOP rung — RequiredStoreSchemaVersion is StorageVersion.SchemaVersion and a + fully-migrated store must map to exactly that, or the version banner reports a mismatch on a store + that is current. */ + if (hasPgBufferUsage) + { + return 93; + } + + /* V92 (#2544, replication): connected standbys and their distance from the primary. Was the top rung + until V93 landed and keeps its own arm, because a store stopped between the two is a real state + during an interrupted migration. */ + if (hasPgReplicationStats) + { + return 92; + } + + /* V91 (#2543): per-column planner statistics. Was the top rung until V92 landed and keeps its own + arm, because a store stopped between the two is a real state during an interrupted migration. */ + if (hasPgColumnStats) + { + return 91; + } + + /* V90 (#2544, locks): lock state by mode, type and relation. Was the top rung until V91 landed and + keeps its own arm, because a store stopped between the two is a real state during an interrupted + migration. */ + if (hasPgLockStats) + { + return 90; + } + + /* V89 (#2545): which extensions this target has, could have, or cannot have. Was the top rung until + V90 landed and keeps its own arm, because a store stopped between the two is a real state during + an interrupted migration. */ + if (hasPgExtensionAvailability) + { + return 89; + } + + /* V88 (#2544): the write side — checkpoints, background writer, WAL. Was the top rung until V89 + landed and keeps its own arm, because a store stopped between the two is a real state during an + interrupted migration. */ + if (hasPgWriteStats) + { + return 88; + } + + /* V87 (#2564): whether a PostgreSQL target could capture plans at all. Was the top rung until V88 + landed and is kept as its own arm rather than folded away, because a store stopped between the two + is a real state during an interrupted migration. */ + if (hasPgPlanCaptureReadiness) + { + return 87; + } + + if (hasPgSessionStates) + { + return 86; + } + + /* #2542: below the top arm now, and still not redundant — a store stopped between the V84 and V85 + rungs is a real state during an interrupted migration, and without its own arm it would report 84 + while carrying V85's table. */ + if (hasPgTableBloatStats) + { + return 85; + } + + /* #2541: the index-usage rung. Below V85, above V83, for the reason every arm here is ordered + newest-first — a V86 store also has V84's table, so testing V84 first would report every current + store as two rungs behind and raise an upgrade banner against a store that is fine. */ + if (hasPgIndexUsageStats) + { + return 84; + } + + /* #2539: the arm below STAYS — a store migrated to exactly 82 must map to 82 rather than falling + through. The table is named only in the probe line, not in this prose, per the V71 finding. + + A collector-table rung would ordinarily get no arm at all; this one was the TOP rung when it + landed, which is why it has one. Nothing in the viewer reads the table; the read that does is on + the MCP surface. */ + if (hasPgDatabaseStats) + { + return 83; + } + + /* #2530: the arm below STAYS — a store migrated to exactly 81 must map to 81 rather than falling + through. The column is named only in the probe line, not in this prose, per the V71 finding. + + This rung's gate is worth having for what it tells the OPERATOR, not only for the standing + RequiredStoreSchemaVersion invariant: a store below it records nothing that says a target is + PostgreSQL, so every PostgreSQL server in it is indistinguishable from a SQL Server that has + never connected — and every surface that branches on engine kind is therefore reading a fact + that store cannot supply. */ + if (hasServerEngineKind) + { + return 82; + } + + if (hasTempDbMaxSize) + { + return 81; + } + + if (hasCollectionLogFanoutRollup) + { + return 80; + } + if (hasFileGrowthAlert) { return 79; @@ -1102,7 +1482,9 @@ public async Task> GetServersAsync(CancellationToken cancell reader.IsDBNull(2) ? serverName : reader.GetString(2), !reader.IsDBNull(3) && reader.GetBoolean(3), reader.IsDBNull(4) ? null : reader.GetInt32(4), - reader.IsDBNull(5) ? 0m : Convert.ToDecimal(reader.GetValue(5)))); + reader.IsDBNull(5) ? 0m : Convert.ToDecimal(reader.GetValue(5)), + reader.IsDBNull(6) ? null : reader.GetString(6), + reader.IsDBNull(7) ? CollectorEngineCapability.UnknownEngineEdition : reader.GetInt32(7))); } return servers; diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresDisplay.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresDisplay.cs new file mode 100644 index 000000000..384938357 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresDisplay.cs @@ -0,0 +1,1136 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Linq; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// Stored PostgreSQL rows, projected into what a grid may show (#2530). Pure and WPF-free, so every rule +/// below is unit-testable without a window — which matters, because the rules ARE the content: each one +/// exists to stop the grid printing a number that would be read as a measurement when it is not one. +/// +/// +/// -1 is not a value. It is the store's not-applicable sentinel for a duration, a WAL +/// size or a threshold. 0 was rejected for these deliberately — it reads as "started this instant", +/// "retains nothing", "already past its threshold" — so rendering -1 raw would be the same mistake +/// one step later. +/// Absent is not zero. pg_stat_io's write side is NULL across the board on Aurora, +/// because backends there do not write data files. A blank cell says that; a 0 claims a measurement +/// that was never taken. +/// NULL recurrence is "cannot tell", not "once". The blocking read returns NULL when the +/// root's own backend id did not resolve, and those are different claims. +/// Every timestamp is naive UTC in the store and goes through +/// , exactly like every other timestamp the viewer renders. +/// Blocks are not bytes. temp_blks_written counts blocks, and the block size is a +/// compile-time setting of the server, not a constant. It is labelled in blocks rather than multiplied by +/// an assumed 8kB. +/// +/// +internal static class PgDisplay +{ + /// The store's not-applicable sentinel. Named rather than repeated, because reading it as a + /// number anywhere is the defect this class exists to prevent. + internal const long NotApplicable = -1; + + /// What an unmeasured / not-applicable cell shows. An em dash, not "0" and not blank: blank + /// reads as "nothing here yet" and 0 reads as a measurement. + internal const string NotApplicableText = "—"; + + internal static string Bytes(long value) + { + if (value <= NotApplicable) + { + return NotApplicableText; + } + + string[] units = { "B", "KB", "MB", "GB", "TB", "PB" }; + double scaled = value; + var unit = 0; + while (scaled >= 1024 && unit < units.Length - 1) + { + scaled /= 1024; + unit++; + } + + /* The unit has to be chosen against the value AS RENDERED, not against the exact quotient. One byte + short of a megabyte the quotient is 1023.999…, which stays in KB and then rounds to "1024.0 KB" — + a number expressed in its own next unit, which reads as a typo rather than as a size. Re-check + once after rounding; one pass is enough, because the second division can only land on 1.0. */ + if (unit < units.Length - 1 && Math.Round(scaled, 1) >= 1024) + { + scaled /= 1024; + unit++; + } + + return unit == 0 + ? $"{value:N0} B" + : string.Create(CultureInfo.CurrentCulture, $"{scaled:N1} {units[unit]}"); + } + + /// A signed byte delta, for the "is this growing" columns. Zero is stated rather than blanked: + /// "flat" is an answer, and an empty cell would read as "not measured". + internal static string ByteDelta(long from, long to) + { + if (from <= NotApplicable || to <= NotApplicable) + { + return NotApplicableText; + } + + var delta = to - from; + return delta == 0 ? "flat" : (delta > 0 ? "+" : "-") + Bytes(Math.Abs(delta)); + } + + /// A block count, LABELLED as blocks. PostgreSQL's block size is a compile-time setting of + /// the server, so converting to bytes here would put a fabricated figure next to measured ones. + internal static string Blocks(long value) => + value <= NotApplicable + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{value:N0} blocks"); + + internal static string Count(long value) => + value <= NotApplicable ? NotApplicableText : value.ToString("N0", CultureInfo.CurrentCulture); + + internal static string CountDelta(long from, long to) + { + var delta = to - from; + return delta == 0 ? "flat" : (delta > 0 ? "+" : "") + delta.ToString("N0", CultureInfo.CurrentCulture); + } + + /// A duration in milliseconds, or the not-applicable dash for the -1 sentinel. + internal static string Milliseconds(long ms) + { + if (ms <= NotApplicable) + { + return NotApplicableText; + } + + return ms < 1000 + ? $"{ms:N0} ms" + : TimeSpan.FromMilliseconds(ms).ToString(ms < 3_600_000 ? @"mm\:ss" : @"d\.hh\:mm\:ss", CultureInfo.CurrentCulture); + } + + internal static string Timestamp(DateTime? utc) => + utc is { } value + ? ViewerTimeHelper.ForDisplay(value).ToString("yyyy-MM-dd HH:mm", CultureInfo.CurrentCulture) + : string.Empty; + + /// A percentage of a total, or the dash when the total is zero — which is "nothing happened", + /// not "0%". + internal static string Percent(long part, long total) => + total <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{part * 100.0 / total:N1}%"); + + // ── Vacuum ─────────────────────────────────────────────────────────────────────────────────── + + internal sealed class XminRow + { + public string Source { get; init; } = ""; + public string IsWinnerText { get; init; } = ""; + public string XminAge { get; init; } = ""; + public string PeakXminAge { get; init; } = ""; + public string WinnerShare { get; init; } = ""; + public string MeasuredAt { get; init; } = ""; + public string Holder { get; init; } = ""; + public string Detail { get; init; } = ""; + } + + internal static XminRow Xmin(DarlingPgXminReader.PgXminRow row) => new() + { + Source = row.Source, + IsWinnerText = row.IsWinner ? "yes" : "", + XminAge = Count(row.XminAge), + PeakXminAge = Count(row.PeakXminAge), + /* The share, not the raw count. A slot that won 58 of 60 samples is a standing problem someone + needs to own; a session that won twice is a query that ran long and finished. Reporting only the + current holder makes those two look the same. */ + WinnerShare = row.Samples <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{row.SamplesAsWinner:N0} of {row.Samples:N0} samples"), + MeasuredAt = Timestamp(row.MeasuredAt), + Holder = row.Holder ?? "", + Detail = row.Detail ?? "", + }; + + internal sealed class AutovacuumRow + { + public string DatabaseName { get; init; } = ""; + public string SchemaName { get; init; } = ""; + public string TableName { get; init; } = ""; + public string AutovacuumState { get; init; } = ""; + public bool AutovacuumDisabled { get; init; } + public string DeadTuples { get; init; } = ""; + public string VacuumThreshold { get; init; } = ""; + public string ThresholdRatio { get; init; } = ""; + public string DeadTupleTrend { get; init; } = ""; + public string InsertsSinceVacuum { get; init; } = ""; + public string InsertVacuumThreshold { get; init; } = ""; + public string LastVacuum { get; init; } = ""; + public string LastAutovacuum { get; init; } = ""; + public string LastAnalyze { get; init; } = ""; + public string LastAutoanalyze { get; init; } = ""; + public string AutovacuumCount { get; init; } = ""; + public string LiveTuples { get; init; } = ""; + public string DeadTuplePct { get; init; } = ""; + public string ModsSinceAnalyze { get; init; } = ""; + public string AnalyzeThreshold { get; init; } = ""; + public string TotalSize { get; init; } = ""; + public string MeasuredAt { get; init; } = ""; + } + + internal static AutovacuumRow Autovacuum(DarlingPgAutovacuumReader.PgAutovacuumRow row) => new() + { + DatabaseName = row.DatabaseName ?? "", + SchemaName = row.SchemaName ?? "", + TableName = row.TableName ?? "", + AutovacuumState = row.AutovacuumDisabled ? "DISABLED" : "enabled", + AutovacuumDisabled = row.AutovacuumDisabled, + DeadTuples = Count(row.DeadTuples), + VacuumThreshold = Count(row.VacuumThreshold), + /* Ratio to the table's OWN threshold, which is what makes a small hot table and a huge cold one + comparable at all. A never-analyzed table can have threshold 0, where a ratio is undefined + rather than infinite. */ + ThresholdRatio = row.VacuumThreshold <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{row.DeadTuples / (double)row.VacuumThreshold:N1}x"), + /* Climbing dead tuples mean autovacuum is losing; a flat figure well past the threshold usually + means it is blocked or switched off. Those need different fixes, so the trend is a column rather + than something to work out from two visits. */ + DeadTupleTrend = CountDelta(row.FirstDeadTuples, row.DeadTuples), + InsertsSinceVacuum = Count(row.InsertsSinceVacuum), + /* -1 here is not "no threshold": it is a major with no autovacuum_vacuum_insert_threshold at all. */ + InsertVacuumThreshold = row.InsertVacuumThreshold <= NotApplicable + ? "n/a on this version" + : Count(row.InsertVacuumThreshold), + /* Manual and automatic runs are separate columns because they answer different questions: a table + whose only vacuums are manual is one somebody is nursing, and that is a finding about the + configuration rather than about the workload. */ + LastVacuum = Timestamp(row.LastVacuum), + LastAutovacuum = Timestamp(row.LastAutovacuum), + LastAnalyze = Timestamp(row.LastAnalyze), + LastAutoanalyze = Timestamp(row.LastAutoanalyze), + /* Zero is the finding, not a missing value: a table autovacuum has NEVER processed is the classic + wraparound route, because relfrozenxid never advances. */ + AutovacuumCount = Count(row.AutovacuumCount), + LiveTuples = Count(row.LiveTuples), + /* The dead-tuple SHARE, which is what bloat actually is. The raw count beside a threshold says + whether vacuum is due; the share says how much of the table is already waste. */ + DeadTuplePct = Percent(row.DeadTuples, row.LiveTuples + row.DeadTuples), + /* The ANALYZE half of the same story, and it is genuinely separate: a table can be vacuumed on + schedule and still have statistics old enough to give the planner the wrong row counts. */ + ModsSinceAnalyze = Count(row.ModsSinceAnalyze), + AnalyzeThreshold = Count(row.AnalyzeThreshold), + TotalSize = Bytes(row.TotalBytes), + MeasuredAt = Timestamp(row.MeasuredAt), + }; + + internal sealed class WraparoundRow + { + public string DatabaseName { get; init; } = ""; + public string FrozenXidAge { get; init; } = ""; + public double PctTowardEmergencyVacuum { get; init; } + public double PctTowardWraparound { get; init; } + public string XidsRemaining { get; init; } = ""; + public string MinMultiXidAge { get; init; } = ""; + public double PctTowardMultixactEmergency { get; init; } + public double PctTowardMultixactWraparound { get; init; } + public string MultiXidsRemaining { get; init; } = ""; + public string FreezeMaxAge { get; init; } = ""; + public string MultixactFreezeMaxAge { get; init; } = ""; + public string WindowPeakFrozenXidAge { get; init; } = ""; + public string WindowPeakMinMultiXidAge { get; init; } = ""; + public string ConnectionsAllowed { get; init; } = ""; + public string MeasuredAt { get; init; } = ""; + } + + internal static WraparoundRow Wraparound(DarlingPgWraparoundReader.PgWraparoundRow row) => new() + { + DatabaseName = row.DatabaseName, + FrozenXidAge = Count(row.FrozenXidAge), + PctTowardEmergencyVacuum = Math.Round(row.PctTowardEmergencyVacuum, 1), + PctTowardWraparound = Math.Round(row.PctTowardWraparound, 1), + XidsRemaining = Count(row.XidsRemaining), + MinMultiXidAge = Count(row.MinMultiXidAge), + PctTowardMultixactEmergency = Math.Round(row.PctTowardMultixactEmergency, 1), + /* The multixact side has its OWN shutdown ceiling, reached independently of the XID one - a + workload heavy on row-level share locks or subtransactions gets there first, and nothing on the + XID columns would say so. */ + PctTowardMultixactWraparound = Math.Round(row.PctTowardMultixactWraparound, 1), + MultiXidsRemaining = Count(row.MultiXidsRemaining), + /* The percentages beside these are graded against THIS cluster's settings, not a constant, so the + settings themselves are a column: two databases at "80% to emergency" are different distances from + trouble when their freeze_max_age differs, and nothing else on the row would say so. */ + FreezeMaxAge = Count(row.AutovacuumFreezeMaxAge), + MultixactFreezeMaxAge = Count(row.AutovacuumMultixactFreezeMaxAge), + WindowPeakFrozenXidAge = Count(row.WindowPeakFrozenXidAge), + WindowPeakMinMultiXidAge = Count(row.WindowPeakMinMultiXidAge), + /* datallowconn=false is why a database can sit at 99% of wraparound and never be vacuumed by + anything that connects to it. It is the finding, not a footnote. */ + ConnectionsAllowed = row.AllowsConnections ? "allowed" : "NOT ALLOWED", + MeasuredAt = Timestamp(row.MeasuredAt), + }; + + // ── Waits, I/O, replication ────────────────────────────────────────────────────────────────── + + internal sealed class WaitRow + { + public string WaitType { get; init; } = ""; + public string WaitEvent { get; init; } = ""; + public string TotalWaits { get; init; } = ""; + public double TotalWaitTimeMs { get; init; } + public double AvgWaitTimeMs { get; init; } + } + + internal static WaitRow Wait(DarlingPgWaitReader.PgWaitRow row) => new() + { + WaitType = row.WaitType, + WaitEvent = row.WaitEvent, + TotalWaits = Count(row.TotalWaits), + TotalWaitTimeMs = Math.Round(row.TotalWaitTimeMs, 1), + AvgWaitTimeMs = Math.Round(row.AvgWaitTimeMs, 3), + }; + + internal sealed class IoRow + { + public string BackendType { get; init; } = ""; + public string ObjectType { get; init; } = ""; + public string Context { get; init; } = ""; + + /// What this context MEANS, from the shared reader — the one copy the MCP surface also + /// prints. Shown as the Context cell's tooltip rather than a column: it is a paragraph, and it is + /// the dimension with no SQL Server counterpart, so an operator meeting `bulkread` for the first + /// time needs it and one who knows it does not want it eating the grid. + public string ContextMeaning { get; init; } = ""; + + public string Reads { get; init; } = ""; + public double ReadTimeMs { get; init; } + + /// Per-read latency — the figure that separates "a lot of I/O" from "slow I/O", which have + /// completely different remedies. + public string AvgReadMs { get; init; } = ""; + + public string Hits { get; init; } = ""; + + /// A hit ratio scoped to THIS combination, the only scope where it means anything: a + /// server-wide ratio averages bulkread's deliberate misses together with normal-context ones and + /// understates both. + public string HitPct { get; init; } = ""; + + /// This combination's share of the window's total read TIME — how the grid's order was + /// decided, made legible. A row at 4% of the reads and 60% of the read time is the finding. + public string PctOfTotalReadTime { get; init; } = ""; + + public string Extends { get; init; } = ""; + public string ExtendTimeMs { get; init; } = ""; + public string Evictions { get; init; } = ""; + public string Reuses { get; init; } = ""; + public string Writes { get; init; } = ""; + public string WriteTimeMs { get; init; } = ""; + /// The per-operation block size. PostgreSQL 18 removed it, so it reads "n/a" there rather + /// than 0 — on 18 a read is no longer one block and there is no single size to report. + public string OpSize { get; init; } = ""; + + /// Bytes actually moved. MEASURED on PostgreSQL 18, which reports byte totals directly; + /// below 18 it is count x block size, which is exact there because one operation moves one block. + /// The two are not the same quantity — 18's vectored reads make the old estimate undercount by an + /// order of magnitude — so the estimate says so in the cell rather than passing for a + /// measurement. + public string ReadVolume { get; init; } = ""; + + public string WriteVolume { get; init; } = ""; + public string StatsReset { get; init; } = ""; + } + + /// + /// Projects the whole I/O result at once, because two of its columns are SHARES of the window and a + /// per-row projection cannot compute a denominator it never sees. Ordering the grid by read time and + /// then not saying what share that is leaves the reader to divide by hand. + /// + internal static List IoRows(IReadOnlyList rows) + { + ArgumentNullException.ThrowIfNull(rows); + + var totalReadTime = rows.Sum(r => r.ReadTimeMs); + return rows.Select(r => Io(r, totalReadTime)).ToList(); + } + + internal static IoRow Io(DarlingPgIoReader.PgIoRow row, double totalReadTimeMs = 0) => new() + { + BackendType = row.BackendType ?? "", + ObjectType = row.ObjectType ?? "", + Context = row.Context ?? "", + ContextMeaning = DarlingPgIoReader.ContextMeaning(row.Context), + Reads = Count(row.Reads), + ReadTimeMs = Math.Round(row.ReadTimeMs, 1), + AvgReadMs = row.Reads > 0 + ? string.Create(CultureInfo.CurrentCulture, $"{row.ReadTimeMs / row.Reads:N3}") + : NotApplicableText, + Hits = Count(row.Hits), + HitPct = Percent(row.Hits, row.Hits + row.Reads), + PctOfTotalReadTime = totalReadTimeMs > 0 + ? string.Create(CultureInfo.CurrentCulture, $"{row.ReadTimeMs / totalReadTimeMs * 100:N1}%") + : NotApplicableText, + Extends = Count(row.Extends), + ExtendTimeMs = string.Create(CultureInfo.CurrentCulture, $"{Math.Round(row.ExtendTimeMs, 1):N1}"), + Evictions = Count(row.Evictions), + Reuses = Count(row.Reuses), + /* "not tracked", not 0. On Aurora the whole pg_stat_io write side is NULL because backends there + do not write data files, and the read carries WriteCountersTracked precisely so this cell can + tell that apart from a server that wrote nothing. */ + Writes = row.WriteCountersTracked ? Count(row.Writes) : "not tracked", + WriteTimeMs = row.WriteCountersTracked + ? Math.Round(row.WriteTimeMs, 1).ToString("N1", CultureInfo.CurrentCulture) + : "not tracked", + OpSize = row.ByteCountersTracked ? "n/a" : Bytes(row.OpBytes), + /* Measured beats derived, and the estimate is labelled rather than silently equated with one. + "not measured" is the third state: a server with neither op_bytes nor the 18 columns, which is + not the same as having moved no bytes. */ + ReadVolume = row.ByteCountersTracked + ? Bytes((long)row.ReadBytes) + : (row.OpBytes > 0 ? Bytes(row.Reads * row.OpBytes) + " (est.)" : "not measured"), + WriteVolume = row.ByteCountersTracked + ? Bytes((long)row.WriteBytes) + : (row.OpBytes > 0 && row.WriteCountersTracked + ? Bytes(row.Writes * row.OpBytes) + " (est.)" + : "not measured"), + StatsReset = Timestamp(row.StatsReset), + }; + + internal sealed class SlotRow + { + public string SlotName { get; init; } = ""; + public string SlotType { get; init; } = ""; + public string ActiveText { get; init; } = ""; + public string WalStatus { get; init; } = ""; + public string RetainedWal { get; init; } = ""; + public string RetainedWalTrend { get; init; } = ""; + public string SafeWalSize { get; init; } = ""; + public string XminAge { get; init; } = ""; + public string CatalogXminAge { get; init; } = ""; + public string InactiveSince { get; init; } = ""; + public string DatabaseName { get; init; } = ""; + public string Plugin { get; init; } = ""; + public string InvalidationReason { get; init; } = ""; + + /// The stored conflicting flag: a logical slot whose needed rows were vacuumed away + /// by a recovery conflict. It is a DIFFERENT failure from invalidation-by-WAL-size and needs a + /// different response (the subscriber must be re-created either way, but the cause is + /// hot_standby_feedback rather than max_slot_wal_keep_size), so it is its own column + /// rather than folded into the invalidation reason. + public string Conflicting { get; init; } = ""; + + public string MeasuredAt { get; init; } = ""; + public bool IsInvalidated { get; init; } + public bool IsInactive { get; init; } + } + + internal static SlotRow Slot(DarlingPgSlotReader.PgSlotRow row) => new() + { + SlotName = row.SlotName, + SlotType = row.SlotType ?? "", + ActiveText = row.IsActive ? "active" : "INACTIVE", + WalStatus = row.WalStatus ?? "", + RetainedWal = Bytes(row.RetainedWalBytes), + RetainedWalTrend = ByteDelta(row.FirstRetainedWalBytes, row.RetainedWalBytes), + SafeWalSize = Bytes(row.SafeWalSizeBytes), + XminAge = Count(row.XminAge), + CatalogXminAge = Count(row.CatalogXminAge), + InactiveSince = Timestamp(row.InactiveSince), + DatabaseName = row.DatabaseName ?? "", + Plugin = row.Plugin ?? "", + InvalidationReason = row.InvalidationReason ?? "", + Conflicting = row.Conflicting ? "CONFLICTING" : "", + MeasuredAt = Timestamp(row.MeasuredAt), + /* An invalidated slot has already lost its WAL: the replica behind it needs rebuilding, and that + is a different day from an inactive slot that is merely accumulating. A CONFLICTING slot is in + the same category — its subscriber cannot continue — so it shares the highlight even though the + cause and the fix are different. */ + IsInvalidated = !string.IsNullOrEmpty(row.InvalidationReason) + || row.Conflicting + || string.Equals(row.WalStatus, "lost", StringComparison.OrdinalIgnoreCase), + IsInactive = !row.IsActive, + }; + + // ── Activity ───────────────────────────────────────────────────────────────────────────────── + + internal sealed class ChainRow + { + public string CapturedAt { get; init; } = ""; + public int RootPid { get; init; } + + /// The collector's synthetic backend id, which is what "seen as root" is counted on — pids + /// are reused, so two captures agreeing on a pid are not necessarily the same backend. 0 is + /// the vanished-blocker sentinel and shows as unknown rather than as an id. + public string RootBackendId { get; init; } = ""; + + public string Databases { get; init; } = ""; + public int TotalVictims { get; init; } + public int DirectVictims { get; init; } + public string Depth { get; init; } = ""; + public string WorstVictimWait { get; init; } = ""; + public string RootState { get; init; } = ""; + public string RootXactDuration { get; init; } = ""; + public string RootUsername { get; init; } = ""; + public string RootApplicationName { get; init; } = ""; + public string SamplesAsRoot { get; init; } = ""; + public string RootQuery { get; init; } = ""; + public string WorstVictimQuery { get; init; } = ""; + public string Caveats { get; init; } = ""; + } + + internal static ChainRow Chain(DarlingPgBlockingReader.PgBlockingChainRow row) + { + var caveats = new List(); + if (row.ChainMayBeTruncated) + { + caveats.Add("chain hit the 32-level walk cap"); + } + if (row.QueryTextMayBeTruncated) + { + caveats.Add("query text truncated at track_activity_query_size"); + } + + return new ChainRow + { + CapturedAt = Timestamp(row.CapturedAt), + RootPid = row.RootPid, + RootBackendId = row.RootBackendId == 0 + ? "unknown" + : row.RootBackendId.ToString("N0", CultureInfo.CurrentCulture), + Databases = string.Join(", ", row.Databases.Where(d => !string.IsNullOrEmpty(d))), + TotalVictims = row.TotalVictims, + DirectVictims = row.DirectVictims, + Depth = row.MaxDepth.ToString(CultureInfo.CurrentCulture) + (row.ChainMayBeTruncated ? " (capped)" : ""), + WorstVictimWait = Milliseconds(row.WorstVictimWaitMs), + /* idle in transaction is the state worth naming on its own: the root is not running anything, + so nothing will finish and release the lock without intervention. */ + RootState = row.RootIsIdleInTransaction ? "idle in transaction" : row.RootState ?? "", + RootXactDuration = Milliseconds(row.RootXactDurationMs), + RootUsername = row.RootUsername ?? "", + RootApplicationName = row.RootApplicationName ?? "", + /* NULL is "cannot tell" — the blocker's own row had already left pg_stat_activity, so it landed + on the collector's synthetic backend id and counting it would conflate unrelated incidents. */ + SamplesAsRoot = row.SamplesAsRoot is { } samples + ? string.Create(CultureInfo.CurrentCulture, $"{samples:N0} captures") + : "unknown", + RootQuery = row.RootQuery ?? "", + WorstVictimQuery = row.WorstVictimQuery ?? "", + Caveats = string.Join("; ", caveats), + }; + } + + internal sealed class CycleRow + { + public string CapturedAt { get; init; } = ""; + public int ParticipantCount { get; init; } + public string Pids { get; init; } = ""; + public string DatabaseName { get; init; } = ""; + public string ApplicationName { get; init; } = ""; + public int BlockedBehindCount { get; init; } + public string BlockedBehindPids { get; init; } = ""; + } + + internal static CycleRow Cycle(DarlingPgBlockingReader.PgBlockingCycleRow row) => new() + { + CapturedAt = Timestamp(row.CapturedAt), + ParticipantCount = row.ParticipantCount, + Pids = string.Join(", ", row.Pids), + DatabaseName = row.DatabaseName ?? "", + ApplicationName = row.ApplicationName ?? "", + BlockedBehindCount = row.BlockedBehindCount, + BlockedBehindPids = string.Join(", ", row.BlockedBehindPids), + }; + + internal sealed class StatementRow + { + /// A string, not a number. PostgreSQL queryid is an int8 that routinely + /// exceeds 2^53, and rendering it through anything that rounds is the one field whose entire + /// purpose is joining back to pg_stat_statements (#2548). + public string QueryId { get; init; } = ""; + + /// The database OID, which is half the grain of this read — one queryid appears once per + /// database it ran in, and two rows with the same statement text are not a duplicate. + public string DatabaseId { get; init; } = ""; + + public string Calls { get; init; } = ""; + public string TotalExecTimeMs { get; init; } = ""; + public string AvgExecTimeMs { get; init; } = ""; + public double MaxExecTimeMs { get; init; } + public string RowsReturned { get; init; } = ""; + + /// Buffer-cache hit share for this statement, over BOTH cache tiers Aurora reports — the + /// shared buffers and the Optimized Read cache — against both miss sources. Reported as one ratio + /// rather than four block counters, which are the same fact in a form nobody reads correctly. + public string CacheHitPct { get; init; } = ""; + + public string TempRead { get; init; } = ""; + public string TempWritten { get; init; } = ""; + public string PeakMemory { get; init; } = ""; + public string WalWritten { get; init; } = ""; + public string QueryText { get; init; } = ""; + } + + internal static StatementRow Statement(DarlingPgStatementReader.PgStatementRow row) => new() + { + QueryId = row.QueryId.ToString(CultureInfo.InvariantCulture), + DatabaseId = row.DatabaseId.ToString("N0", CultureInfo.CurrentCulture), + Calls = Count(row.Calls), + TotalExecTimeMs = Count(row.TotalExecTimeMs), + AvgExecTimeMs = row.Calls <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{row.TotalExecTimeMs / (double)row.Calls:N1}"), + MaxExecTimeMs = Math.Round(row.MaxExecTimeMs, 1), + RowsReturned = Count(row.RowsReturned), + /* #2625: the Aurora-only halves are NULL on a self-hosted target, where there is no Optimized + Reads tier and no storage volume to read from. Treating them as zero would still produce a + NUMBER here — a cache-hit ratio computed over shared blocks alone, presented in the same column + as Aurora's four-way one, with nothing saying the two mean different things. So the ratio is + computed only when the split is present, and reads as not-applicable when it is not. */ + CacheHitPct = row.OrcacheBlocksHit is { } orcacheHit && row.StorageBlocksRead is { } storageRead + ? Percent( + row.SharedBlocksHit + orcacheHit, + row.SharedBlocksHit + orcacheHit + row.SharedBlocksRead + storageRead) + : NotApplicableText, + /* BLOCKS, and labelled as such — both of them. The block size is a compile-time setting of the + server, so multiplying by an assumed 8kB would put a fabricated byte count next to real ones. + Temp READ sits beside temp written because a spill that is written and never read back is a + different (and cheaper) event than one the query then re-reads. */ + TempRead = Blocks(row.TempBlocksRead), + TempWritten = Blocks(row.TempBlocksWritten), + /* Aurora-only too: core PostgreSQL has no per-statement peak-memory figure at all, so this is + not-applicable rather than a zero-byte grant. */ + PeakMemory = row.MaxPeakMemBytes is { } peak ? Bytes(peak) : NotApplicableText, + WalWritten = Bytes(row.WalBytes), + /* Null is "no text captured for this queryid yet" — text refreshes hourly, and a major-version + upgrade re-keys every queryid — which is a different statement from an empty query. */ + QueryText = row.QueryText ?? "(statement text not captured yet)", + }; + + internal sealed class DatabaseRow + { + public string DatabaseName { get; init; } = ""; + public string TempFiles { get; init; } = ""; + public string TempBytes { get; init; } = ""; + public string Deadlocks { get; init; } = ""; + public string CacheHitPct { get; init; } = ""; + public string XactCommit { get; init; } = ""; + public string XactRollback { get; init; } = ""; + public string RollbackPct { get; init; } = ""; + public int SampleCount { get; init; } + + /// The stored stats_reset timestamp. Shown beside the caveat that counts resets in + /// the window, because "reset once" and "reset at 14:02" send you to different places. + public string StatsReset { get; init; } = ""; + + public string Caveats { get; init; } = ""; + } + + internal static DatabaseRow Database(DarlingPgDatabaseReader.PgDatabaseRow row) + { + var caveats = new List(); + /* Every total beside it is a LOWER BOUND when the counters were reset inside the window, so this + has to sit on the row rather than in a footnote nobody reads. */ + if (row.StatsResetCount > 0) + { + caveats.Add($"counters reset {row.StatsResetCount}x in window"); + } + if (row.CounterRewindCount > 0) + { + caveats.Add($"counters rewound {row.CounterRewindCount}x"); + } + + return new DatabaseRow + { + DatabaseName = row.DatabaseName ?? "", + TempFiles = Count(row.TempFiles), + TempBytes = Bytes(row.TempBytes), + Deadlocks = Count(row.Deadlocks), + CacheHitPct = Percent(row.BlksHit, row.BlksHit + row.BlksRead), + XactCommit = Count(row.XactCommit), + XactRollback = Count(row.XactRollback), + RollbackPct = Percent(row.XactRollback, row.XactCommit + row.XactRollback), + SampleCount = row.SampleCount, + StatsReset = Timestamp(row.StatsReset), + Caveats = string.Join("; ", caveats), + }; + } + + internal sealed class TableBloatRow + { + public string DatabaseName { get; init; } = ""; + public string SchemaName { get; init; } = ""; + public string TableName { get; init; } = ""; + public string HeapSize { get; init; } = ""; + + /// The estimate, or the not-applicable dash when it was SUPPRESSED. Never the raw number + /// with a caption beside it: a percentage rendered in a grid cell is read as a measurement whatever + /// text sits next to it, and this one can be 81 percentage points wrong. + public string BloatEstimate { get; init; } = ""; + + public string BloatPctEstimate { get; init; } = ""; + + /// The MEASURED fallback, from the server's own counters and needing no width model. + /// Always populated, so a suppressed estimate still leaves the row saying something true. + public string DeadTuplePct { get; init; } = ""; + + public string ToastSize { get; init; } = ""; + public string IndexSize { get; init; } = ""; + public string HeapGrowth { get; init; } = ""; + + /// Why the estimate is or is not shown, in a few words. The full reason is in the MCP + /// read; this is what fits in a column and still stops somebody acting on a blank. + public string Confidence { get; init; } = ""; + + public bool EstimateSuppressed { get; init; } + + /// Only ever true for an estimate that was PUBLISHED - a suppressed row has no percentage + /// to be high, and painting it red would assert the very thing the suppression denies. + public bool IsHighBloat { get; init; } + } + + /// + /// Projects one bloat row for the grid. + /// + /// The suppression decision comes from + /// - the SAME call the MCP read makes - rather than from a second copy of the rule here. Two matching + /// literals in two projects is exactly how the two surfaces would come to disagree about whether a + /// number is publishable, silently and in the direction that shows a figure the other surface had + /// already withheld. + /// + internal static TableBloatRow TableBloat(DarlingPgTableBloatReader.PgTableBloatRow row) + { + var suppressed = DarlingPgTableBloatReader.EstimateIsUnpublishable(row); + + var confidence = row.EstimateUnavailable + ? "no column statistics - check the monitoring login can SELECT this table" + : row.LiveTuples < 0 + ? "never analyzed - no row count to compare against" + : suppressed + ? "statistics stale - ANALYZE this table, then re-read" + : row.PgstattupleAvailable + ? "estimate; confirm exactly with pgstattuple" + : "estimate; pgstattuple is not installed here"; + + return new TableBloatRow + { + DatabaseName = row.DatabaseName ?? "", + SchemaName = row.SchemaName ?? "", + TableName = row.TableName ?? "", + HeapSize = Bytes(row.HeapBytes), + BloatEstimate = suppressed ? NotApplicableText : Bytes(row.BloatBytesEstimate), + BloatPctEstimate = suppressed + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, $"{row.BloatPctEstimate:N2}%"), + DeadTuplePct = row.LiveTuples < 0 + ? NotApplicableText + : Percent(row.DeadTuples, row.LiveTuples + row.DeadTuples), + ToastSize = Bytes(row.ToastBytes), + IndexSize = Bytes(row.IndexBytes), + HeapGrowth = ByteDelta(row.FirstHeapBytes, row.HeapBytes), + Confidence = confidence, + EstimateSuppressed = suppressed, + IsHighBloat = !suppressed && row.BloatPctEstimate >= 50m, + }; + } + + internal sealed class IndexUsageRow + { + public string DatabaseName { get; init; } = ""; + public string TableName { get; init; } = ""; + public string IndexName { get; init; } = ""; + public string IndexSize { get; init; } = ""; + public string ScansInWindow { get; init; } = ""; + + /// The server's own lifetime counter, shown BESIDE the windowed figure rather than + /// instead of it: they answer different questions, and an operator who checks in psql must not + /// find this grid appearing to contradict the server. + public string TotalScans { get; init; } = ""; + + public string LastScan { get; init; } = ""; + public string BlockAccesses { get; init; } = ""; + public int SampleCount { get; init; } + + /// The short form of the droppability answer. Deliberately phrased as an answer to a + /// question rather than as an instruction: no cell in this grid ever reads "drop it". + public string Droppability { get; init; } = ""; + + public string IndexDefinition { get; init; } = ""; + public bool IsInvalid { get; init; } + + /// Unscanned AND unblocked AND watched long enough - all three, so the amber never lands + /// on a constraint index or on one we have only just met. + public bool IsUnscanned { get; init; } + } + + internal static IndexUsageRow IndexUsage(DarlingPgIndexUsageReader.PgIndexUsageRow row) + { + var blocked = row.IsPrimaryKey || row.SupportsConstraint || row.IsUnique || row.IsReplicaIdentity; + var unscanned = row.IsValid && !blocked && row.ScansInWindow == 0 && row.SampleCount >= 2; + + var droppability = !row.IsValid + ? "INVALID - never used by the planner, still maintained by writes" + : row.IsPrimaryKey + ? "no - backs the PRIMARY KEY" + : blocked + ? (row.IsReplicaIdentity ? "no - is the REPLICA IDENTITY" : "no - backs a constraint") + : row.SampleCount < 2 + ? "too early to say - only one sample" + : row.ScansInWindow > 0 + ? "in use this window" + : row.IsPartial + ? "candidate - but PARTIAL, check the predicate still matches" + : row.IsExpression + ? "candidate - but EXPRESSION, check the expression still matches" + : "candidate - widen the window past your slowest job first"; + + return new IndexUsageRow + { + DatabaseName = row.DatabaseName ?? "", + TableName = row.TableName ?? "", + IndexName = row.IndexName ?? "", + IndexSize = Bytes(row.IndexBytes), + ScansInWindow = Count(row.ScansInWindow), + TotalScans = Count(row.TotalScans), + /* NULL means two different things - PostgreSQL 15 and below do not record it at all, and on + 16+ it is an index never scanned since the reset - so the dash stands for "not recorded" + rather than being filled with a fabricated date. The MCP read spells out which. */ + LastScan = Timestamp(row.LastScan), + BlockAccesses = Count(row.BlocksHit + row.BlocksRead), + SampleCount = row.SampleCount, + Droppability = droppability, + IndexDefinition = row.IndexDefinition ?? "", + IsInvalid = !row.IsValid, + IsUnscanned = unscanned, + }; + } + + // ── Session states ─────────────────────────────────────────────────────────────────────────── + + /// + /// The share of the samples that saw a session in which it must have been the oldest xmin holder + /// before the grid paints it as one. + /// + /// Half, and it is a COPY of DarlingMcpPgSessionStatesTools.SustainedHolderSampleShare + /// rather than a reference to it: the viewer does not reference the service project, so the constant + /// cannot be reached from here. The duplication is deliberate and pinned — + /// ViewerPostgresTabsTests asserts this projection's flags agree with the MCP's own severity + /// band on the same row, in the one assembly that references both — because the alternative is two + /// surfaces quietly disagreeing about whether a session is starving vacuum. + /// + /// The threshold itself is the argument from the read: every write transaction is momentarily + /// the oldest holder, so a single sighting proves only that the instance does writes. Sustained + /// holding is the finding. + /// + internal const double SustainedHolderSampleShare = 0.5; + + /// + /// How long a session must have sat idle in transaction, while pinning NOTHING, before the grid says + /// anything about it. Five minutes, the copy of + /// DarlingMcpPgSessionStatesTools.IdleWithoutHorizonAttentionMs — see above for why it is a copy. + /// It costs vacuum nothing, so the amber is not the vacuum argument: at five minutes it is holding + /// a connection and whatever locks the transaction already took, which is a forgotten commit. + /// + internal const long IdleWithoutHorizonAttentionMs = 5 * 60 * 1000; + + /// What the horizon-age cell shows for the -1 sentinel. WORDS, not the em dash the rest + /// of this class uses for a missing measurement, and that difference is the entire feature: -1 here is a + /// positive finding — this session pinned nothing — and the dash would file it under "not measured", + /// which is the reading that gets a harmless session killed. + internal const string PinsNothingText = "pins nothing"; + + internal sealed class SessionStateRow + { + public int Pid { get; init; } + + /// The collector's synthetic (backend_start, pid) identity, shown because a pid alone is + /// not one: pids are reused, and this is the id that matches the same backend on the blocking + /// grid. + public long BackendId { get; init; } + + public string DatabaseName { get; init; } = ""; + public string Username { get; init; } = ""; + public string ApplicationName { get; init; } = ""; + public string ClientAddr { get; init; } = ""; + public string BackendType { get; init; } = ""; + public string LastState { get; init; } = ""; + public string LastWait { get; init; } = ""; + + /// The leading SQL keyword, whitelisted at collection. Not a truncation of the statement: + /// pg_stat_activity.query carries literal parameter values, so no raw text is stored anywhere. + public string LastCommandTag { get; init; } = ""; + + /// The normalised statement identity, to join against the Activity tab's query grid. Blank + /// rather than 0 when absent — PostgreSQL 13 has no such column, and on 14+ it is NULL whenever + /// compute_query_id is off, neither of which is a query whose id happens to be zero. + public string LastQueryId { get; init; } = ""; + + /// The column this panel exists for. An age in transactions, and the words in + /// when the session held neither a snapshot nor a transaction id in + /// any sample. + public string PeakHorizonAge { get; init; } = ""; + + public string PeakXminAge { get; init; } = ""; + public string PeakXidAge { get; init; } = ""; + public string PeakXactDuration { get; init; } = ""; + public string PeakStateDuration { get; init; } = ""; + public string PeakQueryDuration { get; init; } = ""; + + /// How long the BACKEND has existed, against how long it has held its transaction. The + /// pair separates two bugs that look identical in the transaction duration alone: a connection + /// created ten minutes ago that has been idle in transaction for all ten is a pool handing out a + /// session nobody finished with; a three-day-old worker that has held one for ten minutes is a + /// code path that forgot to commit. + public string BackendAge { get; init; } = ""; + + /// How often this session was the oldest holder, as a fraction of the samples that saw it — + /// never a bare count. One sighting in a hundred is normal write traffic; ninety-eight in a hundred + /// is the reason vacuum reclaims nothing, and a count alone cannot tell them apart. + public string HorizonHoldShare { get; init; } = ""; + + public string IdleInTransactionShare { get; init; } = ""; + public int SampleCount { get; init; } + public string FirstSeenAt { get; init; } = ""; + public string LastSeenAt { get; init; } = ""; + + /// What the rest of the instance looked like in this backend's most recent sample. Two + /// idle-in-transaction sessions out of six connections is a different server from two out of four + /// thousand, and the row cannot be read without it. + public string InstanceContext { get; init; } = ""; + + /// Every qualifier on the row in one column — redaction, and a capture that hit the + /// collector's per-capture cap so the stored set is a worst-first sample of a larger one. + public string Caveats { get; init; } = ""; + + /// Sustained holder: the red. Only ever true for a session seen holding the oldest xmin + /// across at least of the samples that saw it, so a passing + /// sighting — which every write transaction produces — is not painted as a cause. + public bool IsSustainedHorizonHolder { get; init; } + + /// Idle in transaction past the attention threshold: the amber. A different finding from + /// the red rather than a weaker one — the red is a session PROVEN to be setting the horizon, this + /// is one holding a connection and its locks for minutes, so the two must never share a colour. + /// Not gated on whether it pinned anything: is_horizon_holder means OLDEST on the + /// instance, so several real long transactions can each take turns being oldest and none of them + /// reach the red. The horizon columns say separately whether this one pinned. + public bool IsLongIdleTransaction { get; init; } + + /// The row's state columns came back NULL because the monitoring login lacks pg_monitor. + /// Not a severity: nothing on this row is a trustworthy observation about the database, and painting + /// it either healthy or critical would invent one. + public bool StateUnknown { get; init; } + } + + /// + /// Projects one session-state row for the grid. + /// + /// PeakHorizonAge of -1 renders as words, not as the dash. Everywhere else in + /// this class the -1 sentinel means "not measured" and the em dash says so. Here it means the + /// opposite of missing: the session held neither a snapshot nor a transaction id in any sample, which + /// was MEASURED, and it is the finding that separates a session starving vacuum from one costing it + /// nothing. Rendering it as a dash — or worse, as a number — is how a monitoring tool talks somebody + /// into killing a harmless backend. + /// + /// The three flags are re-derived here rather than read off a severity string, because the + /// XAML row triggers bind to booleans and the MCP's band is a word. They are pinned against that band in + /// the tests so the two surfaces cannot drift apart. + /// + internal static SessionStateRow SessionState(DarlingPgSessionStatesReader.PgSessionStateRow r) + { + /* Redaction is asked FIRST and suppresses both other flags, for the same reason the read's finding + does: with the state columns NULL, the inputs to every judgement below are missing, and a flag + derived from them would report an absent GRANT as an observation about the workload. */ + var unknown = r.StateWasRedacted; + + var sustained = !unknown + && r.SampleCount > 0 + && r.HorizonHolderSamples >= Math.Max(1, (int)Math.Ceiling(r.SampleCount * SustainedHolderSampleShare)); + + /* Deliberately NOT gated on PeakHorizonAge < 0, and the gate that used to be here was a real + cross-surface bug. is_horizon_holder means "the OLDEST holder on the instance", not "holds + anything" - so on a busy instance several genuinely long idle transactions each take turns + being oldest, none of them clears the sustained threshold, and every one of them pinned + something. Requiring "pinned nothing" here painted all of them Healthy while the MCP band + called them Warning. The band is "long idle transaction", full stop; whether it pins is an + orthogonal fact the horizon columns already carry. */ + var longIdleTransaction = !unknown + && r.IdleInTransactionSamples > 0 + && r.PeakStateDurationMs >= IdleWithoutHorizonAttentionMs; + + var caveats = new List(); + if (r.StateWasRedacted) + { + caveats.Add("state columns REDACTED — the monitoring login lacks pg_monitor here"); + } + + if (r.CaptureWasTruncated) + { + caveats.Add("a capture hit the per-capture row cap, so this is a worst-first sample of more"); + } + + /* wait_event_type and wait_event are two halves of one name ("Lock: transactionid") and are worth + little apart: the type alone is a category, the event alone repeats across categories. */ + var wait = string.Join(": ", new[] { r.LastWaitEventType, r.LastWaitEvent } + .Where(part => !string.IsNullOrWhiteSpace(part))); + + return new SessionStateRow + { + Pid = r.Pid, + BackendId = r.BackendId, + DatabaseName = r.DatabaseName ?? "", + Username = r.Username ?? "", + ApplicationName = r.ApplicationName ?? "", + ClientAddr = r.ClientAddr ?? "", + BackendType = r.BackendType ?? "", + LastState = r.LastState ?? "", + LastWait = wait, + LastCommandTag = r.LastCommandTag ?? "", + LastQueryId = r.LastQueryId is { } queryId + ? queryId.ToString("N0", CultureInfo.CurrentCulture) + : string.Empty, + PeakHorizonAge = r.PeakHorizonAge < 0 + ? PinsNothingText + : string.Create(CultureInfo.CurrentCulture, $"{r.PeakHorizonAge:N0} transactions"), + /* These two keep the ordinary dash. They are the components of the horizon age above, and an + absent one is the narrower statement "this session held no snapshot" / "no transaction id" — + which the horizon column has already said in words. */ + PeakXminAge = Count(r.PeakXminAge), + PeakXidAge = Count(r.PeakXidAge), + /* PEAKS, not averages. The finding is how far this went, and averaging a hundred samples of a + transaction that grew monotonically reports about half of what actually happened. */ + PeakXactDuration = Milliseconds(r.PeakXactDurationMs), + PeakStateDuration = Milliseconds(r.PeakStateDurationMs), + PeakQueryDuration = Milliseconds(r.PeakQueryDurationMs), + BackendAge = Milliseconds(r.PeakBackendDurationMs), + HorizonHoldShare = r.SampleCount <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, + $"{r.HorizonHolderSamples:N0} of {r.SampleCount:N0} samples"), + IdleInTransactionShare = r.SampleCount <= 0 + ? NotApplicableText + : string.Create(CultureInfo.CurrentCulture, + $"{r.IdleInTransactionSamples:N0} of {r.SampleCount:N0} samples"), + SampleCount = r.SampleCount, + FirstSeenAt = Timestamp(r.FirstSeenAt), + LastSeenAt = Timestamp(r.LastSeenAt), + InstanceContext = string.Create(CultureInfo.CurrentCulture, + $"{r.IdleInTransactionSessions:N0} idle in xact / {r.ActiveSessions:N0} active / " + + $"{r.TotalSessions:N0} sessions; {r.ReportableSessions:N0} reportable"), + Caveats = string.Join("; ", caveats), + IsSustainedHorizonHolder = sustained, + IsLongIdleTransaction = longIdleTransaction, + StateUnknown = unknown, + }; + } + + /// One write-side metric, as a display row. + internal sealed class WriteStatRow + { + public string Group { get; init; } = ""; + public string Metric { get; init; } = ""; + + /// The change across the window, or an em dash when the value is NULL. The dash is + /// deliberate and is NOT rendered as 0 — a metric can be null because this PostgreSQL major does not + /// expose it, or because its statistics family was reset inside the window, and both of those are + /// "we do not know" rather than "nothing happened". + public string Value { get; init; } = ""; + + public string Note { get; init; } = ""; + } + + /// + /// Projects the single write-side row into one display row per metric (#2544), in causal reading order: + /// checkpoints, then what the background writer did between them, then the WAL they were driven by. + /// + /// A metric whose value is NULL is kept rather than dropped, and carries the reason in its note. + /// Hiding it would answer the reader's next question ("where is buffers_backend?") with silence, and on + /// a mixed-version fleet the answer differs per target: PostgreSQL 17 moved that counter to + /// pg_stat_io, and 18 removed the WAL timing columns outright. + /// + internal static List WriteStatsRows(DarlingPgWriteStatsReader.PgWriteStatsRow row) + { + ArgumentNullException.ThrowIfNull(row); + + return new List + { + Metric("Checkpoints", "Timed", row.CheckpointsTimed, + "Checkpoints that began because checkpoint_timeout elapsed. This is the healthy kind."), + Metric("Checkpoints", "Requested", row.CheckpointsRequested, + "Began because WAL volume demanded one. Climbing against Timed is the classic " + + "max_wal_size-too-small signal."), + Metric("Checkpoints", "Completed", row.CheckpointsDone, + "PostgreSQL 18+. Blank on earlier majors, which do not expose it."), + MetricMs("Checkpoints", "Write Time (ms)", row.CheckpointWriteTimeMs, + "Time spent writing buffers during checkpoints."), + MetricMs("Checkpoints", "Sync Time (ms)", row.CheckpointSyncTimeMs, + "Time spent in fsync during checkpoints. High here with low write time points at the " + + "storage rather than at the volume."), + Metric("Checkpoints", "Buffers Written", row.BuffersWrittenCheckpoint, + "Buffers written by the checkpointer."), + Metric("Checkpoints", "SLRU Written", row.SlruWritten, + "PostgreSQL 18+. Blank on earlier majors."), + Metric("Restartpoints", "Timed", row.RestartpointsTimed, + "A standby's equivalent of a checkpoint. PostgreSQL 17+, and zero on a primary."), + Metric("Restartpoints", "Requested", row.RestartpointsRequested, + "PostgreSQL 17+, and zero on a primary."), + Metric("Restartpoints", "Completed", row.RestartpointsDone, + "PostgreSQL 17+. A standby where requested climbs and completed does not is falling behind " + + "on replay."), + Metric("Background Writer", "Buffers Cleaned", row.BuffersClean, + "Buffers the background writer wrote out ahead of demand."), + Metric("Background Writer", "Hit Max Written", row.MaxwrittenClean, + "Cleaning rounds that stopped early because bgwriter_lru_maxpages was reached. Sustained " + + "non-zero means the background writer is CAPPED rather than idle."), + Metric("Background Writer", "Buffers Allocated", row.BuffersAlloc, + "Buffers handed out — the denominator for how hard the pool is being churned."), + Metric("Background Writer", "Buffers Written By Backends", row.BuffersBackend, + "Queries writing their own buffers because nothing else kept up. Blank on PostgreSQL 17+, " + + "which moved this to pg_stat_io — see the grid above, not a zero here."), + Metric("Background Writer", "Backend fsyncs", row.BuffersBackendFsync, + "Blank on PostgreSQL 17+ for the same reason."), + Metric("WAL", "Records", row.WalRecords, "WAL records generated."), + Metric("WAL", "Full Page Images", row.WalFpi, + "Full-page writes. Spiking right after each checkpoint is the checkpoint_timeout-too-low " + + "shape, and it is only legible next to the checkpoint counts above."), + MetricBytes("WAL", "Bytes", row.WalBytes, "WAL volume generated across the window."), + Metric("WAL", "Buffers Full", row.WalBuffersFull, + "Times a backend had to flush WAL because wal_buffers was full."), + Metric("WAL", "Writes", row.WalWrite, + "Blank on PostgreSQL 18+, which removed the WAL write/sync counters."), + Metric("WAL", "Syncs", row.WalSync, "Blank on PostgreSQL 18+."), + MetricMs("WAL", "Write Time (ms)", row.WalWriteTimeMs, "Blank on PostgreSQL 18+."), + MetricMs("WAL", "Sync Time (ms)", row.WalSyncTimeMs, "Blank on PostgreSQL 18+."), + }; + + /* Three overloads rather than one taking a formatted string, so a caller cannot accidentally pass + "0" where the value was null. The em dash is produced HERE, once, from an actual null. */ + static WriteStatRow Metric(string group, string metric, long? value, string note) => + Row(group, metric, value?.ToString("N0", CultureInfo.CurrentCulture), note); + + static WriteStatRow MetricMs(string group, string metric, double? value, string note) => + Row(group, metric, value?.ToString("N1", CultureInfo.CurrentCulture), note); + + static WriteStatRow MetricBytes(string group, string metric, decimal? value, string note) => + Row(group, metric, value?.ToString("N0", CultureInfo.CurrentCulture), note); + + static WriteStatRow Row(string group, string metric, string? formatted, string note) => new() + { + Group = group, + Metric = metric, + /* EM DASH, never "0". A null here means either that this PostgreSQL major does not expose the + counter or that its statistics family was reset inside the window - both "unknown", and both + the opposite of "nothing happened". */ + Value = formatted ?? "\u2014", + Note = note, + }; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresTabs.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresTabs.cs new file mode 100644 index 000000000..82662b57a --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPostgresTabs.cs @@ -0,0 +1,230 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Linq; +using PerformanceMonitor.Collectors; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// One PostgreSQL inner tab: which inner-tab index it occupies, what its header reads, and which +/// COLLECTORS serve it. Panels are declared by collector NAME rather than by store table, so the table can +/// only ever be what says it is — the name is also the key +/// takes, so a panel's data and a panel's +/// explanation for having none are keyed on the same string by construction. +/// +/// Its position in ViewerServerTab.xaml's InnerTabs. +/// A stable identifier for pins and for the SQL Server registry's deep links to collide +/// with deliberately (overview means the same place on either engine). +/// The tab strip label, which must match the XAML <TabItem Header=>. +/// Every collector whose stored rows this tab renders, in panel order. +/// Why these panels sit together, when that is not obvious from the header. Rendered +/// above the tab's panels. Null for a tab whose single panel needs no framing. +internal sealed record ViewerPostgresTab( + int InnerTabIndex, + string Id, + string Header, + IReadOnlyList Collectors, + string? Note); + +/// +/// The PostgreSQL inner-tab registry — the viewer's half of #2530, and the direct counterpart of the web +/// dashboard's POSTGRES_TABS in server-tabs.js (#2547). Same seven tabs, same ids, same +/// grouping, so the two front ends do not teach an operator two different shapes for one engine. +/// +/// Why a registry rather than "read the XAML". The SQL Server surface has nineteen inner tabs +/// declared in XAML and dispatched by integer index, and nothing there is machine-readable enough to ask +/// "does every PostgreSQL collector reach a screen?". That question is the whole point of this issue: the +/// first nine PostgreSQL collectors spent three releases MCP-only precisely because no check could see that the +/// graphical surface had fallen behind the data. This list is what such a check reads, and +/// ViewerPostgresTabsTests derives the other side of it from — so a +/// THIRTEENTH PostgreSQL collector turns the pin red naming itself, rather than quietly shipping invisible. +/// +/// Seven tabs against SQL Server's nineteen is the design. Parity was never the constraint; +/// the missing tabs (tempdb, Query Store, trace flags, plan cache, system_health, Always On) have no +/// PostgreSQL analogue to fill, and the signals that DO matter here — wraparound headroom, the xmin +/// horizon, vacuum backlog, WAL retention — have no SQL Server analogue either. A PostgreSQL surface built +/// by asking "what does the SQL Server viewer have" would have missed most of them. +/// +/// The Aurora-only panel is SHOWN, not hidden. pg_wait_stats gates on +/// CollectorTargetInfo.IsAurora — it reads Aurora's own wait instrumentation, and +/// pg_wait_sampling is the stock-PostgreSQL answer to the same question — so on stock PostgreSQL it +/// can never have content. (pg_statement_stats was in this paragraph until #2625, when it learned to +/// read the vanilla pg_stat_statements view; its panel now fills on every PostgreSQL target.) Since +/// #2532 the reason is a sentence +/// () naming the server, the engine, the +/// collector and the exact Aurora surface, ending "and never will". The defect #2530 is about is +/// UNEXPLAINED emptiness, not emptiness: a panel that says that is the opposite of the twelve blank SQL +/// Server tabs it replaces. Hiding them would also make the tab strip a different shape on two PostgreSQL +/// servers in one fleet, and would make the one Aurora-specific capability the product has invisible from +/// a stock instance — while re-deriving in the viewer a gate the collectors already decide. +/// +internal static class ViewerPostgresTabs +{ + /// + /// The seven tabs, in strip order. Indices continue the SQL Server run rather than displacing it: for a + /// PostgreSQL server every SQL Server TabItem is collapsed and these seven are shown, so both sets + /// keep their fixed indices and ViewerServerTab's existing index constants — which drill-down + /// navigation keys on — are untouched by this feature. + /// + public static readonly IReadOnlyList All = new[] + { + new ViewerPostgresTab( + ViewerServerTab.PgOverviewInnerTabIndex, + "overview", + "Overview", + new[] { "pg_extension_availability", "pg_server_config" }, + /* No collector of its own: it reports on every one of them, from collection_log joined to the + catalog. + That join is the point — a collector gated off for this engine writes NO collection_log row + at all (pre-dispatch filtering, deliberately, so a permanent gap does not manufacture ~2,880 + fake SUCCESS rows a day), so a grid built from the log alone would show a PostgreSQL + operator one row short and never mention the missing collector. Derived from the catalog, it + appears with the reason it is missing. */ + "Every PostgreSQL collector for this server: when it last ran, what it returned, and — for a " + + "collector this engine cannot run at all — why it never will. The extension panel below is the " + + "third capability axis (#2545) and the only one that is ACTIONABLE: engine kind and edition " + + "are walls, but an extension that is available and not installed is one command away."), + + new ViewerPostgresTab( + ViewerServerTab.PgActivityInnerTabIndex, + "activity", + "Activity", + new[] { "pg_blocking", "pg_lock_stats", "pg_statement_stats", "pg_database_stats", "pg_kernel_stats", "pg_plan_capture", "pg_deadlocks" }, + /* pg_database_stats sits under the statement grid rather than on a tab of its own because it + answers the question the statement grid raises and cannot answer: a statement whose time + makes no sense from its row count usually spilled, and pg_stat_database's temp counters are + the only evidence of that we collect. */ + "What ran, what waited on what, and what it cost. Blocking is a periodic SAMPLE rather than an " + + "event log, so the capture counts above the grid are its denominator — three chains means " + + "something different in 60 captures than in 4."), + + new ViewerPostgresTab( + ViewerServerTab.PgVacuumInnerTabIndex, + "vacuum", + "Vacuum", + new[] { "pg_session_states", "pg_xmin_horizon", "pg_autovacuum_stats", "pg_wraparound_stats", "pg_plan_capture_readiness" }, + /* One tab, four panels, in causal order. Read separately each of the four looks survivable. + + Session states is FIRST rather than on a tab of its own because it is the link UPSTREAM of + what was previously the first panel, and the order here is the causal one. pg_xmin_horizon + names the CLASS of thing holding the horizon — a session, a slot, standby feedback, a + prepared transaction — and #2540 is the panel that names WHICH session, which is where the + fix has to be made: the horizon panel can tell an operator the problem is a backend, and + only this one can tell them which application opened it. + + It also carries the correction the rest of the tab cannot make. A long idle-in-transaction + session is the shape everybody recognises and it is NOT automatically a cause: measured on a + live instance, a READ COMMITTED transaction that only read, and one whose UPDATE matched no + rows, both sat idle in transaction indefinitely holding neither a snapshot nor a transaction + id. Those sessions starve vacuum of nothing, and a tab that opened with a vacuum backlog + would have an operator killing them. */ + "One story in four panels: a session holds a transaction open, that transaction holds the xmin " + + "horizon, the horizon starves vacuum, and vacuum falling behind ends in wraparound. Read on " + + "their own each of these looks survivable. Only the first panel can name the session, and " + + "only it says whether that session pins anything at all — an open transaction that holds no " + + "snapshot and no transaction id costs vacuum nothing, however long it has been idle."), + + new ViewerPostgresTab( + ViewerServerTab.PgWaitsInnerTabIndex, + "waits", + "Waits", + /* Two instruments for one question, and which of them has rows says what the server + offers: pg_wait_stats is Aurora-native, pg_wait_sampling is the extension (#2603) that + brings wait analysis to everything else. */ + new[] { "pg_wait_stats", "pg_wait_sampling" }, + null), + + new ViewerPostgresTab( + ViewerServerTab.PgIoInnerTabIndex, + "io", + "I/O", + new[] { "pg_io_stats", "pg_write_stats", "pg_buffer_usage" }, + /* The write side sits on the I/O tab rather than a tab of its own because it is the same + subject from the other end. pg_stat_io says WHO issued an I/O and in what context; the write + panel says whether the server is keeping up with the writes it was given. They also complete + each other on 17+, where buffers_backend left pg_stat_bgwriter and its successor information + lives only in pg_stat_io - so on a modern target the two panels together are the answer that + either alone stopped being. */ + "Two ends of the same subject: what issued the I/O, and whether the server kept up with it. " + + "The write panel reports the CHANGE across the window rather than the counters' levels, " + + "because a cumulative total since the last stats reset answers nothing on its own."), + + new ViewerPostgresTab( + ViewerServerTab.PgReplicationInnerTabIndex, + "replication", + "Replication", + new[] { "pg_replication_slots", "pg_replication_stats" }, + /* Slots and connections are different facts and the tab carries both because either alone + misleads: a slot with no standby attached is the classic way to fill a disk, and a standby + streaming without a slot can be cut off the moment the primary recycles WAL it still needed. */ + "Two halves of one question. A SLOT is a promise to retain WAL and exists whether or not " + + "anybody is attached; a CONNECTION is a standby actually streaming. The connection panel " + + "reports the worst each standby reached across the window, not just where it is now - a " + + "replica that falls far behind and recovers looks perfect in any single sample."), + + new ViewerPostgresTab( + ViewerServerTab.PgStorageInnerTabIndex, + "storage", + "Storage", + new[] { "pg_table_bloat_stats", "pg_index_usage_stats", "pg_index_bloat", "pg_column_stats", "pg_predicate_stats" }, + /* One tab, because both panels answer the same question — where is the space going and is it + earning its keep — and the two remedies compete for the same maintenance window. Bloat is + deliberately NOT on the Vacuum tab despite being what vacuum lag costs: that tab is the + CAUSE chain read in causal order, and dropping the damage into the middle of it would break + the sequence that makes those three panels one story. The bloat panel's own note points back + at it instead. + + Bloat sits above index usage because it is the more urgent of the two and the more + dangerous to act on: a bloat percentage is an ESTIMATE and the panel has to say so before + anyone reads a number off it. */ + "Where the space went, and whether it is earning its keep. The bloat figures are ESTIMATES " + + "computed from column-width statistics — the table itself is never read — so confirm one with " + + "pgstattuple before rewriting anything. An index nothing scans is a candidate, never a " + + "conclusion: check the constraint and validity columns beside it first. The column-statistics " + + "panel is the INPUT those bloat estimates are computed from, and it answers a different " + + "question of its own: why the planner chose what it chose."), + }; + + /// The first PostgreSQL tab's index — what a PostgreSQL server's tab strip selects, since + /// every index below it belongs to a collapsed SQL Server tab. + public static int FirstInnerTabIndex => All[0].InnerTabIndex; + + /// True when belongs to this registry. + public static bool Owns(int innerTabIndex) => All.Any(t => t.InnerTabIndex == innerTabIndex); + + /// + /// A tab's framing note by ID — empty for a tab that has none, and empty rather than throwing for an id + /// this registry does not carry, because a missing note must never take a tab down with it. Looked up by + /// id rather than by position so the strip order stays free to change. + /// + public static string NoteFor(string tabId) => + All.FirstOrDefault(t => string.Equals(t.Id, tabId, StringComparison.Ordinal))?.Note ?? string.Empty; + + /// + /// The store table a declared collector writes, from the collector's own definition. Null for a name + /// the catalog does not know, which ViewerPostgresTabsTests refuses — the registry may not name + /// a collector that does not exist. + /// + public static string? TableOf(string collectorName) => + CollectorCatalog.Find(collectorName)?.TargetTable; + + /// + /// Every PostgreSQL collector the catalog ships, derived rather than listed. The pins compare this + /// against in BOTH directions: a collector missing from the registry is a screen that + /// was never built, and a registry entry that is not a PostgreSQL collector is a panel wired to + /// something that will never fill it. + /// + public static IReadOnlyList PostgresCollectors() => + CollectorCatalog.All + .Where(d => d.TargetEngine == CollectorTargetEngine.PostgreSql) + .OrderBy(d => d.Name, StringComparer.Ordinal) + .ToList(); +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerPreferences.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPreferences.cs index 1047854c1..0e43b4677 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerPreferences.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerPreferences.cs @@ -8,9 +8,9 @@ using System; using System.Collections.Generic; -using System.Diagnostics; using System.IO; using System.Text.Json; +using PerformanceMonitor.Common; namespace PerformanceMonitor.Darling.Viewer; @@ -92,6 +92,9 @@ public sealed class ViewerPreferencesStore { private static readonly JsonSerializerOptions s_jsonOptions = new() { WriteIndented = true }; + /// The source every diagnostic from this store is filed under. + private const string LogSource = "ViewerPreferencesStore"; + private readonly string _filePath; /// Override the on-disk location (tests pass a temp file); null uses . @@ -103,6 +106,26 @@ public ViewerPreferencesStore(string? filePath = null) /// The resolved settings file path (surfaced mainly so tests and diagnostics can name it). public string FilePath => _filePath; + /// + /// What the last on this instance found. Same three-way split as + /// , and for the same reason: a missing + /// viewer-preferences.json is a first run and must stay silent, while a present one that could not be + /// read is a configuration the viewer is currently ignoring. + /// + public SettingsFileState LastLoadState { get; private set; } = SettingsFileState.Absent; + + /// Why the last could not read the file, or null when it could. + public string? LastLoadProblem { get; private set; } + + /// + /// The settings the last could not read, by NAME, and empty when there were none + /// (#2456). Non-empty means the rest of the file loaded normally and only these reverted to their + /// defaults — the distinction alone cannot make, and the reason the + /// startup dialog can now say which settings were lost instead of only where the parse stopped. + /// + public IReadOnlyList LastLoadUnreadableMembers { get; private set; } = + Array.Empty(); + /// %APPDATA%\PerformanceMonitorDarling\viewer-preferences.json. public static string DefaultFilePath() { @@ -113,42 +136,29 @@ public static string DefaultFilePath() } /// - /// Reads the settings, returning defaults when the file does not exist yet or cannot be read/parsed — - /// the viewer never blocks on a first run or a corrupt file. Loaded values are clamped to valid ranges. + /// Reads the preferences, returning defaults when the file does not exist yet or cannot be + /// read/parsed. Loaded values are clamped to valid ranges. A present-but-unreadable file is reported + /// rather than silently treated as absent (#2434). /// public ViewerPreferences Load() { - try - { - if (!File.Exists(_filePath)) - { - return new ViewerPreferences(); - } - - var json = File.ReadAllText(_filePath); - var preferences = JsonSerializer.Deserialize(json, s_jsonOptions); - return (preferences ?? new ViewerPreferences()).Normalize(); - } - catch (Exception ex) - { - /* The viewer writes no application log of its own (its diagnostics go to Debug output), so a - corrupt-file fallback is a Debug trace, not a log entry — and never a crash on startup. */ - Debug.WriteLine($"ViewerPreferencesStore: failed to load '{_filePath}', using defaults: {ex.Message}"); - return new ViewerPreferences(); - } + var read = ViewerSettingsFile.Load(_filePath, LogSource, s_jsonOptions); + LastLoadState = read.State; + LastLoadProblem = read.Problem; + LastLoadUnreadableMembers = read.UnreadableMembers ?? Array.Empty(); + return read.Value!.Normalize(); } - /// Writes the settings as indented JSON, creating the app-data directory on first save. - public void Save(ViewerPreferences preferences) + /// + /// Writes the preferences as indented JSON, creating the app-data directory on first save, and returns + /// whether the write happened. Collapsing a sidebar group is a whole-file rewrite of this document and + /// nobody thinks of it as a save, which is precisely why the unreadable case is copied aside first and + /// a refusal is reported rather than swallowed (#2434). + /// + public bool Save(ViewerPreferences preferences) { ArgumentNullException.ThrowIfNull(preferences); - var directory = Path.GetDirectoryName(_filePath); - if (!string.IsNullOrEmpty(directory)) - { - Directory.CreateDirectory(directory); - } - - File.WriteAllText(_filePath, JsonSerializer.Serialize(preferences, s_jsonOptions)); + return ViewerSettingsFile.Save(_filePath, preferences, LogSource, s_jsonOptions); } } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerSeatIndicator.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerSeatIndicator.cs new file mode 100644 index 000000000..0125021c9 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerSeatIndicator.cs @@ -0,0 +1,148 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// Which seat this viewer is connected with, as one status-bar field (#2479, items 3 and 4). +/// +/// Why this exists. #2400 asked why "Current Active Queries" refuses on a read-write-looking +/// connection, and #2403 answered it by rewriting the individual refusals. That improved every refusal and +/// left the question intact: you still discover your seat one surface at a time, at the moment a write is +/// denied, which is the worst moment to learn it. The viewer has known the answer since V8 — one +/// has_table_privilege probe at connect — and never showed it anywhere you could look on purpose. +/// +/// Why it matters more in UAT than it did before. The two defaults are deliberately opposite: +/// a local seat takes postgres.connectAs, which defaults to admin (read-write), while a seat +/// reached over the LAN takes postgres.network.role, which defaults to viewer (read-only). A +/// UAT tester connecting from their own desktop is the second case by construction, so the read-only seat +/// is the EXPECTED default rather than a misconfiguration — and nothing in the product said so. +/// +/// Unreachable is its own state, not a read-only verdict. DetectReadOnlyAsync fails safe +/// to read-only for an unexpected RESPONSE, but raises for a +/// failure to REACH the store — a distinction #2117 built deliberately and this field must not undo. A +/// status line claiming "read-only" when the service is simply down would send an operator to fix a role +/// they do not need to touch. So the state machine carries four values and the two failures never share +/// one. +/// +internal enum ViewerSeatState +{ + /// Not probed yet. The field's initial value, and where a schema-skew refusal leaves it — + /// the store answered, but the seat question was never asked, so nothing is claimed about it. + Unknown, + + /// The store could not be reached, so there is no seat to report. Never collapsed into + /// : see the type remarks. + Unreachable, + + /// The probe said this connection cannot write the operator-config tables — either the + /// restricted role, or a response that could not be read as a yes (the V8 fail-safe). + ReadOnly, + + /// The probe said this connection can write the operator-config tables. + ReadWrite, +} + +/// The status-bar text and tooltip for a . Pure, so the wording is +/// pinned by tests rather than by reading a XAML file. +internal static class ViewerSeatIndicator +{ + /// + /// The status-bar field. Prefixed "Seat:" to match the other four fields ("Collectors:", "Database:", + /// "Collection:", "Servers:"), which is what makes it findable without a legend. + /// + public static string Text(ViewerSeatState state) => state switch + { + ViewerSeatState.ReadOnly => "Seat: read-only", + ViewerSeatState.ReadWrite => "Seat: read-write", + ViewerSeatState.Unreachable => "Seat: not connected", + _ => "Seat: --", + }; + + /// + /// The tooltip, which is where the ANSWER lives rather than in a doc. + /// + /// orders the two defaults by which one is likelier to + /// have decided this seat; it never claims which one DID. A non-loopback store host does not prove a + /// remote seat — a bring-your-own store on another host reached from the service host reads + /// identically — so the wording says the role arrived in the connection string and then names both + /// defaults, which is true either way. + /// + public static string ToolTip(ViewerSeatState state, bool storeIsOnThisMachine) => state switch + { + ViewerSeatState.ReadOnly => ReadOnlyToolTip(storeIsOnThisMachine), + ViewerSeatState.ReadWrite => ReadWriteToolTip, + ViewerSeatState.Unreachable => UnreachableToolTip, + _ => UnknownToolTip, + }; + + internal const string LocalDefaultSentence = + "A seat on the service host connects as postgres.connectAs, which defaults to \"admin\" — read-write."; + + internal const string NetworkDefaultSentence = + "A seat reached over the LAN connects as postgres.network.role, which defaults to \"viewer\" — read-only."; + + /* Said in full on the read-only tooltip because it is the actual #2400 question. The refusals say it + per-dialog already; a reader who has not tripped one yet has no way to know a "write" here is a row + in the MONITORING store rather than something aimed at their SQL Server. */ + internal const string NothingIsWrittenSentence = + "Nothing is ever written to a monitored SQL Server. The actions a read-only seat cannot take " + + "(Refresh on Current Active Queries, Get Actual Plan, Purge Now, Generate now, dismissing an " + + "alert, editing mute rules, adding a server) work by enqueueing a request for the service in the " + + "MONITORING store, and that row is the write a read-only seat cannot make."; + + private static string ReadOnlyToolTip(bool storeIsOnThisMachine) + { + /* Order by likelihood, claim neither. The likelier default goes first so the sentence that + explains THIS seat is the one the reader hits first. */ + var defaults = storeIsOnThisMachine + ? LocalDefaultSentence + " " + NetworkDefaultSentence + : NetworkDefaultSentence + " " + LocalDefaultSentence; + + var where = storeIsOnThisMachine + ? "The store is on this machine." + : "The store is not on this machine, so the role came from the connection string the service handed out."; + + return "This viewer is connected with a read-only role, so write affordances are hidden or disabled." + + Environment.NewLine + Environment.NewLine + + where + " " + defaults + + " The two defaults are deliberately opposite — the local seat is the operator's, the remote " + + "seat is a laptop — so a remote seat being read-only is the expected default, not a fault." + + Environment.NewLine + Environment.NewLine + + NothingIsWrittenSentence + + Environment.NewLine + Environment.NewLine + + "To change it: set postgres.network.role to \"admin\" on the SERVICE host and restart the " + + "service (then re-run --export-viewer-config) for a LAN seat, or set postgres.connectAs to " + + "\"admin\" and restart the viewer for a seat on the service host."; + } + + /* static readonly rather than const so the paragraph breaks come from Environment.NewLine, exactly as + the read-only tooltip's do. A const cannot call it, and two tooltips breaking lines two different + ways is the kind of drift nobody notices until a test pins one of them. */ + internal static readonly string ReadWriteToolTip = + "This viewer is connected with a read-write role, so it can enqueue requests for the service " + + "(Refresh on Current Active Queries, Get Actual Plan, Purge Now, adding a server, editing mute " + + "rules) by writing to the MONITORING store." + + Environment.NewLine + Environment.NewLine + + "Nothing is ever written to a monitored SQL Server. The service reads those under the same " + + "least-privilege login the collectors use." + + Environment.NewLine + Environment.NewLine + + LocalDefaultSentence + " " + NetworkDefaultSentence; + + /* Deliberately says nothing about the seat. The probe never ran, so there is no verdict to report, + and the full-window message already carries the diagnosis and the connection details. */ + internal const string UnreachableToolTip = + "The Darling store could not be reached, so this viewer has not probed which seat it has. This is " + + "NOT a read-only seat — it is no connection at all. See the message on the main panel for the " + + "store address it tried and why it failed."; + + internal const string UnknownToolTip = + "This viewer has not probed which seat it has yet."; +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerStore.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerStore.cs index 2d8b59174..3738026f7 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerStore.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerStore.cs @@ -8,7 +8,6 @@ using System; using System.Collections.Generic; -using System.Diagnostics; using System.IO; using System.Linq; using System.Text.Json; @@ -91,6 +90,9 @@ public sealed class ViewerServerStore { private static readonly JsonSerializerOptions s_jsonOptions = new() { WriteIndented = true }; + /// The source every diagnostic from this store is filed under. + private const string LogSource = "ViewerServerStore"; + private readonly string _filePath; private readonly IViewerServerSecretStore _secrets; private readonly List _servers; @@ -107,6 +109,39 @@ public ViewerServerStore(string? filePath = null, IViewerServerSecretStore? secr /// The resolved registry file path (surfaced for tests and diagnostics). public string FilePath => _filePath; + /// + /// What the load in the constructor found. is a first run and + /// says nothing; means the operator's registry is on disk, + /// could not be read, and is being ignored — which is the difference between "you have not added any + /// servers yet" and "your servers are still in that file" (#2434). + /// + public SettingsFileState LastLoadState { get; private set; } = SettingsFileState.Absent; + + /// Why the registry could not be read, or null when it could. + public string? LastLoadProblem { get; private set; } + + /// + /// Present so all three stores answer the same three questions, and ALWAYS empty here — which is the + /// point rather than an oversight. The member recovery #2456 added only edits a root JSON object, and + /// this file's root is an array: dropping a bad element would silently delete a monitored server from + /// the operator's registry, which is the data loss #2434 exists to prevent wearing a repair's clothes. + /// The registry stays all-or-nothing, and the guard has a control test that pins it. + /// + public IReadOnlyList LastLoadUnreadableMembers { get; private set; } = + Array.Empty(); + + /// + /// Whether the last write of the registry reached disk (#2434). Every mutator here ends in the same + /// , and several of them keep return types that already mean something else — + /// ToggleFavorite answers "is it favourite now", ImportServersFromFile answers with + /// counts — so the answer lives here rather than being crammed into those. A caller that is about to + /// tell the user something happened can ask; one doing incidental cleanup need not. + /// + /// True until a write is attempted, so "nothing has failed" is the starting position rather + /// than a claim about a write nobody made. + /// + public bool LastSaveSucceeded { get; private set; } = true; + /// %APPDATA%\PerformanceMonitorDarling\viewer-servers.json. public static string DefaultFilePath() { @@ -349,36 +384,30 @@ private void ReplaceOrAdd(ViewerServerEntry entry) } } + /// + /// Reads the registry, beginning empty when there is nothing usable to read — a corrupt registry must + /// never block startup. What changed with #2434 is what happens NEXT: beginning empty and then writing + /// that empty list back over the file was how an unreadable registry became a lost one, and the very + /// first Add Server did it. The state is recorded, the failure is reported to the viewer's log, and + /// copies the file aside before it replaces it. + /// private List LoadFromDisk() { - try - { - if (!File.Exists(_filePath)) - { - return new List(); - } - - var json = File.ReadAllText(_filePath); - return JsonSerializer.Deserialize>(json, s_jsonOptions) - ?? new List(); - } - catch (Exception ex) - { - /* A corrupt or unreadable registry must never block startup — begin empty and log. */ - Debug.WriteLine($"ViewerServerStore: failed to load '{_filePath}', starting empty: {ex.Message}"); - ViewerLogger.Warn("ViewerServerStore", $"Failed to load '{_filePath}': {ex.Message}"); - return new List(); - } + var read = ViewerSettingsFile.Load>(_filePath, LogSource, s_jsonOptions); + LastLoadState = read.State; + LastLoadProblem = read.Problem; + LastLoadUnreadableMembers = read.UnreadableMembers ?? Array.Empty(); + return read.Value!; } - private void Save() + /// + /// Persists the registry, and reports whether it reached disk. Every mutator here calls it — Add, Edit, + /// Delete, favourite, tag, import — so this is the whole-file replacement that stands behind an + /// ordinary click, exactly as the display-mode dropdown does for viewer-settings.json. + /// + private bool Save() { - var directory = Path.GetDirectoryName(_filePath); - if (!string.IsNullOrEmpty(directory)) - { - Directory.CreateDirectory(directory); - } - - File.WriteAllText(_filePath, JsonSerializer.Serialize(_servers, s_jsonOptions)); + LastSaveSucceeded = ViewerSettingsFile.Save(_filePath, _servers, LogSource, s_jsonOptions); + return LastSaveSucceeded; } } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CollectionHealth.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CollectionHealth.cs index 9a60b8485..0a38e481b 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CollectionHealth.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CollectionHealth.cs @@ -84,7 +84,9 @@ private async void PurgeNow_Click(object sender, RoutedEventArgs e) { MessageBox.Show( "Purging asks the service to run the retention purge, which it does by running a command — a " + - "read-only viewer seat can't enqueue commands. Reconnect with a read-write profile to purge.", + "read-only viewer seat can't enqueue commands. The command is queued in the MONITORING STORE, " + + "and the purge only ever deletes from the store — never from a monitored server. Reconnect " + + "with a read-write store profile to purge.", "Read-Only Viewer", MessageBoxButton.OK, MessageBoxImage.Information); return; } diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CurrentActiveQueries.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CurrentActiveQueries.cs index 903c848fe..f503c5e6e 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CurrentActiveQueries.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.CurrentActiveQueries.cs @@ -32,6 +32,18 @@ namespace PerformanceMonitor.Darling.Viewer; /// public partial class ViewerServerTab { + /// + /// #2400: the read-only refusal names WHICH write is blocked, because the previous wording ("reconnect + /// with a read-write profile") reads as "grant this viewer write access to my production SQL Server" — + /// which is the one thing it does not mean, and a reasonable person declines rather than asks. The write + /// is an INSERT into config_command in the monitoring store; the monitored server is only ever READ, by + /// the service, over its own connection. + /// + private const string ReadOnlySeatMessage = + "Read-only viewer — the live fetch queues a request for the service in the MONITORING STORE, which a " + + "read-only seat can't write to. Nothing is sent to or written on the monitored server. Reconnect " + + "with a read-write store profile to fetch live active queries."; + /// Refresh button — fires the live fetch (never auto-fires; a live server hit is explicit). private async void RefreshCurrentActiveQueries_Click(object sender, RoutedEventArgs e) => await LoadCurrentActiveQueriesAsync(); @@ -42,8 +54,7 @@ private async Task LoadCurrentActiveQueriesAsync() rather than attempting an enqueue that would throw. */ if (_dataService.IsReadOnly) { - CurrentActiveQueriesStatus.Text = - "Read-only viewer — reconnect with a read-write profile to fetch live active queries."; + CurrentActiveQueriesStatus.Text = ReadOnlySeatMessage; return; } @@ -76,8 +87,7 @@ rather than attempting an enqueue that would throw. */ catch (ViewerReadOnlyException) { /* Grants changed under us between the IsReadOnly check and the enqueue. */ - CurrentActiveQueriesStatus.Text = - "Read-only viewer — reconnect with a read-write profile to fetch live active queries."; + CurrentActiveQueriesStatus.Text = ReadOnlySeatMessage; } catch (Exception ex) { diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.HypotheticalIndex.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.HypotheticalIndex.cs new file mode 100644 index 000000000..f989f7c50 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.HypotheticalIndex.cs @@ -0,0 +1,190 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Globalization; +using System.Text.Json; +using System.Threading.Tasks; +using System.Windows; +using System.Windows.Controls; +using System.Windows.Input; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// The one place a human can ask "would an index on this column actually help?" (#2612). +/// +/// +/// It hangs off the predicate-statistics grid and nowhere else, which is the shape the feature was scoped +/// to: on demand only, never scheduled, driven from a specific predicate row somebody is already +/// looking at. That grid names columns filtered on heavily with nothing supporting them; this turns each +/// one from a candidate into an answer. +/// +/// +/// +/// Until this existed the command had no caller at all — it shipped reachable only by hand-writing a row +/// into the command queue, which is how it was tested and is not a feature. +/// +/// +/// +/// The row's own two reasons are not the same question. Poor selectivity means an index might help. +/// A large estimate error means the planner is working from a wrong row count, and an index will not fix a +/// plan built on one — the experiment will usually say so, but the dialog says it first, because a "no" +/// that arrives after a round trip teaches less than one that arrives before it. +/// +/// +public partial class ViewerServerTab +{ + private async void TestHypotheticalIndex_Click(object sender, RoutedEventArgs e) + { + if (PgPredicateStatsGrid.SelectedItem is not DarlingPgPredicateStatsReader.PgPredicateStatRow row) + { + MessageBox.Show( + "Select a predicate row first. The experiment is about one column on one table, taken from " + + "the row you are looking at.", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Information); + return; + } + + if (string.IsNullOrWhiteSpace(row.SchemaName) || string.IsNullOrWhiteSpace(row.TableName) || string.IsNullOrWhiteSpace(row.ColumnName)) + { + MessageBox.Show( + "This row does not name a schema, table and column, so there is no candidate to test. That " + + "happens when the predicate could not be attributed to a single column — an expression, or a " + + "join condition.", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Information); + return; + } + + /* Said BEFORE the round trip, not after. The grid separates the two reasons a column looks + interesting and only one of them is this experiment's business; letting somebody spend a + planner round trip to be told that is a worse way to learn it. */ + var estimateWarning = row.WorstEstimateErrorRatio >= 10 + ? "\n\nNote: this predicate's worst estimate error is " + + row.WorstEstimateErrorRatio.ToString("N1", CultureInfo.CurrentCulture) + + "x, which means the planner is working from a wrong row count. An index does not fix a plan " + + "built on a bad estimate — statistics or correlated columns are the likelier story, and the " + + "answer below may be a confident 'no' for that reason." + : string.Empty; + + var confirm = MessageBox.Show( + $"Ask {_server.DisplayName} whether the planner would use an index on " + + $"{row.SchemaName}.{row.TableName} ({row.ColumnName})?\n\n" + + "Nothing is executed and no index is built. The candidate is visible only inside one session " + + "on the server, the statement is PLANNED rather than run, and the session is reset before the " + + "answer comes back." + + estimateWarning, + "Test an index", MessageBoxButton.OKCancel, MessageBoxImage.Question); + + if (confirm != MessageBoxResult.OK) + { + return; + } + + var args = JsonSerializer.Serialize(new + { + /* A STRING, like every other queryid that crosses a wire here: it is signed 64-bit and a JSON + number would round it in a double-decoding parser into an id that resolves to no statement. */ + queryid = row.QueryId.ToString(CultureInfo.InvariantCulture), + schemaName = row.SchemaName, + tableName = row.TableName, + columns = new[] { row.ColumnName }, + databaseName = row.DatabaseName, + }); + + try + { + Mouse.OverrideCursor = Cursors.Wait; + + var result = await _dataService.RunCommandAsync( + "test_hypothetical_index", _server.ServerId, args, requestedBy: Environment.UserName, + timeout: TimeSpan.FromSeconds(90)); + + ShowHypotheticalIndexResult(result, row); + } + catch (Exception ex) + { + MessageBox.Show( + $"The experiment could not be run: {ex.Message}", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Warning); + } + finally + { + Mouse.OverrideCursor = null; + } + } + + /// + /// Reports the verdict. A no is presented as an ANSWER rather than as a failure — it is the one + /// that saves somebody a maintenance window, and dressing it as an inconclusive run would waste it. + /// + private static void ShowHypotheticalIndexResult( + CommandResult? result, + DarlingPgPredicateStatsReader.PgPredicateStatRow row) + { + var subject = $"{row.SchemaName}.{row.TableName} ({row.ColumnName})"; + + if (result is null) + { + MessageBox.Show( + "The service did not answer within 90 seconds. The experiment plans a statement twice, which " + + "is normally instant — a wait this long usually means the service is not running or cannot " + + "reach that server. Nothing was left behind on it either way: the session is reset even " + + "when the call is abandoned.", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Warning); + return; + } + + if (!string.Equals(result.Status, "succeeded", StringComparison.OrdinalIgnoreCase)) + { + MessageBox.Show( + $"The experiment did not run for {subject}.\n\n{Explanation(result.ResultJson) ?? result.ResultStatus ?? "No reason was given."}", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Warning); + return; + } + + MessageBox.Show( + $"{subject}\n\n{Explanation(result.ResultJson) ?? "The experiment ran but returned no explanation."}", + "Test an index", MessageBoxButton.OK, MessageBoxImage.Information); + } + + /// + /// The service's own sentence, which already says whether the planner would switch and by how much. + /// Re-deriving a verdict here would give the Viewer and every other caller two ways to describe one + /// result, and they would eventually disagree. + /// + private static string? Explanation(string? resultJson) + { + if (string.IsNullOrWhiteSpace(resultJson)) + { + return null; + } + + try + { + using var document = JsonDocument.Parse(resultJson); + + foreach (var name in new[] { "explanation", "error", "message" }) + { + if (document.RootElement.TryGetProperty(name, out var value) + && value.ValueKind == JsonValueKind.String) + { + return value.GetString(); + } + } + } + catch (JsonException) + { + /* A result we cannot parse is reported as absent rather than as raw JSON: the caller's fallback + sentence is more use to somebody than a brace. */ + } + + return null; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.Postgres.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.Postgres.cs new file mode 100644 index 000000000..522f70b5f --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.Postgres.cs @@ -0,0 +1,947 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.Linq; +using System.Threading.Tasks; +using System.Windows; +using System.Windows.Controls; +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Darling.Storage; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// The seven PostgreSQL inner tabs (#2530) — which tab set a server gets, what each tab loads, and the +/// projections that turn stored rows into something a grid may show. +/// +/// Why the projections exist rather than binding the reader records directly. Every one of +/// these reads carries at least one value a grid must not print raw. -1 is the store's +/// not-applicable sentinel for a duration or a size (0 was rejected because it reads as "started this +/// instant" / "retains nothing"); pg_stat_io's write counters are genuinely ABSENT on Aurora rather +/// than zero, because backends there do not write data files; recurrence is NULL when the blocker's own row +/// had already left pg_stat_activity, and "cannot tell" is a different claim from "seen once"; and +/// every timestamp is naive UTC that has to go through like every +/// other timestamp the viewer renders. A DataGrid bound straight at the record would show -1, a +/// misleading 0, and UTC. +/// +/// Every panel says why it is empty. asks +/// first — the same sentence the MCP surface and +/// the web dashboard print, naming the server, the engine, the collector and the exact surface it would have +/// read, ending "and never will". That is what makes the two Aurora-only panels worth SHOWING on stock +/// PostgreSQL rather than hiding: the defect #2530 is about is unexplained emptiness, not emptiness, and a +/// tab set that changed shape between two PostgreSQL servers in one fleet would be its own confusion. +/// +public partial class ViewerServerTab +{ + /// + /// A row limit for the per-tab PostgreSQL grids. Generous — these are already server-side ranked reads + /// — but bounded, because a fleet-sized autovacuum backlog is tens of thousands of tables and a WPF + /// DataGrid asked to realise all of them is a hung UI thread rather than a slow one. + /// + private const int PgGridRowLimit = 200; + + /// + /// Picks the engine's tab set, once, in the constructor. Only a POSITIVE PostgreSQL claim switches: + /// is false for an absent kind, an unrecognised token and a + /// server that has never connected, none of which is evidence for either engine. Falling back to the + /// SQL Server tabs there is the pre-#2530 behaviour, unchanged — and it has to be, because the + /// unclaimed population is every server that has not reconnected since the engine-kind rung landed. + /// + private void ApplyEngineTabSet() + { + var engine = _server.EngineDescription; + if (engine is not null) + { + ServerEngineText.Text = engine; + ServerEngineBadge.Visibility = Visibility.Visible; + } + + if (!_server.IsPostgres) + { + return; + } + + /* Collapse the SQL Server run and reveal the PostgreSQL one. Visibility rather than removal, so + both sets keep their fixed indices: LoadInnerTabAsync dispatches on SelectedIndex and every + drill-down in the viewer navigates by one of the index constants above. */ + for (var i = 0; i < InnerTabs.Items.Count; i++) + { + if (InnerTabs.Items[i] is TabItem item) + { + item.Visibility = ViewerPostgresTabs.Owns(i) ? Visibility.Visible : Visibility.Collapsed; + } + } + + /* A collapsed TabItem can still be the SELECTED one — WPF hides the header and shows the content + anyway — so the selection has to move explicitly. Index 0 (the SQL Server Overview lanes) would + otherwise be what a PostgreSQL server opens on. */ + InnerTabs.SelectedIndex = ViewerPostgresTabs.FirstInnerTabIndex; + + /* The per-server database filter drives the SQL Server database-scoped reads and nothing else. On a + PostgreSQL server it would sit there offering to filter views that never consult it, which is a + worse answer than not offering. */ + DatabaseFilterButton.Visibility = Visibility.Collapsed; + + /* IsPostgres is only true for a token the describer recognises, so `engine` has words here by + construction; the coalesce is the compiler's, not a state this can reach. */ + PgEngineBanner.Text = $"{_server.DisplayName} runs {engine ?? "PostgreSQL"}."; + + /* Looked up by ID, never by position: the registry's order is the STRIP order and is free to change, + while these four assignments are about which tab gets which framing. Indexing All[0..2] would + have silently put the Vacuum note above the Activity grids the first time someone reordered. */ + PgOverviewNote.Text = ViewerPostgresTabs.NoteFor("overview"); + PgActivityNote.Text = ViewerPostgresTabs.NoteFor("activity"); + PgVacuumNote.Text = ViewerPostgresTabs.NoteFor("vacuum"); + PgStorageNote.Text = ViewerPostgresTabs.NoteFor("storage"); + } + + // ───────────────────────────────────────────────────────────────────────────────────────────── + // Empty-state prose + // ───────────────────────────────────────────────────────────────────────────────────────────── + + /// + /// What a panel says when it has nothing to show. Three different answers, because they need three + /// different responses from the person reading them: a collector this engine can never run (the shared + /// capability sentence — nothing to do), a zero-row window for a collector where zero IS the answer + /// (nothing wrong), and a zero-row window for a collector that should have data (something to chase). + /// Returns empty for a panel with rows: the rows speak for themselves and a standing sentence above a + /// full grid is noise. + /// + private string PanelNote(string collectorName, int rowCount, string healthyEmptyText) + { + var gap = CollectorEngineCapability.NotCollectedMessage( + _server.ServerName, _server.EngineEdition, _server.EngineKind, collectorName); + + if (gap is not null) + { + return gap; + } + + return rowCount > 0 ? string.Empty : healthyEmptyText; + } + + /// True when a panel's collector cannot run here — used to skip the read entirely rather than + /// spend a round trip proving a permanent gap is still permanent. + private bool PgCollectorIsGatedOff(string collectorName) => + CollectorEngineCapability.NotCollectedMessage( + _server.ServerName, _server.EngineEdition, _server.EngineKind, collectorName) is not null; + + // ───────────────────────────────────────────────────────────────────────────────────────────── + // Tab loaders + // ───────────────────────────────────────────────────────────────────────────────────────────── + + /// + /// Overview — every PostgreSQL collector for this server, from collection_log composed against + /// the CATALOG. Catalog-driven so a collector that is gated off for this engine (and therefore writes + /// no log row at all) is still a visible row carrying its own explanation, instead of the one line an + /// operator most needs being the one line missing. + /// + private async Task LoadPgOverviewAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + var collectors = ViewerPostgresTabs.PostgresCollectors(); + + var facts = await _dataService.GetPostgresCollectorLogFactsAsync( + _server.ServerId, startUtc, endUtc, collectors.Select(c => c.Name).ToList()); + + PgCollectorHealthGrid.ItemsSource = + ViewerDataService.BuildPostgresCollectorHealth(_server, collectors, facts); + + await LoadPgExtensionsAsync(startUtc, endUtc); + await LoadPgServerConfigAsync(); + } + + /// + /// The extension capability axis (#2545), under the collector grid because it answers the question that + /// grid raises and cannot: a collector missing for want of an extension is a SETUP step rather than a + /// permanent gap, and this names the step. + /// + /// The note leads with the ACTIONABLE count — extensions this product can use that are available + /// on the server and simply not installed — because that number is the entire reason to look. Zero of + /// them is also worth saying: it means the gap is a platform limit rather than a missed install. + /// + private async Task LoadPgExtensionsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_extension_availability")) + { + PgExtensionsGrid.ItemsSource = null; + PgExtensionsNote.Text = PanelNote("pg_extension_availability", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgExtensionAvailabilityAsync(_server.ServerId, startUtc, endUtc); + + PgExtensionsGrid.ItemsSource = rows; + + var actionable = rows.Count(r => r.IsMonitoringRelevant && string.Equals(r.State, "available", StringComparison.Ordinal)); + var outdated = rows.Count(r => r.IsMonitoringRelevant && string.Equals(r.State, "outdated", StringComparison.Ordinal)); + + PgExtensionsNote.Text = PanelNote("pg_extension_availability", rows.Count, + "The extension collector runs daily, so a server added in the last day has nothing here yet.") + + (rows.Count == 0 + ? string.Empty + : $" {actionable} extension(s) this product can use are available on this server and NOT " + + "installed — each is one CREATE EXTENSION away." + + (outdated > 0 + ? $" {outdated} are installed but behind the version the server offers, which is worth " + + "fixing with ALTER EXTENSION … UPDATE: a stale extension can be missing columns a " + + "collector reads, and that surfaces as a confusing error rather than as a gap." + : string.Empty) + + " \u201cInstalled\u201d is scoped to the database this server entry connects to — " + + "pg_extension is per-database while the server's offer is cluster-wide."); + } + + /// + /// The server's own configuration (#2658), under the extension axis because it is the same kind of fact + /// one layer in: extensions say what this server CAN do, settings say what it was told to do. + /// + /// Not scoped to the toolbar window, unlike every other panel on this tab. A configuration + /// is the state NOW rather than something that happened during an interval, and filtering it by the + /// window would return nothing for a server whose HOURLY collector last ran just outside it — which on + /// this screen reads as "this server has no configuration" rather than "widen the window". + /// + /// The note leads with pending_restart when there is one, because that is the only row here that + /// reports a DISAGREEMENT rather than a value: the file has been changed and reloaded, the running + /// server is still on the old value, and nothing else in the product would ever mention it. + /// + private async Task LoadPgServerConfigAsync() + { + if (PgCollectorIsGatedOff("pg_server_config")) + { + PgServerConfigGrid.ItemsSource = null; + PgServerConfigNote.Text = PanelNote("pg_server_config", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgServerConfigAsync(_server.ServerId); + + /* Non-default only in the grid: several hundred parameters sorted alphabetically is a dump, and the + ones somebody chose are the answer. The full set stays one MCP call away for anyone who wants it. */ + var chosen = rows.Where(r => !r.IsDefault).ToList(); + PgServerConfigGrid.ItemsSource = chosen; + + var pendingRestart = rows.Where(r => r.PendingRestart).Select(r => r.Name).ToList(); + + PgServerConfigNote.Text = PanelNote("pg_server_config", chosen.Count, + "This collector runs HOURLY, so a server added in the last hour has nothing here yet.") + + (chosen.Count == 0 + ? string.Empty + : $" {chosen.Count} setting(s) differ from the compiled-in default; the rest are omitted " + + "rather than truncated.") + + (pendingRestart.Count > 0 + ? $" {pendingRestart.Count} setting(s) are PENDING RESTART — " + + string.Join(", ", pendingRestart) + + " — meaning the configuration file has been changed and reloaded but the running " + + "server is still using the previous value. The file and the server disagree until the " + + "next restart, at which point behaviour changes with no deployment to explain it." + : string.Empty); + } + + /// + /// Activity — blocking (denominator, chains, cycles) and query shapes (statements over per-database + /// counters). All five reads fire together: the sub-tabs are two views of one load, so switching + /// between them needs no second round trip. + /// + private async Task LoadPgActivityAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var countsTask = _dataService.GetPgBlockingCaptureCountsAsync(_server.ServerId, startUtc, endUtc); + var chainsTask = _dataService.GetPgBlockingChainsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + var cyclesTask = _dataService.GetPgBlockingCyclesAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + /* Aurora-only, so on stock PostgreSQL the read is skipped and the panel carries the capability + sentence instead. Skipped rather than run-and-discarded: the answer is decided by the collectors' + own gate and cannot change between now and the query returning. */ + var statementsGatedOff = PgCollectorIsGatedOff("pg_statement_stats"); + var statementsTask = statementsGatedOff + ? Task.FromResult(new List()) + : _dataService.GetPgTopQueriesAsync(_server.ServerId, startUtc, endUtc); + + var databasesTask = _dataService.GetPgDatabaseStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + await Task.WhenAll(countsTask, chainsTask, cyclesTask, statementsTask, databasesTask); + + var counts = countsTask.Result; + var chains = chainsTask.Result; + var cycles = cyclesTask.Result; + + PgBlockingChainsGrid.ItemsSource = chains.Select(PgDisplay.Chain).ToList(); + PgBlockingCyclesGrid.ItemsSource = cycles.Select(PgDisplay.Cycle).ToList(); + PgBlockingCyclesExpander.Visibility = cycles.Count > 0 ? Visibility.Visible : Visibility.Collapsed; + + /* The denominator, always — not only when the grid is empty. "Three chains" means something + different in a window of 60 captures than in a window of 4, and the edge table cannot tell an + absent capture from a capture that found nothing: both are an absence of rows. This is the whole + reason the read carries capture counts from collection_log. */ + PgBlockingNote.Text = PgCollectorIsGatedOff("pg_blocking") + ? PanelNote("pg_blocking", 0, string.Empty) + : counts.CapturesTotal == 0 + ? "No blocking capture ran for this server in this window, so this grid being empty says " + + "nothing about whether anything blocked. Check the Overview tab for the pg_blocking " + + "collector's status." + : $"{counts.CapturesWithBlocking:N0} of {counts.CapturesTotal:N0} captures in this window " + + $"saw blocking{(chains.Count == 0 && cycles.Count == 0 ? " — and none of them produced a chain or a cycle in view" : "")}."; + + PgTopQueriesGrid.ItemsSource = statementsTask.Result.Select(PgDisplay.Statement).ToList(); + PgStatementsNote.Text = PanelNote("pg_statement_stats", statementsTask.Result.Count, + "No statement accumulated execution time in this window."); + + PgDatabaseStatsGrid.ItemsSource = databasesTask.Result.Select(PgDisplay.Database).ToList(); + PgDatabasesNote.Text = PanelNote("pg_database_stats", databasesTask.Result.Count, + "No database counter moved in this window."); + + await LoadPgLockStatsAsync(startUtc, endUtc); + await LoadPgWaitSamplingAsync(startUtc, endUtc); + await LoadPgKernelStatsAsync(startUtc, endUtc); + await LoadPgPredicateStatsAsync(startUtc, endUtc); + await LoadPgPlanCaptureAsync(startUtc, endUtc); + await LoadPgDeadlocksAsync(startUtc, endUtc); + } + + /// + /// The Locks sub-tab (#2544) — lock state by mode and relation, beside Blocking rather than inside it. + /// + /// The two answer different questions about the same event: Blocking has the blocked/blocker + /// pairs, and this has the MODE, which is what decides the remedy. An ungranted + /// AccessExclusiveLock is a DDL queue and everything arriving behind it will also queue; + /// RowExclusiveLock contention is ordinary write traffic. Identical pair shape, opposite + /// advice. + /// + /// The note leads with the QUEUE count rather than the row count, because a granted lock is not a + /// finding — a healthy server holds thousands — and a panel that opened with "412 rows" would bury the + /// three that matter. + /// + private async Task LoadPgLockStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_lock_stats")) + { + PgLockStatsGrid.ItemsSource = null; + PgLockStatsNote.Text = PanelNote("pg_lock_stats", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgLockStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgLockStatsGrid.ItemsSource = rows; + + var queued = rows.Where(r => !r.Granted).ToList(); + var totalCaptures = rows.Count == 0 ? 0 : rows[0].TotalCaptures; + + PgLockStatsNote.Text = rows.Count == 0 + ? PanelNote("pg_lock_stats", 0, + "No lock sample was taken for this server in this window, so an empty grid says nothing " + + "about whether anything contended. Check the pg_lock_stats collector on the Overview tab.") + : queued.Count == 0 + ? $"No lock was waiting in any of {totalCaptures:N0} captures — every lock held in this " + + "window was granted, which is the healthy answer rather than a missing read." + : $"{queued.Count:N0} lock queue(s) across {totalCaptures:N0} captures. These are SAMPLES, " + + "so the capture columns are the denominator: a queue seen once in 60 captures is a blip, " + + "and one seen in 55 is a standing problem. The Mode column decides the remedy — an " + + "ungranted AccessExclusiveLock is a DDL everything else is queued behind. A blank " + + "Relation with an OID is a lock in a different database, not a missing name."; + } + + /// + /// Wait events attributed to the query that waited (#2603). + /// + /// An empty grid here has THREE causes and they need different actions, so the note names + /// which one it is: the module is not loaded (an install step), the collector is gated off, or the + /// server genuinely waited on nothing. Collapsing those into "no data" is how a missing + /// extension reads as a healthy server. + /// + private async Task LoadPgWaitSamplingAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_wait_sampling")) + { + PgWaitSamplingGrid.ItemsSource = null; + PgWaitSamplingNote.Text = PanelNote("pg_wait_sampling", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgWaitSamplingAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgWaitSamplingGrid.ItemsSource = rows; + + var attributed = rows.Count(r => r.QueryId != 0); + var reset = rows.Any(r => r.CounterReset); + + PgWaitSamplingNote.Text = rows.Count == 0 + ? PanelNote("pg_wait_sampling", 0, + "No wait samples for this server in this window. The usual cause is that " + + "pg_wait_sampling is not in shared_preload_libraries — check the Extensions panel, " + + "which says whether it is installed, available or absent. If it IS loaded, an empty " + + "grid means the server waited on nothing worth sampling, which is the healthy answer.") + : $"{rows.Count:N0} wait event(s), {attributed:N0} attributed to a query. " + + "Est. Wait is samples multiplied by the profile period — an estimate from a sampling " + + "profiler, not a measured duration, so treat it as a ranking rather than a stopwatch. " + + "A Query ID of 0 is a background process rather than an unknown query, and CPU/Running " + + "is the backend on processor rather than waiting." + + (reset + ? " One or more counters RESET inside this window (a restart, or " + + "pg_wait_sampling_reset_profile), so those rows cover only the time since the reset." + : string.Empty); + } + + /// + /// The kernel's CPU and disk per query (#2603). + /// + /// The note has to say what zero read bytes means, because it is the reading most likely to + /// be got wrong: these counters measure I/O that reached the DEVICE, so a cached read is genuinely + /// zero and that is the healthy case, not a broken instrument. + /// + private async Task LoadPgKernelStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_kernel_stats")) + { + PgKernelStatsGrid.ItemsSource = null; + PgKernelStatsNote.Text = PanelNote("pg_kernel_stats", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgKernelStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgKernelStatsGrid.ItemsSource = rows; + + var reset = rows.Any(r => r.CounterReset); + + PgKernelStatsNote.Text = rows.Count == 0 + ? PanelNote("pg_kernel_stats", 0, + "No kernel statistics for this server in this window. The usual cause is that " + + "pg_stat_kcache is not installed - the Extensions panel says whether it is available " + + "here. Where it IS installed, this separates a query that was WAITING from one that " + + "was burning processor, which elapsed time alone cannot do.") + : $"{rows.Count:N0} statement(s) by OS CPU. Read bytes are I/O that reached the DEVICE, so " + + "zero with high CPU is a cached workload behaving well rather than a missing " + + "measurement. Query ID joins the statement grid above." + + (reset + ? " One or more counters RESET inside this window, so those rows cover only the " + + "time since the reset." + : string.Empty); + } + + /// + /// Predicates evaluated and how badly the planner estimated them (#2603). + /// + /// The note leads with the SAMPLE RATE because it is the number most likely to mislead: the + /// extension defaults to 1/max_connections, so a small count means the sampler fired rarely rather + /// than the predicate being rare. + /// + private async Task LoadPgPredicateStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_predicate_stats")) + { + PgPredicateStatsGrid.ItemsSource = null; + PgPredicateStatsNote.Text = PanelNote("pg_predicate_stats", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgPredicateStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgPredicateStatsGrid.ItemsSource = rows; + + var sampled = rows.Count == 0 ? 1.0 : rows[0].SampleRate; + var misestimated = rows.Count(r => r.WorstEstimateErrorRatio >= 10); + + PgPredicateStatsNote.Text = rows.Count == 0 + ? PanelNote("pg_predicate_stats", 0, + "No predicate statistics for this server in this window. pg_qualstats is created PER " + + "DATABASE, so it may be installed in one database and not another - the Extensions " + + "panel says which. Note also that its sample_rate defaults to 1/max_connections, so a " + + "lightly-used database can genuinely record nothing.") + : $"{rows.Count:N0} predicate(s), sampled at {sampled:P2} of executions - these counts are a " + + "SAMPLE, so a small number means the sampler fired rarely rather than the predicate " + + "being rare. Filtered % with no supporting index is the index candidate. " + + $"{misestimated:N0} predicate(s) show an estimate error of 10x or worse, which is a " + + "different problem: the planner does not understand that column, and an index will not " + + "fix a plan built on a wrong row count."; + } + + /// + /// Plans captured by auto_explain (#2566). + /// + /// An empty grid here almost always means a MISSING GRANT rather than a quiet server, so the + /// note says so: reading the log needs pg_read_server_files plus an explicit EXECUTE on pg_read_file, + /// and on Aurora or RDS the route does not exist at all. + /// + private async Task LoadPgPlanCaptureAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_plan_capture")) + { + PgCapturedPlansGrid.ItemsSource = null; + PgCapturedPlansNote.Text = PanelNote("pg_plan_capture", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgPlanCaptureAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgCapturedPlansGrid.ItemsSource = rows; + + var orphans = rows.Count(r => r.QueryId == 0); + + PgCapturedPlansNote.Text = rows.Count == 0 + ? PanelNote("pg_plan_capture", 0, + "No captured plans. The Plan Capture Readiness panel says which precondition is missing - " + + "usually auto_explain is not loaded, or the monitoring login cannot read the server log. " + + "That read needs pg_read_server_files AND an explicit GRANT EXECUTE on pg_read_file; the " + + "role alone does not carry it. On Aurora and RDS there is no filesystem to read, so this " + + "panel stays empty by design.") + : $"{rows.Count:N0} plan shape(s), ranked by TOTAL time - a plan that takes 40ms constantly " + + "costs more than one that took 900ms once. Captures counts how often the collector SAW " + + "the plan, not how often it ran. Plan JSON is redacted at collection: query text is " + + "dropped and literals are replaced, so nothing here carries customer values." + + (orphans > 0 + ? $" {orphans:N0} plan(s) have no query id, which means log_line_prefix lacks %Q - they " + + "cannot be joined to a statement until that is fixed." + : string.Empty); + } + + /// + /// Deadlocks reported in the window (#2661), last on this tab because they are blocking at its limit: a + /// chain the server had to break by cancelling somebody. + /// + /// An empty grid is the healthy answer AND the shape of an unreadable log, which is why the + /// note names the other check rather than leaving it. Deadlock reports need nothing configured on the + /// target — unlike plan capture there is no setting that suppresses them — so the only precondition is + /// being able to read the server log, which the plan-capture panel above already reports on because it + /// reads the same file. pg_stat_database's cumulative deadlock counter is the independent test: if that + /// moved and this is empty, the log is the problem rather than the server. + /// + private async Task LoadPgDeadlocksAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_deadlocks")) + { + PgDeadlocksGrid.ItemsSource = null; + PgDeadlocksNote.Text = PanelNote("pg_deadlocks", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgDeadlocksAsync(_server.ServerId, startUtc, endUtc); + + PgDeadlocksGrid.ItemsSource = rows; + + PgDeadlocksNote.Text = PanelNote("pg_deadlocks", rows.Count, + "No deadlock was reported in this window. That is the healthy answer, and it is also what an " + + "unreadable server log looks like — the plan-capture panel above reads the same file and says " + + "which it is.") + + (rows.Count == 0 + ? string.Empty + : " Sightings counts how often the collector saw the SAME report while it stayed inside " + + "the log tail it re-reads; a deadlock that genuinely recurred is its own row, because " + + "the process IDs differ."); + } + + /// + /// Vacuum — the sessions holding a transaction open, the xmin horizon, the autovacuum backlog and the + /// freeze headroom, in that causal order. One load for all four: they are one story, and reading them a + /// tab apart is how each of them ends up looking survivable. + /// + /// Session states leads because it is the only panel that can name the SESSION behind a pinned + /// horizon, and the only one that can say a long idle-in-transaction session pins nothing at all. + /// + private async Task LoadPgVacuumAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var sessionsTask = _dataService.GetPgSessionStatesAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + var xminTask = _dataService.GetPgXminHorizonAsync(_server.ServerId, startUtc, endUtc); + var autovacuumTask = _dataService.GetPgAutovacuumAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + var wraparoundTask = _dataService.GetPgWraparoundAsync(_server.ServerId, startUtc, endUtc); + var planCaptureTask = _dataService.GetPgPlanCaptureReadinessAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + await Task.WhenAll(sessionsTask, xminTask, autovacuumTask, wraparoundTask, planCaptureTask); + + PgSessionStatesGrid.ItemsSource = sessionsTask.Result.Select(PgDisplay.SessionState).ToList(); + + /* The healthy-empty sentence has to carry BOTH halves. Zero rows here is a real all-clear — the + collector stores nothing when every transaction is short — and it is an all-clear about a SAMPLE, + so a transaction that opened and closed between two captures left no trace to find. Saying only + the first half would overstate it; saying only the second would read as broken collection. */ + PgSessionStatesNote.Text = PgCollectorIsGatedOff("pg_session_states") + ? PanelNote("pg_session_states", 0, string.Empty) + : sessionsTask.Result.Count == 0 + ? "No session held a transaction open past the collector's floor in this window. Zero rows " + + "is the HEALTHY answer here rather than a missing read — the collector stores nothing " + + "when every transaction is short. It is also a SAMPLE taken once per collection cycle, " + + "not an event log: PostgreSQL records nothing about session state unless something " + + "asks, so a transaction that opened and closed between two captures is genuinely " + + "invisible here." + : "Duration is NOT evidence that a session is starving vacuum — read the Pinned Horizon " + + "column, where \"" + PgDisplay.PinsNothingText + "\" means the session held neither a " + + "snapshot nor a transaction id in any sample and terminating it would not reclaim one " + + "dead row. This is a SAMPLE at the collection interval, so a transaction that opened " + + "and closed between two captures never appears, and a grey row is one PostgreSQL " + + "redacted because the monitoring login lacks pg_monitor."; + + PgXminHorizonGrid.ItemsSource = xminTask.Result.Select(PgDisplay.Xmin).ToList(); + PgXminNote.Text = PanelNote("pg_xmin_horizon", xminTask.Result.Count, + "Nothing held the xmin horizon back in this window — no long-running transaction, replication " + + "slot, standby feedback or prepared transaction pinned an old xmin. Zero rows is the healthy " + + "answer here, not a missing read."); + + PgAutovacuumGrid.ItemsSource = autovacuumTask.Result.Select(PgDisplay.Autovacuum).ToList(); + PgAutovacuumNote.Text = PanelNote("pg_autovacuum_stats", autovacuumTask.Result.Count, + "No table reported an autovacuum backlog in this window. This collector runs on the WRITER " + + "only — pg_stat_user_tables reports zeros on a replica — so an empty grid on a reader is " + + "expected rather than informative."); + + PgWraparoundGrid.ItemsSource = wraparoundTask.Result.Select(PgDisplay.Wraparound).ToList(); + PgWraparoundNote.Text = PanelNote("pg_wraparound_stats", wraparoundTask.Result.Count, + "No per-database freeze headroom has been collected in this window."); + + PgPlanCaptureGrid.ItemsSource = planCaptureTask.Result + .Select(r => new + { + r.Facet, + /* Rendered as words, not a checkbox or a bare bool. The reader is being told whether a + PRECONDITION holds, and "False" beside a remedy sentence reads as a failure rather than as + a step not yet taken. */ + Satisfied = r.IsSatisfied ? "yes" : "no", + Observed = r.Observed ?? "(not reported)", + Detail = r.Detail ?? string.Empty, + }) + .ToList(); + + /* Three states, and they are genuinely different answers. Gated off is the engine sentence. Zero + rows means the collector has not run yet on a server that only just started collecting - NOT that + capture is impossible, which is the mistake worth heading off, because "no rows about readiness" + and "not ready" look identical. And when rows exist the panel says whether every facet is + satisfied, because the useful summary is the AND of them: one unsatisfied facet is enough to mean + no plans. */ + PgPlanCaptureNote.Text = PgCollectorIsGatedOff("pg_plan_capture_readiness") + ? PanelNote("pg_plan_capture_readiness", 0, string.Empty) + : planCaptureTask.Result.Count == 0 + ? "Plan-capture readiness has not been collected for this server yet. This is an hourly " + + "collector, so a server added in the last hour has nothing here — it does NOT mean plan " + + "capture is unavailable." + : planCaptureTask.Result.All(r => r.IsSatisfied) + ? "Every precondition for auto_explain capture is satisfied on this server, and a " + + "captured plan carries the query id that joins it back to pg_stat_statements. That " + + "means plans CAN be captured — not that this product is reading them, which is a " + + "separate thing: auto_explain writes to the server log, which a SQL connection " + + "cannot read." + : "At least one precondition is unmet, so either no execution plans are being captured " + + "by auto_explain here or the ones that are cannot be attributed. Read the rows in " + + "order — extension_available, library_loaded, capture_threshold, plan_text_setting, " + + "plan_attribution — each names the specific step and, on Aurora/RDS, whether it " + + "needs a parameter-group change and a reboot rather than a SET. plan_attribution is " + + "the one that is easy to miss: auto_explain puts no query id in the plan itself, so " + + "without %Q in log_line_prefix every captured plan is an orphan."; + } + + /// Waits — Aurora's cumulative wait counters. Shown on stock PostgreSQL too, where the panel + /// carries the capability sentence rather than a blank rectangle; see the type header. + private async Task LoadPgWaitsAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var rows = PgCollectorIsGatedOff("pg_wait_stats") + ? new List() + : await _dataService.GetPgWaitStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgWaitStatsGrid.ItemsSource = rows.Select(PgDisplay.Wait).ToList(); + PgWaitsNote.Text = PanelNote("pg_wait_stats", rows.Count, + "No wait time was recorded for this server in this window."); + } + + /// I/O — pg_stat_io, differenced over the window. + private async Task LoadPgIoAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var rows = await _dataService.GetPgIoAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgIoStatsGrid.ItemsSource = PgDisplay.IoRows(rows); + PgIoNote.Text = PanelNote("pg_io_stats", rows.Count, + "No backend / object / context combination did any I/O in this window. This view needs " + + "PostgreSQL 16 or newer, where pg_stat_io exists; below that the collector does not run and " + + "the Overview tab says so."); + + await LoadPgWriteStatsAsync(startUtc, endUtc); + await LoadPgBufferUsageAsync(startUtc, endUtc); + } + + /// + /// The buffer-pool panel (#2544) — what the memory is actually holding, the third end of the same + /// subject as the two grids above it. + /// + /// Needs the pg_buffercache extension. When it is absent the collector records an + /// ObjectMissing outcome and this panel says so AND says the remedy is one CREATE EXTENSION — + /// which is checkable on the Overview tab's extension panel, so the two halves meet. + /// + private async Task LoadPgBufferUsageAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_buffer_usage")) + { + PgBufferUsageGrid.ItemsSource = null; + PgBufferUsageNote.Text = PanelNote("pg_buffer_usage", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgBufferUsageAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgBufferUsageGrid.ItemsSource = rows; + + var total = rows.Count == 0 ? 0L : rows[0].PoolBuffersTotal; + var used = rows.Count == 0 ? 0L : rows[0].PoolBuffersUsed; + + PgBufferUsageNote.Text = rows.Count == 0 + ? "Nothing recorded. This panel needs the pg_buffercache extension — without it the collector " + + "reports the object as missing rather than failing, and the Overview tab's extension panel " + + "says whether it is available on this server and one CREATE EXTENSION away." + : $"Latest snapshot: {used:N0} of {total:N0} buffers in use " + + $"({(total == 0 ? 0 : 100.0 * used / total):N1}% of the pool). Residency is a LEVEL, so this " + + "is the newest sample rather than a window average — averaging what was resident over a day " + + "answers nothing. A blank Relation is another database's table or a shared catalog, not a " + + "missing name: the pool is cluster-wide while pg_class is per-database. Avg Usage near 0 " + + "means a relation is holding memory it is not earning."; + } + + /// + /// The write-side panel under the I/O grid (#2544) — checkpoints, background writer and WAL as the + /// CHANGE across the window. + /// + /// Three states, and they must not read alike. A gated-off collector says so; a null read means + /// fewer than two samples, which is a real and temporary state on a freshly added server rather than a + /// quiet one; and a row whose ResetDuringWindow is set has had at least one statistics family + /// reset underneath it, so those metrics are blank rather than wrong. + /// + private async Task LoadPgWriteStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_write_stats")) + { + PgWriteStatsGrid.ItemsSource = null; + PgWriteStatsNote.Text = PanelNote("pg_write_stats", 0, string.Empty); + return; + } + + var row = await _dataService.GetPgWriteStatsAsync(_server.ServerId, startUtc, endUtc); + + PgWriteStatsGrid.ItemsSource = row is null ? null : PgDisplay.WriteStatsRows(row); + PgWriteStatsNote.Text = row is null + ? "Write-side counters need TWO collections before a change exists between them, so a server " + + "added in the last cycle has nothing here yet. This is not the same as a quiet server, " + + "which would show zeroes." + : row.ResetDuringWindow + ? "At least one statistics family was RESET inside this window, so its counters went " + + "backwards. Those metrics are left blank rather than differenced across the reset — a " + + "difference taken across one reports an enormous number that looks like a catastrophe " + + "and means nothing. The families reset independently, so the untouched ones below are " + + "still accurate." + : "Change across the window, not the counters' cumulative levels. Requested checkpoints " + + "climbing against timed ones is the max_wal_size-too-small signal; full-page images " + + "spiking right after each checkpoint points at checkpoint_timeout instead. A blank " + + "value is a metric this PostgreSQL version does not expose, which is not zero."; + } + + /// Replication — slot WAL retention and the xmin each slot pins. + private async Task LoadPgReplicationAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var rows = await _dataService.GetPgSlotsAsync(_server.ServerId, startUtc, endUtc); + + PgReplicationSlotsGrid.ItemsSource = rows.Select(PgDisplay.Slot).ToList(); + PgReplicationNote.Text = PanelNote("pg_replication_slots", rows.Count, + "This server has no replication slots, so nothing is retaining WAL or pinning an xmin on their " + + "account. Zero rows is the healthy answer here, not a missing read."); + + await LoadPgReplicationStatsAsync(startUtc, endUtc); + } + + /// + /// Connected standbys (#2544), beneath the slots grid — the other half of "is replication healthy". + /// + /// Zero rows on a REPLICA is correct, not a fault. pg_stat_replication is the + /// primary-side view, so a standby reports nothing unless it is cascading to a downstream of its own. + /// The note says which of those it is looking at rather than reporting an absence of replication. + /// + /// The note leads with the WORST distance reached, not the current one: a replica that drifts + /// hundreds of megabytes behind every afternoon and recovers by evening reads as perfectly healthy in + /// every single sample, and it is the one most likely to be useless at the moment somebody needs to fail + /// over to it. + /// + private async Task LoadPgReplicationStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_replication_stats")) + { + PgReplicationStatsGrid.ItemsSource = null; + PgReplicationStatsNote.Text = PanelNote("pg_replication_stats", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgReplicationStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgReplicationStatsGrid.ItemsSource = rows; + + var worst = rows.Count == 0 ? 0L : rows.Max(r => r.WorstReplayBytesBehind ?? 0L); + var flapping = rows.Count(r => r.TotalSamples > 0 && r.Samples < r.TotalSamples); + + PgReplicationStatsNote.Text = rows.Count == 0 + ? "No standby was streaming from this server in this window. On a REPLICA that is the expected " + + "answer — pg_stat_replication is the primary-side view and reports nothing unless this " + + "server is cascading to a downstream of its own. On a primary it means nothing is " + + "replicating from it, which is either correct or the finding." + : $"{rows.Count:N0} standby connection(s). The worst any of them fell behind in this window was " + + $"{worst:N0} bytes of unapplied WAL — that column, not the current one, is what catches a " + + "replica that drifts far behind and recovers before anybody looks. Rank on BYTES rather than " + + "the lag columns: measured against a stalled standby, the time lag read 2.8 seconds for a " + + "33.7 MB backlog, because it times the round trip of the last replayed record rather than " + + "sizing the backlog." + + (flapping > 0 + ? $" {flapping:N0} standby(s) appeared in fewer samples than were taken, which means they " + + "have been DISCONNECTING — every other column shows that as healthy." + : string.Empty); + } + + /// + /// Storage - the per-table bloat estimate and per-index usage. Both reads fire together: they are one + /// tab answering one question (where the space went, and whether it is earning its keep), so moving + /// between the two grids needs no second round trip. + /// + private async Task LoadPgStorageAsync() + { + var (startUtc, endUtc) = GetWindowUtc(); + + var bloatTask = _dataService.GetPgTableBloatAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + var indexTask = _dataService.GetPgIndexUsageAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + await Task.WhenAll(bloatTask, indexTask); + + var bloat = bloatTask.Result; + var indexes = indexTask.Result; + + var bloatRows = bloat.Select(PgDisplay.TableBloat).ToList(); + PgTableBloatGrid.ItemsSource = bloatRows; + + /* The suppression count is stated ON THE PANEL rather than left to the per-row Confidence column, + because the usual cause is a single instance-wide permissions gap: every row is suppressed for + the same reason, and saying it once above the grid is what gets it fixed. Counted from the + PROJECTED rows so the sentence and the grid cannot disagree about which rows were suppressed. */ + var suppressed = bloatRows.Count(r => r.EstimateSuppressed); + + PgTableBloatNote.Text = PgCollectorIsGatedOff("pg_table_bloat_stats") + ? PanelNote("pg_table_bloat_stats", 0, string.Empty) + : bloatRows.Count == 0 + ? "No table of at least 1 MB was measured on this server in this window. Below that floor " + + "bloat is not an actionable amount of space, so zero rows here is the healthy answer " + + "rather than a missing read." + : suppressed == 0 + ? "Bloat figures are ESTIMATES computed from column-width statistics - the table itself " + + "is never read. Confirm one with pgstattuple before rewriting anything." + : $"Bloat figures are ESTIMATES computed from column-width statistics. {suppressed} of " + + $"{bloatRows.Count} row(s) have NO publishable estimate and show a dash rather than " + + "a number - see the Confidence column. If most of them say 'no column statistics', " + + "the monitoring login cannot SELECT these tables: pg_stats is filtered by SELECT " + + "privilege and pg_monitor does not grant it, so granting pg_read_all_data " + + "(PostgreSQL 14+) fixes the whole instance at once."; + + PgIndexUsageGrid.ItemsSource = indexes.Select(PgDisplay.IndexUsage).ToList(); + PgIndexUsageNote.Text = PgCollectorIsGatedOff("pg_index_usage_stats") + ? PanelNote("pg_index_usage_stats", 0, string.Empty) + : indexes.Count == 0 + ? "No index of at least 64 KB was recorded on this server in this window. Below that floor " + + "an index costs effectively nothing to keep, so zero rows here is the healthy answer " + + "rather than a missing read." + : "Scans are cumulative since each database's statistics were last reset. An index with no " + + "scans is a CANDIDATE, never a conclusion: check the Can It Go? column, and widen the " + + "window past the slowest scheduled job you have before acting - a monthly report looks " + + "exactly like a dead index over seven days."; + + await LoadPgColumnStatsAsync(startUtc, endUtc); + await LoadPgIndexBloatAsync(startUtc, endUtc); + } + + /// + /// Measured index bloat (#2561), under index usage — the two are halves of one question. + /// + /// The note leads with RECLAIMABLE BYTES rather than a worst-density figure, because density + /// alone ranks the wrong thing: a tiny index at 20% looks alarming and is worth kilobytes. It also has + /// to say that a healthy index measures near 90 rather than 100, or the first person to read the density + /// column concludes every index in the fleet is 10% bloated. + /// + private async Task LoadPgIndexBloatAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_index_bloat")) + { + PgIndexBloatGrid.ItemsSource = null; + PgIndexBloatNote.Text = PanelNote("pg_index_bloat", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgIndexBloatAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgIndexBloatGrid.ItemsSource = rows; + + var reclaimable = 0L; + foreach (var r in rows) + { + reclaimable += r.EstimatedReclaimableBytes ?? 0L; + } + + var skipped = rows.Count(r => r.SkippedReason is not null); + + PgIndexBloatNote.Text = rows.Count == 0 + ? "Nothing recorded. This panel needs the pgstattuple extension — without it the collector " + + "reports the function as missing rather than failing, and the Overview tab's extension panel " + + "says whether it is available on this server and one CREATE EXTENSION away. Only B-TREE " + + "indexes are measured; pgstatindex raises on GIN, BRIN and hash." + : $"MEASURED, not estimated from column statistics — every page of each index was read. About " + + $"{reclaimable:N0} bytes look reclaimable across {rows.Count:N0} index(es), and that is what " + + "the grid is ranked by: density alone ranks the wrong thing, since a tiny index at 20% is " + + "worth kilobytes next to a large one at 70%. **Leaf density is the server's raw figure and " + + "is not 100-minus-bloat** — a freshly built index measures around 90, so the reclaimable " + + "estimate is computed against that floor rather than against a full page." + + (skipped > 0 + ? $" {skipped:N0} index(es) were too large to read and are listed FIRST with their reason: " + + "their bloat is unknown rather than zero, and they are the likeliest big win." + : string.Empty); + } + + /// + /// The column-statistics panel (#2543) — the planner inputs that explain WHY a plan was chosen, and the + /// same statistics the bloat estimate above is computed from. + /// + /// Zero rows has two causes and the note must not collapse them. pg_stats filters on + /// has_column_privilege, so a monitoring login without SELECT on a table sees nothing for it — + /// measured: a pg_monitor-only role gets zero rows where a superuser gets all of them. Row-level + /// security empties it the same way. Neither is an absence of problems, and reporting "no statistics" as + /// though the data were clean is the exact claim the miss vocabulary exists to prevent. + /// + private async Task LoadPgColumnStatsAsync(DateTime startUtc, DateTime endUtc) + { + if (PgCollectorIsGatedOff("pg_column_stats")) + { + PgColumnStatsGrid.ItemsSource = null; + PgColumnStatsNote.Text = PanelNote("pg_column_stats", 0, string.Empty); + return; + } + + var rows = await _dataService.GetPgColumnStatsAsync(_server.ServerId, startUtc, endUtc, PgGridRowLimit); + + PgColumnStatsGrid.ItemsSource = rows; + + var skewed = rows.Count(r => r.TopValueFrequency >= 0.25); + + PgColumnStatsNote.Text = rows.Count == 0 + ? "No column statistics were collected. That is NOT the same as clean statistics, and it has two " + + "causes worth telling apart: pg_stats is filtered by SELECT privilege, so a monitoring login " + + "without it on a table sees nothing for that table (row-level security empties the view the " + + "same way) — or the server genuinely has no table above the 1 MB floor this collects at." + : $"Ranked by suspicion, not alphabetically. {skewed:N0} column(s) have a single value covering " + + "a quarter or more of the table, which is the PostgreSQL analogue of parameter sniffing: a " + + "plan that suits most values is catastrophic for that one. Low correlation on a wide column " + + "is the other shape, and it is why an index scan was rejected on a column that obviously " + + "has an index. Distinct is NEGATIVE when it is a ratio of row count — -1 means nearly every " + + "row is unique, not minus one value. Most-common VALUES and histogram bounds are " + + "deliberately not collected: they hold raw column data."; + } +} diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml index d3afe94a4..5a6927626 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml @@ -1,4 +1,4 @@ - + + + + + + + + + + + + + + + + + + + + + @@ -409,6 +435,15 @@ + + + + @@ -789,6 +824,7 @@ VerticalAlignment="Center"/> + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml.cs index dd3155453..db01f6462 100644 --- a/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml.cs +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerServerTab.xaml.cs @@ -67,9 +67,27 @@ resource run (Latches & Spinlocks was formerly at position 3, System Events dead /* #1496 Long Queries — the opt-in completion trace surface, appended after the diagnostics tail so no existing inner-tab index shifts (drill-down navigation keys on these constants). MUST remain the last - in ViewerServerTab.xaml's InnerTabs. */ + SQL SERVER in ViewerServerTab.xaml's InnerTabs. */ internal const int LongQueriesInnerTabIndex = 18; + /* ── The PostgreSQL run (#2530) ─────────────────────────────────────────────────────────────── + Indices 19-24 CONTINUE the same TabControl rather than living in a second one, and every index + above stays exactly where it was. That is the whole reason for the layout: for a PostgreSQL server + the nineteen SQL Server TabItems are collapsed and these six are shown, so the two sets never + renumber each other and no drill-down that keys on a constant above can be broken by this feature. + + ViewerPostgresTabs is the registry these belong to — headers, panel-to-collector mapping and the + framing notes live there, and ViewerPostgresTabsTests holds it to CollectorCatalog in both + directions so a tenth PostgreSQL collector cannot ship without a screen. Keep these constants, the + registry's InnerTabIndex values and the order in the XAML in step; two pins check it. */ + internal const int PgOverviewInnerTabIndex = 19; + internal const int PgActivityInnerTabIndex = 20; + internal const int PgVacuumInnerTabIndex = 21; + internal const int PgWaitsInnerTabIndex = 22; + internal const int PgIoInnerTabIndex = 23; + internal const int PgReplicationInnerTabIndex = 24; + internal const int PgStorageInnerTabIndex = 25; + private readonly ViewerDataService _dataService; private readonly DarlingServer _server; private readonly ViewerServerStore? _serverStore; @@ -137,6 +155,13 @@ chart chrome applies without any per-control theme plumbing. */ /* CPU Scheduler sub-tab chart (cpu_scheduler_stats parity): pressure trend + latest-snapshot grid. */ InitializeCpuSchedulerChart(); + /* #2530: a PostgreSQL target gets the PostgreSQL tab set. Done here, once, from the server record + the sidebar already loaded — not from an async probe — because the tab strip must be right on the + FIRST paint: the web half's review found that moving this decision into an async callback broke + the page's render model twice over, and the viewer has an even shorter path to get it right + since the engine kind arrives with the server row itself. */ + ApplyEngineTabSet(); + /* Memory inner-tab charts (copied from Lite): same up-front theme + hover for the five Memory charts (Overview trend, Clerks, Grant sizing/activity, Pressure events). */ InitializeMemoryCharts(); @@ -295,6 +320,14 @@ visible tab renders (the viewer's visible-only rule), so its offset wins. Loaded /// private async Task UpdatePermissionDeniedBadgeAsync() { + /* #2530: Collection Health is a SQL Server tab and is collapsed for a PostgreSQL server, which + gets the same answer on its own Overview tab instead. Writing a header nobody can see would be + harmless; issuing the read every refresh to do it would not. */ + if (_server.IsPostgres) + { + return; + } + try { var denied = await _dataService.GetPermissionDeniedCollectorCountAsync(_server.ServerId); @@ -392,6 +425,34 @@ not loaded on tab-switch. Explicit no-op so selecting it doesn't fall through to case LongQueriesInnerTabIndex: await LoadLongQueriesAsync(); break; + + /* #2530: the PostgreSQL run. Explicit arms, never the default: falling through to + LoadOverviewChartsAsync would point the SQL Server correlated-lane reads (CPU %, wait + ms/sec, buffer pool, file I/O latency — four collectors that cannot run on this engine) + at a PostgreSQL server and paint four permanently empty lanes, which is the exact defect + this issue is about. */ + case PgOverviewInnerTabIndex: + await LoadPgOverviewAsync(); + break; + case PgActivityInnerTabIndex: + await LoadPgActivityAsync(); + break; + case PgVacuumInnerTabIndex: + await LoadPgVacuumAsync(); + break; + case PgWaitsInnerTabIndex: + await LoadPgWaitsAsync(); + break; + case PgIoInnerTabIndex: + await LoadPgIoAsync(); + break; + case PgReplicationInnerTabIndex: + await LoadPgReplicationAsync(); + break; + case PgStorageInnerTabIndex: + await LoadPgStorageAsync(); + break; + case OverviewInnerTabIndex: default: await LoadOverviewChartsAsync(); diff --git a/Darling/PerformanceMonitor.Darling.Viewer/ViewerSettingsFile.cs b/Darling/PerformanceMonitor.Darling.Viewer/ViewerSettingsFile.cs new file mode 100644 index 000000000..78c8258a4 --- /dev/null +++ b/Darling/PerformanceMonitor.Darling.Viewer/ViewerSettingsFile.cs @@ -0,0 +1,160 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Text.Json; +using PerformanceMonitor.Common; + +namespace PerformanceMonitor.Darling.Viewer; + +/// +/// The one read and the one write behind both of the viewer's per-user JSON settings files — +/// (viewer-settings.json) and +/// (viewer-preferences.json) — sitting on the shared (#2434). +/// +/// Both stores previously ended their load in a bare catch whose only output was a +/// trace. Debug.WriteLine carries +/// ("DEBUG") and is removed by the compiler from a +/// Release build, so in the viewer anyone actually runs, a settings file that could not be read produced +/// no record at all — not a log line, not a dialog, nothing. And neither Save merged: each serialized its +/// whole in-memory object over the file, so a load that had silently fallen back to defaults, followed by +/// any save at all, replaced every setting in the file with a default. Changing the time-display dropdown +/// once was enough. +/// +/// So there are two rules here, and the second is the one that turns an annoyance into data loss. +/// A file that is present and unreadable is always reported, to the log the viewer now has +/// ( — the "the viewer writes no application log of its own" comment these +/// stores carried predates it). And nothing replaces a file it could not read until a copy of it exists: +/// when even the copy cannot be made, the save is refused and says so, because leaving the file alone +/// beats replacing it when the alternative is permanent. +/// +/// An ABSENT file goes through both paths in total silence. A first run has nothing to preserve and +/// nothing to explain, and a warning there would be pure noise — keeping absent apart from unreadable is +/// half the reason the guard exists. +/// +/// Generic in the value rather than written per store, because there are three of these files and +/// the third — 's registry of monitored servers — is the one where losing +/// the file costs the operator the most and is a JSON ARRAY rather than an object. Everything the guard +/// decides is decided by the same code for all three. +/// +internal static class ViewerSettingsFile +{ + /// + /// The (file, problem) pairs already reported this session, so one broken file is one log line rather + /// than one per read. + /// + /// Nothing reads these files once. The theme is applied from viewer-settings.json before the + /// window exists, MainWindow loads it again to seed itself, the Settings window loads it on open and + /// MainWindow re-loads it on close, and the control-plane migration loads it too — so an unreported + /// duplicate would put five identical ERROR lines in the log for one defect, which is a poor reward for + /// whoever went looking. The first line is the one that carries the fact; the rest carry nothing. + /// + /// Keyed on the problem as well as the path deliberately: if the file changes mid-session and + /// breaks in a NEW way, that is a new fact and it is said again. + /// + private static readonly HashSet s_reported = new(StringComparer.Ordinal); + + private static bool FirstReportOf(string filePath, string? problem) + { + lock (s_reported) + { + return s_reported.Add($"{filePath}|{problem}"); + } + } + + /// + /// Reads into , substituting a default instance + /// when there is nothing usable to read. The returned State is what a caller needs to decide + /// whether the defaults it just got are a legitimate first run or a configuration that is still on + /// disk and could not be understood; the log line for the second case is written here, so a call site + /// that never looks at the state is still not silent. + /// + internal static SettingsObjectRead Load(string filePath, string logSource, JsonSerializerOptions options) + where T : class, new() + { + var read = SettingsFileGuard.ReadObject(filePath, options); + + if (read.State == SettingsFileState.Unreadable && FirstReportOf(filePath, read.Problem)) + { + /* Two different facts, so two different sentences. A file the reader could not use AT ALL costs + every setting in it; a file it read after dropping named members costs only those. Saying the + first when the second is true is the overstatement #2456 was filed to end — and saying the + second when the first is true would be worse, because it implies the rest survived. */ + ViewerLogger.Error(logSource, + read.UnreadableMembers is { Count: > 0 } + ? $"'{filePath}': {read.Problem}. Every other setting in the file loaded normally. " + + "The file has not been changed; the next save copies it aside before replacing it." + : $"'{filePath}' could not be read ({read.Problem}), so every setting it holds is at " + + "its default for this session. The file has not been changed; the next save copies " + + "it aside before replacing it."); + } + + return read.Value is null ? read with { Value = new T() } : read; + } + + /// + /// Serializes over , and returns whether it + /// actually reached disk. + /// + /// The bool is the point. This is a whole-object replacement, so a save that fails silently and + /// a save that worked are indistinguishable to a caller that cannot ask — which is how a UI ends up + /// saying it saved something it did not. Every failure here is logged AND reported, and the two call + /// sites that persist a setting from an ordinary click surface it to the user rather than dropping + /// it into a log nobody reads after a dialog said it worked. + /// + /// It returns false rather than throwing because the handlers that call it — a dropdown + /// selection change, a sidebar group collapsing — do not wrap it, and a full disk is not a reason to + /// take the viewer down. + /// + internal static bool Save(string filePath, T value, string logSource, JsonSerializerOptions options) + where T : class + { + var permit = SettingsFileGuard.PermitReplace(filePath, DateTime.Now, options); + + if (!permit.Allowed) + { + ViewerLogger.Error(logSource, + $"'{filePath}' could not be read ({permit.Problem}) and no copy of it could be made, so " + + "it has been left untouched rather than overwritten with defaults, and nothing was saved. " + + "Fix the file, or move it aside by hand, and try again."); + return false; + } + + if (permit.QuarantinedTo is not null) + { + /* Deliberately does not claim WHAT the replacement is written from. Since #2456 that depends on + how much of the file the load recovered — everything but the named members, or nothing at all + — and this permit is asked at save time by a caller holding an object it did not necessarily + load. The one fact that matters here is true either way: the original bytes are in the copy. */ + ViewerLogger.Warn(logSource, + $"'{Path.GetFileName(filePath)}' could not be read ({permit.Problem}), so this save " + + "replaces it. The unreadable original was copied to " + + $"'{Path.GetFileName(permit.QuarantinedTo)}' first — the settings it held are recoverable " + + "from there."); + } + + try + { + var directory = Path.GetDirectoryName(filePath); + if (!string.IsNullOrEmpty(directory)) + { + Directory.CreateDirectory(directory); + } + + File.WriteAllText(filePath, JsonSerializer.Serialize(value, options)); + return true; + } + catch (Exception ex) + { + ViewerLogger.Error(logSource, $"'{filePath}' could not be written, so nothing was saved", ex); + return false; + } + } +} diff --git a/Darling/README.md b/Darling/README.md index 97ad4e15e..2ed644e99 100644 --- a/Darling/README.md +++ b/Darling/README.md @@ -1,1174 +1,1298 @@ -# Performance Monitor Darling — Headless Edition - -Darling is the headless, centralized edition of Performance Monitor: a 24/7 Windows service that collects from your SQL Servers into a central PostgreSQL (optionally TimescaleDB) store, plus a detached desktop viewer that reads that store. No desktop app has to stay open for collection to happen, and every viewer seat reads the same central data. - -It runs the **same monitoring brain as the Lite edition** — one shared codebase, two storage engines: - -- `PerformanceMonitor.Collectors` owns all 48 collector definitions — 41 for SQL Server and 7 for PostgreSQL: the exact query sent to monitored servers, the result-row mappings, the delta rules, the default cadences and retention horizons, and the ignored-wait-types list. Lite writes those rows to DuckDB; Darling writes the same rows to PostgreSQL via binary COPY. Each definition declares which engine it targets, and a collector never runs against the other one — see [PostgreSQL targets](#postgresql-targets). -- `PerformanceMonitor.Alerting` owns the shared alert engine — the same thresholds, edge-trigger gates, cooldowns, and dedup fingerprints Lite uses. -- The analysis/recommendations pipeline (the same inference engine behind both apps' Recommendations tabs and the `analyze_server` MCP tool) runs on a schedule inside the service. - -A collector, alert, or analysis change lands once in the shared libraries and both editions get it. A Darling install monitoring a server even derives the **same `server_id`** Lite would for that server, because the identity rule (`host[:database][:RO]`, hashed) is shared too. - -> **Status: in development.** Darling builds and runs from source (it is wired into the solution and CI), but is not yet packaged into the signed release artifacts. Expect the surface documented here to grow. - ---- - -## When to Choose Darling vs. Lite - -| | **Lite** | **Darling** | -|---|---|---| -| Collection runs | While the desktop app is open (or in the tray) | 24/7 as a Windows service | -| Data lives | Locally per seat (DuckDB + Parquet) | Centrally (PostgreSQL / TimescaleDB) | -| Execution plans | Not stored (fetched live when you view a query) | Captured and stored, TOAST-compressed (`capturePlans`, default on) | -| Viewers | The app is the viewer | Any number of viewer seats read the central store | -| Setup | Download and run | Provision PostgreSQL, edit `darling.json`, install the service | -| Best for | Quick triage, consultants, a handful of servers | Always-on team monitoring, larger estates, one shared store | -| Configuration | Settings UI | One JSON file (no UI) | - -Nothing is installed on the monitored SQL Servers by either edition beyond two lightweight Extended Events ring-buffer sessions and, when it is unset, a one-time `blocked process threshold` bootstrap (see [What the Service Does on Monitored Servers](#what-the-service-does-on-monitored-servers)). - ---- - -## Quick Start - -### Prerequisites - -- **Windows** for the service host (Windows-service lifetime, DPAPI password protection) and for the viewer (WPF). Monitored servers can be SQL Server 2016–2025, Azure SQL Managed Instance, AWS RDS for SQL Server, or Azure SQL Database. -- **A PostgreSQL store — bundled or your own.** In managed mode (the shipped default, see [Managed Bundled PostgreSQL](#managed-bundled-postgresql)) the service runs its own bundled PostgreSQL 18 + TimescaleDB and no database provisioning is needed. To bring your own instead, PostgreSQL 16 or newer is recommended (developed and validated against PostgreSQL 18) with a database and a login the service can create tables in — and if that store has TimescaleDB, size its background workers before you rely on compression, because the stock PostgreSQL defaults cannot run the policies (see [Background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont)). -- **TimescaleDB is optional and auto-adopted.** If the extension is installed (or pre-created by an administrator) in the store database, the service detects it at startup and automatically converts the collector tables to hypertables with compression; without it, the service runs in plain-PostgreSQL mode, which is fully supported. No configuration flag either way. -- **.NET 10** to build and run. - -Build from the repository root: - -``` -dotnet build Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -``` - -``` -dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release -``` - -### Configure darling.json - -The service reads one JSON file. It resolves the path in this order: - -1. An explicit path (when a component is handed one) -2. The `DARLING_CONFIG` environment variable -3. `darling.json` next to the service binary - -Copy the shipped `darling.sample.json` (it lands next to the built binary) to `darling.json` and edit. Comments and trailing commas are allowed; property names are case-insensitive. - -Minimal working example — one server, integrated auth, bring-your-own PostgreSQL. (With the bundled store instead, replace the `postgres` block with `"postgres": { "managed": true }` and skip provisioning entirely — see [Managed Bundled PostgreSQL](#managed-bundled-postgresql).) - -```json -{ - "postgres": { - "connectionString": "Host=localhost;Port=5432;Username=darling;Database=darling" - }, - "servers": [ - { - "name": "SQL2022", - "host": "SQL2022", - "auth": "integrated", - "excludedDatabases": [] - } - ] -} -``` - -**Integrated auth (recommended).** The service connects to monitored servers as the Windows account the service runs under — there is no separate Windows credential to configure. Grant that account the [permissions below](#permissions-on-monitored-servers). The default install's virtual service account reaches *remote* servers as the collector machine's computer account (`DOMAIN\$`), so for integrated auth you will usually [run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) instead. - -**SQL auth.** Set `"auth": "sql"`, a `username`, and an `encryptedPassword` produced by the `--encrypt-password` verb: - -``` -PerformanceMonitor.Darling.Service.exe --encrypt-password -``` - -It prompts for the password on stdin (so the plaintext never lands in your shell history) and prints a base64 DPAPI blob. Paste that blob into the server's `"encryptedPassword"`. The blob is protected with **DPAPI LocalMachine scope**, so an administrator can encrypt it interactively and the service account can decrypt it later on the same machine — but it is machine-bound: run `--encrypt-password` **on the machine that will run the service**, and re-encrypt if you move `darling.json` to another machine. A plaintext `"password"` also works as a dev convenience, but the service logs a warning every time it is used. The same slot also takes an **`env:NAME` or `file:/path` reference** (#1804): the service reads the named environment variable or the file's (trimmed) contents at connect time, nothing secret lands in `darling.json`, and no warning is logged — the supported shape on non-Windows hosts, and compose-`secrets:`-friendly everywhere. A missing or empty reference target is a configuration error naming both the setting and the target, never a silent empty password. - -**excludedDatabases** (per server) removes databases from collection: per-database collectors skip them and the exclusion is spliced into the collector queries — the same filter Lite applies. There is a second, separate `alerts.excludedDatabases` list that excludes databases from blocking/deadlock/long-running-query **alert evaluation** without affecting collection. - -### Validate the Config (Pre-flight) - -Before installing the service, check that `darling.json` is well-formed and that every monitored server is reachable with the configured credentials: - -``` -PerformanceMonitor.Darling.Service.exe --test-connection -``` - -(`--validate-config` is an alias.) It validates the file, then connects to and probes each server, printing a `[PASS]`/`[FAIL]` line per server (SQL major version, engine edition, and whether the account has msdb access for failed-job alerts). It exits `0` only when the file is valid **and** every server is reachable, so it doubles as a deployment gate. - -A PostgreSQL target reports what matters there instead — version, writer or reader, Aurora or not, and **how many of the PostgreSQL collectors will actually run against it**, naming the ones that will not: - -``` - [PASS] aurora-writer: PostgreSQL 17 (server_version_num 170007), writer, Aurora — all 8 PostgreSQL collectors apply - [PASS] aurora-reader: PostgreSQL 17 (server_version_num 170007), reader (in recovery), Aurora — 7 of 8 PostgreSQL collectors apply (skipped: pg_autovacuum_stats) - [PASS] selfhosted: PostgreSQL 15 (server_version_num 150012), reader (in recovery), not Aurora — 4 of 8 PostgreSQL collectors apply (skipped: pg_autovacuum_stats, pg_io_stats, pg_statement_stats, pg_wait_stats) -``` - -That count comes from the same [engine and version gate](#postgresql-targets) the collector runner uses, not a separate list, so it is the real answer rather than an estimate — and it is the answer at *pre-flight*, before an empty table has to be explained weeks later. Add an explicit config path as a second argument if `darling.json` is not next to the exe and `DARLING_CONFIG` is not set. This is the same probe the Viewer's **Test Connection** button runs through the service. - -One identity caveat: the verb connects as **you**, the console user — not as the service account. For `"auth": "integrated"` servers a `[PASS]` proves the server is reachable and the config is well-formed, but the grants that matter at runtime are the *service account's*: the per-server connect lines in the service log are the real proof (see [Run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa)). - -### Run It — Console Mode - -The same executable serves interactive debugging and service installation; the Windows-service lifetime is a no-op when run from a console. - -``` -Darling\PerformanceMonitor.Darling.Service\bin\Release\net10.0\PerformanceMonitor.Darling.Service.exe -``` - -Watch the log output: you should see the config load (`Loaded configuration from ...`), the store migrate (`Postgres store ready (schema v44, ...)` — the number is whatever the current migration count is), the TimescaleDB detection result, per-server connects, and then per-collector run lines with row counts. - -### Run on Linux (Docker Compose or systemd) {#1804} - -The service is cross-platform .NET; only the **bundled zero-admin store** and DPAPI are Windows-specific. On Linux you pair the service with the official TimescaleDB image (compose, the recommended shape) or point it at PostgreSQL you already run (systemd), keeping `postgres.managed = false` either way. The Viewer stays a Windows desktop app — Linux hosts read the **web dashboard**, which the container exposes. - -**Compose (the whole stack as one deployment)** — everything lives in [`Darling/compose/`](compose/): - -```bash -cd Darling/compose -cp darling.sample.json darling.json # edit: servers, alerting, tokens -# one secret per file — see secrets/README.md for the exact list -docker compose up -d -``` - -Web dashboard on `http://:5153` behind its token, MCP (if enabled) on `:5152` behind its bearer token. The port mappings are the exposure boundary: the container-aware bind gate honors `web.network`/`mcp.network` under `managed = false` **inside a container only**, and the tokens are still mandatory. Three rules worth knowing before they bite: - -- **Nothing secret goes in darling.json.** Every secret slot — the whole `postgres.connectionString`, server `password`s, `smtp.password`, the tokens — takes an `env:NAME` or `file:/run/secrets/` reference. The compose file mounts each secret from `secrets/`. -- **Start with a fresh store volume per deployment.** The control plane is store-authoritative after the first seed, so a reused volume's enable toggles override darling.json — by design. -- **File permissions are yours on Linux.** The Windows build locks config/credentials down with ACLs; here the container boundary is the isolation, and the `secrets/` directory should be `chmod 700` with `600` files (the systemd shape should do the same for `darling.json` itself). - -**systemd + bring-your-own PostgreSQL** — download `PerformanceMonitorDarling-linux-x64-*.tar.gz` from the release, extract to `/opt/darling`, point `DARLING_CONFIG` at your config (connection string to your own PostgreSQL 15+ with TimescaleDB; the service degrades gracefully without TimescaleDB), and run `dotnet PerformanceMonitor.Darling.Service.dll` under a unit like: - -```ini -[Unit] -Description=PerformanceMonitor Darling -After=network-online.target - -[Service] -ExecStart=/usr/bin/dotnet /opt/darling/PerformanceMonitor.Darling.Service.dll -Environment=DARLING_CONFIG=/etc/darling/darling.json -User=darling -Restart=on-failure - -[Install] -WantedBy=multi-user.target -``` - -Use the same `env:`/`file:` secret references (systemd `LoadCredential=` pairs naturally with `file:`), and note `Microsoft.Data.SqlClient` needs `libgssapi-krb5-2` installed (`apt-get install libgssapi-krb5-2`) — the container image carries it already. - -### Install as a Windows Service - -**Scripted (recommended):** the packaged zips ship `install-darling.ps1` beside the service exe. Extract the zip to its final location (e.g. `C:\PerformanceMonitorDarling`), then from an elevated PowerShell in that folder run `.\install-darling.ps1`. It checks the install location and refuses anywhere the service could not read itself (see [below](#the-install-location-has-to-be-machine-scoped)), checks for `darling.json` (copying the sample and stopping for you to edit it on first run), runs the `--test-connection` pre-flight, registers the Event Log source, creates the service under the virtual account (or upgrades an existing install's binPath in place, preserving config/store/credentials), starts it, and creates Desktop + Start Menu **Darling Viewer** shortcuts (pin to taskbar from the Start Menu entry — Windows does not allow programmatic pinning). `uninstall-darling.ps1` reverses it, deliberately leaving the store/config in place unless you pass `-PurgeData`. - -#### The install location has to be machine-scoped - -Extract to a local, machine-scoped path — `C:\PerformanceMonitorDarling` is the documented one. **Not** anywhere under a user profile (`C:\Users\...`, including your Desktop or Downloads), and not a UNC path or a mapped drive. - -The service runs as the unprivileged virtual account `NT SERVICE\PerformanceMonitor Darling`, never LocalSystem, because the bundled PostgreSQL refuses to run with administrative privileges. That account is not you, not SYSTEM, and not Administrators — and a user profile grants access to about those three and nobody else, so the service cannot read its own program files there. It installs cleanly and then fails: `initdb.exe` dies at `0xC0000135` (STATUS_DLL_NOT_FOUND) before it can report anything (#2185). A folder created under `C:\` inherits read + execute for `BUILTIN\Users` instead, which the virtual account is a member of, which is why the documented location works. Network paths fail for a related reason: a virtual account [reaches the network as the computer account](https://learn.microsoft.com/en-us/sql/database-engine/configure-windows/configure-windows-service-accounts-and-permissions#virtual-accounts) rather than as you, and a mapped drive letter belongs to your logon session, which a service does not share. - -`install-darling.ps1` refuses a fresh install in any of these locations rather than leaving you a service that cannot start. A service registered by hand instead — the manual `sc create` path below, which the installer never sees — gets the same diagnosis from the service itself: on start, ahead of reading `darling.json` and long before the store bootstrap, it logs one critical line naming the path, why its own account cannot read it, and where to move it, so the cause is above the failure rather than three messages downstream of it. To move an existing install, stop the service, move the folder, and re-run `install-darling.ps1` from the new location — it updates the service's binPath in place and leaves your `darling.json`, store data, and credentials alone. - -**Manual:** publish (or copy the build output) to a stable path, put `darling.json` next to the exe (or set `DARLING_CONFIG` as a machine environment variable), then register it: - -``` -dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o C:\PerformanceMonitorDarling -``` - -``` -sc create "PerformanceMonitor Darling" binPath= "C:\PerformanceMonitorDarling\PerformanceMonitor.Darling.Service.exe" start= auto obj= "NT SERVICE\PerformanceMonitor Darling" -``` - -``` -sc start "PerformanceMonitor Darling" -``` - -Also register the service's Windows event source once, from the same elevated shell — event-source registration requires elevation, and the virtual service account cannot do it itself (without this, Event Log diagnostics are silently dropped; the file log under `%ProgramData%\PerformanceMonitorDarling\logs` works regardless): - -``` -powershell -NoProfile -Command "New-EventLog -LogName Application -Source 'PerformanceMonitor Darling' -ErrorAction SilentlyContinue" -``` - -The `obj=` clause runs the service under a **virtual service account** (`NT SERVICE\` — password-less, per-service SID, unprivileged; the same convention SQL Server itself uses). That is the right account for SQL-auth monitoring, and with `postgres.managed = true` it is more than a preference: PostgreSQL refuses to execute with administrative privileges, so don't run the service as LocalSystem — a least-privilege account keeps the bundled store's initdb/start path on ground PostgreSQL supports. For integrated auth to monitored servers, [run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) instead. Note the space after `binPath=`, `start=`, and `obj=` — `sc` requires it. - -One managed-mode handoff gotcha: if you test-drove the service from a console first, the bundled store's data directory belongs to *your* account, and the service account may not be able to write it. Point the service at a fresh `postgres.dataDirectory` (or delete the test directory) rather than fighting ACLs. - -#### Run the service as a domain account or gMSA - -With `"auth": "integrated"`, the monitoring identity **is** the service's Log On account — nothing in `darling.json` names a Windows account, and there is no separate credential to set. The default virtual account carries only the *machine* identity onto the network (remote servers see `DOMAIN\$`), so for integrated auth against remote servers you almost always want a real AD service account or, better, a gMSA. Switching is a Windows-side change plus a SQL-side grant, with one file-permission step in the middle that bites everyone who skips it: - -1. **Change the Log On account.** Stop the service, then **Services.msc → PerformanceMonitor Darling → Log On → This account** — that route also grants the account the *Log on as a service* right automatically. Or from an elevated prompt (with `sc config` you grant *Log on as a service* yourself, via secpol.msc or GPO): - - ``` - sc config "PerformanceMonitor Darling" obj= "DOMAIN\svc-account" password= "ThePassword" - ``` - - A gMSA works the same way with an empty password: `obj= "DOMAIN\gmsa-name$" password= ""`. Keep the account **out of the local Administrators group**: with `postgres.managed = true` the bundled PostgreSQL refuses to run with administrative privileges, exactly as it refuses LocalSystem. - -2. **Grant the account on every monitored server** — a Windows login holding the same [permissions below](#permissions-on-monitored-servers) (the `GRANT`s there apply to a Windows login unchanged): - - ```sql - USE [master]; - CREATE LOGIN [DOMAIN\svc-account] FROM WINDOWS; - ``` - -3. **Re-grant the service's own files — the step people miss.** The service deliberately locks its files down to SYSTEM, Administrators, and the account it was *running as*; the new account is on none of those ACLs, and the service will fail to read its config or write its store. One-time, from an elevated prompt, before starting the service: - - ``` - icacls "C:\ProgramData\PerformanceMonitorDarling" /grant "DOMAIN\svc-account:(OI)(CI)F" - icacls "C:\PerformanceMonitorDarling\darling.json" /grant "DOMAIN\svc-account:F" - ``` - - Adjust the second path to wherever `darling.json` sits beside the service exe; the first covers the logs and, in managed mode, the store's data directory. On its next start the service re-asserts the tight ACL itself — now including the new account — so this does not need repeating. - - In managed-store mode there is one more, and it needs **ownership**, not a grant: the store's superuser credential `pg-credential.dpapi` (beside the data directory, under `C:\ProgramData\PerformanceMonitorDarling` by default) is trusted only when *owned* by SYSTEM, Administrators, or the service account — an anti-pre-plant check — and `icacls /grant` changes permissions, never ownership, so after the switch the file is still owned by the *previous* service account and the service refuses it. Hand ownership to Administrators (trusted across any future account change, which is why not the new account itself) and grant the new account on the file directly — its ACL is protected and does **not** inherit the folder grant above: - - ``` - takeown /f "C:\ProgramData\PerformanceMonitorDarling\pg-credential.dpapi" /a - icacls "C:\ProgramData\PerformanceMonitorDarling\pg-credential.dpapi" /grant "DOMAIN\svc-account:F" - ``` - - The sibling role credentials (the admin/viewer/mcp `.dpapi` files) hit the same ownership check but self-heal — a role password can be re-asserted, a superuser's cannot — so expect one-time `discarding and regenerating` warnings on the first start, not faults. - -4. **Start the service and verify from its log** (`%ProgramData%\PerformanceMonitorDarling\logs`): the per-server connect lines are the proof that the *service account's* grants work. `--test-connection` from your console runs as you, not the service account — see the [pre-flight note above](#validate-the-config-pre-flight). - -Nothing else moves: anything encrypted with `--encrypt-password` (SQL-auth server passwords, SMTP) survives the account change, because those blobs are DPAPI **machine**-scope, not account-scope — and collected data is untouched. Later `install-darling.ps1` upgrades preserve a custom Log On account and harden `darling.json` for the account the service actually runs as. - -### What the Service Does on Monitored Servers - -On each successful connect, the service: - -1. **Probes the server** — one query against `sys.dm_os_sys_info` / `SERVERPROPERTY()` for version, engine edition (box / Managed Instance / Azure SQL DB), AWS RDS detection, and msdb access. It is the same detection query Lite runs, so both editions classify a server identically. -2. **Ensures two Extended Events ring-buffer sessions** (created if missing, started if stopped; ~4 MB ring buffer each, no files written on the server): - - `PerformanceMonitor_Deadlock` — `xml_deadlock_report`, server-scoped on on-prem/Managed Instance/RDS; `database_xml_deadlock_report`, database-scoped on Azure SQL Database. - - `PerformanceMonitor_BlockedProcess` — `blocked_process_report`, server-scoped (database-scoped on Azure SQL Database). -3. **Bootstraps the blocked-process threshold** — if `blocked process threshold (s)` is `0`, the service sets it to `5` via `sp_configure`. On AWS RDS `sp_configure` is unavailable; the attempt is tolerated and logged, and you set the threshold through an RDS Parameter Group instead (Azure SQL Database has a fixed 20-second threshold). -4. **Runs the on-connect config snapshots once** (`server_config`, `database_config`, `database_scoped_config`, `trace_flags`, `server_properties`), then runs all scheduled collectors on the shared default cadences. - -Every failure in steps 2–3 is tolerated and logged: the deadlock/blocked-process collectors simply read zero rows until the sessions exist (and blocked-process reports only start arriving once the threshold is set). Monitoring queries connect with a 15-second connect budget and an application name of `PerformanceMonitorDarling`; connection encryption fails closed to `Mandatory` when the configured mode is unrecognized. - -### Permissions on Monitored Servers - -Darling needs the **same target-server grants as Lite**, so the copy-paste block lives in one place for both: **[Permissions in the root README](../README.md#lite--darling-on-premises)** — `VIEW SERVER STATE`, `CONNECT ANY DATABASE`, `VIEW ANY DEFINITION`, `ALTER ANY EVENT SESSION`, and the optional `ALTER TRACE`, `ALTER SETTINGS`, and msdb job-table grants, verified live against SQL Server 2025 with a scratch login carrying exactly them ([#1823](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1823)). That block is authoritative; this section is the Darling-specific reading of it. Keeping one list instead of two is deliberate — a second copy is how the old one went stale. - -**The one Darling-specific line:** for `"auth": "integrated"` the grants go to the Windows account **the service runs as**, so use `CREATE LOGIN [DOMAIN\svc-account] FROM WINDOWS;` in place of the block's `CREATE LOGIN ... WITH PASSWORD`. Everything after it is unchanged. See [Run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) for which account that actually is — it is not the one you ran `--test-connection` as. - -What each grant buys you, and what breaks without it: - -| Grant | Why | If missing | -|---|---|---| -| `VIEW SERVER STATE` | All DMV collectors (wait stats, query stats, memory, CPU, file I/O, sessions, etc.) and the connect probe | Collection fails — this one is required | -| `ALTER ANY EVENT SESSION` | Create/start the two XE sessions | Logged; deadlock and blocked-process collectors read zero rows (an admin can pre-create the sessions instead) | -| `CONNECT ANY DATABASE` | The per-database collectors (`database_scoped_config`, `query_store_health`, `index_object_stats`, `database_size_stats`, `query_store_stats`) enter each database via `EXECUTE [db].sys.sp_executesql` | Databases the login cannot enter are skipped; without the grant that is every user database | -| `VIEW ANY DEFINITION` | Catalog-view row visibility everywhere: `sys.tables` / `sys.indexes` / `sys.objects` for the index and object collectors, `sys.dm_db_partition_stats`, and the AG catalog views (`sys.availability_groups`, `sys.availability_replicas`) | **Silently zero rows** — catalog views hide rows rather than erroring, so missing objects look exactly like empty databases, and a real AG cluster looks identical to a server with no AGs | -| `ALTER SETTINGS` | The `sp_configure` blocked-process-threshold bootstrap | Logged; set the threshold yourself (or via RDS Parameter Group) | -| `ALTER TRACE` | The `default_trace_events` collector — `sys.traces` / `fn_trace_gettable` accept nothing less | `PERMISSIONS` skip in collection health; the default-trace tab stays empty | -| msdb job-table `SELECT`s + `agent_datetime` `EXECUTE` | `running_jobs` / `job_history` / `agent_status` collectors and the failed/long-running-job alerts — all direct table reads; `SQLAgentReaderRole` alone leaves every one failing with error 229 | Skipped gracefully — logged as a permissions skip, alerts return no jobs | -| `DBCC TRACESTATUS` permission | `trace_flags` snapshot | Degrades to zero rows with a warning | - -The msdb grants live inside a system database SQL Server setup can rewrite — re-check them after a CU or version upgrade. - -**Azure SQL Database:** connect to the one database you monitor (set the server entry's `"database"`), using a contained user with `VIEW DATABASE STATE` and `VIEW DEFINITION`, matching the product's existing Azure guidance. The XE sessions are created database-scoped there (`ALTER ANY DATABASE EVENT SESSION`); SQL Agent collectors are skipped automatically. - -Collectors that hit a permission error (SQL errors 229/297/300, plus 8189 from `sys.traces`) log a `PERMISSIONS` row in `collection_log` and retry on their next scheduled run — one denied collector never stops the rest. - -#### Which collectors run on which platform - -Every collector declares its own applicability in code (`AppliesTo(CollectorTargetInfo)`), so this is not a hand-maintained list of 36 rows — the collectors fall into five groups, and a collector outside its supported platform is **skipped before it runs**, not failed and logged every cycle. - -| Runs on | Collectors | Gate | -|---|---|---| -| Everything | wait stats, CPU utilization, memory (stats/clerks/grants), file I/O, tempdb, latches, spinlocks, plan cache, session summary, plus blocking, deadlocks, blocked-process reports, DMV blocking snapshots, perfmon, query snapshots, procedure stats, index/object stats, long-query completions, database config/scoped-config/size, server properties, session stats, waiting tasks | no gate | -| On-prem, Managed Instance, RDS — **not** Azure SQL DB | CPU scheduler stats, default trace events, memory pressure events, server config, system health events, trace flags | `!IsAzureSqlDb` | -| On-prem and Managed Instance, needs msdb | job history | `!IsAzureSqlDb && HasMsdbAccess` | -| On-prem and Managed Instance, needs msdb — **not** RDS | agent status, running jobs | `!IsAzureSqlDb && !IsAwsRds && HasMsdbAccess` | -| SQL Server 2016+ (or any Azure flavour) | query stats, Query Store stats | `SqlMajorVersion >= 13 \|\| IsAzureSqlDb \|\| IsAzureManagedInstance` | - -Notes: - -- **Azure SQL DB** is the most restricted target: the six `!IsAzureSqlDb` collectors read server-scoped DMVs or on-disk artifacts that do not exist there, and the SQL Agent collectors have no Agent to read. Nothing about that is a permission problem, so it is not reported as one. -- **AWS RDS** blocks direct `msdb` job reads specifically; the rest of the SQL Agent surface is unaffected. -- **`HasMsdbAccess`** is probed per server at connect and is exactly `HAS_DBACCESS('msdb')` — *any* access to msdb, not a specific role or table grant. Losing msdb access later moves those collectors from running to skipped without an error storm. A login that can enter msdb but lacks `SELECT` on the job tables passes this probe and is caught one layer down as a `PERMISSIONS` skip instead. -- An unknown version (`SqlMajorVersion == 0`, i.e. detection has not completed yet) is treated as capable rather than skipped, so a collector is never silently dropped because a probe was slow. - -If a tab or column is empty and you expect data, check **Collection Health**: a collector skipped for platform reasons shows no runs at all, whereas one denied by permissions logs `PERMISSIONS` and is classified `NO_PERMISSIONS`. Those are different problems with different fixes — the first is expected on that platform, the second is a grant to add from the table above. - ---- - -## Configuration Reference - -All sections except `postgres` and `servers` are optional — omit a section (or any key) to get the defaults listed here. Defaults deliberately mirror a fresh Lite install. - -### postgres - -Two mutually exclusive modes — setting both `managed: true` and `connectionString` is a validation error: - -| Key | Default | Notes | -|---|---|---| -| `managed` | `false` | `true` runs the bundled PostgreSQL + TimescaleDB (Windows only; see [Managed Bundled PostgreSQL](#managed-bundled-postgresql)). The connection string is derived, never configured. | -| `port` | `5641` | Managed mode only: the loopback port the bundled server listens on. Deliberately uncommon so it coexists with any PostgreSQL (5432) already on the machine. | -| `dataDirectory` | *(null)* | Managed mode only: the cluster's data directory. `null` means `%ProgramData%\PerformanceMonitorDarling\pg`. | -| `connectAs` | `"admin"` | Managed mode only: which least-privilege role the Viewer connects as — `"admin"` (reads everything + manages mute rules and dismisses alerts) or `"viewer"` (read-only; those write actions are hidden/disabled). See [Security & Least-Privilege Roles](#security--least-privilege-roles). Ignored in bring-your-own mode (the connection string picks the role). | -| `connectionString` | *(required unless managed)* | Npgsql connection string for a store you provision yourself, e.g. `Host=localhost;Port=5432;Username=darling;Password=...;Database=darling`. You own that cluster's settings: if it has TimescaleDB, size its [background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont) — managed mode does this for you, this mode does not. | - -### servers (array, at least one entry) - -| Key | Default | Notes | -|---|---|---| -| `name` | `""` | Display name; falls back to `host` | -| `host` | *(required)* | Server/instance to monitor | -| `engine` | `"sqlserver"` | `"sqlserver"` or `"postgres"` (`postgresql` / `pg` / `aurora-postgresql` also accepted). Configuration rather than something probed, because it decides which driver builds the connection string before there is a connection to ask. An omitted or unrecognized value means SQL Server, so every existing `darling.json` keeps its exact present behaviour — see [PostgreSQL targets](#postgresql-targets) | -| `port` | *(driver default)* | PostgreSQL targets on a non-default port. SQL Server carries its port in the host as `host,1433` instead, and that convention is left alone | -| `database` | *(none)* | Azure SQL Database only: the one database this entry monitors (also part of the server's storage identity). PostgreSQL targets connect to the maintenance database and read cluster-wide catalogs | -| `auth` | `"integrated"` | `"integrated"` or `"sql"` | -| `username` | *(none)* | Required for `"sql"` | -| `encryptedPassword` | *(none)* | DPAPI blob from `--encrypt-password` (preferred) | -| `password` | *(none)* | A literal (dev only, warned on every use) or an `env:NAME` / `file:/path` reference (#1804) — references are the supported non-Windows shape and are not warned | -| `readOnlyIntent` | `false` | Route to a readable AG secondary (`ApplicationIntent=ReadOnly`) | -| `trustServerCertificate` | `false` | | -| `encryptMode` | `"Mandatory"` | `Mandatory` / `Strict` / `Optional`; unknown values fail closed to `Mandatory` | -| `multiSubnetFailover` | `false` | | -| `excludedDatabases` | `[]` | Databases excluded from collection | - -### capturePlans (boolean, optional) - -| Key | Default | Notes | -|---|---|---| -| `capturePlans` | `true` | Capture execution plans into `query_stats.query_plan_xml` and `query_store_stats.query_plan_text`. PostgreSQL TOAST compresses the plan text transparently (LZ4 on the managed store) and TimescaleDB chunk compression squeezes it further, so plans are cheap to keep — unlike Lite, which stores to DuckDB/Parquet and deliberately never captures them. Set `false` to skip plan capture (e.g. to shave storage across a very large fleet). | - -### collectSchemaChangeEvents (boolean, optional) - -| Key | Default | Notes | -|---|---|---| -| `collectSchemaChangeEvents` | `true` | Record `Object:Created` / `Object:Altered` / `Object:Deleted` schema-change (DDL) events in the built-in default-trace collector. Set `false` on a noisy or benchmark box where a create/drop-happy workload floods the viewer's **System Events > Default Trace** tab — e.g. HammerDB's TPC-H Query 15 creates and drops a `revenue` view thousands of times, and the collector faithfully records every create/delete. Only the Object DDL slice is suppressed; file auto-grow/shrink, ErrorLog, and security-audit events are still collected. The shared collector's equivalent of the full Dashboard's `@include_object_events`. A file-only knob (not stored in the control plane): edit and restart. | - -### alerts - -The shared alert engine's switches and thresholds. Every default mirrors Lite's alert defaults exactly, so an empty section alerts like a fresh Lite install. `enabled: false` turns off all alert evaluation **and** scheduled-analysis finding notifications (the analysis itself still runs and persists findings). - -| Key | Default | Meaning | -|---|---|---| -| `enabled` | `true` | Master switch for alert evaluation + finding notifications | -| `cpuEnabled` | `true` | | -| `cpuThresholdPercent` | `80` | | -| `cpuMode` | `"total"` | `"total"` = SQL + other processes; `"sql"` = SQL process only | -| `blockingEnabled` | `true` | | -| `blockingCountThreshold` | `1` | Blocked-process count (rolling window) that trips the alert | -| `blockingWaitSecondsThreshold` | `0` | Total blocked wait, in seconds, summed across the latest blocking snapshot; `0` = off. A second gate beside the count one, because a count cannot tell one session blocked for an hour from one blocked for a second. Reports as its own "Blocking Wait Time" alert, and unlike the count gate it is level-triggered: it re-fires every cooldown while the wait stays above the threshold and clears when it drops below | -| `deadlockEnabled` | `true` | | -| `deadlockCountThreshold` | `1` | Deadlock count (rolling window) that trips the alert | -| `poisonWaitEnabled` | `true` | THREADPOOL / RESOURCE_SEMAPHORE / RESOURCE_SEMAPHORE_QUERY_COMPILE | -| `poisonWaitThresholdMs` | `500` | Average ms per wait | -| `longRunningQueryEnabled` | `true` | | -| `longRunningQueryThresholdMinutes` | `30` | | -| `tempDbSpaceEnabled` | `true` | | -| `tempDbSpaceThresholdPercent` | `80` | | -| `lowDiskEnabled` | `true` | Volume free space; graded CRITICAL when critically low | -| `lowDiskThresholdPercent` | `10` | Fire below X% free; `0` disables this dimension (clamped 0–100) | -| `lowDiskThresholdGb` | `5` | Fire below X GB free; `0` disables this dimension | -| `longRunningJobEnabled` | `true` | SQL Agent job running long vs. its history | -| `longRunningJobMultiplier` | `3` | Fires at 3x the job's historical average | -| `failedJobEnabled` | `true` | Live msdb check for recently failed jobs | -| `failedJobLookbackMinutes` | `60` | Clamped 1–1440 | -| `cooldownMinutes` | `5` | Minimum minutes between repeats of the same alert condition (clamped 1–120) | -| `excludedDatabases` | `[]` | Excluded from blocking/deadlock/long-running-query **alert evaluation** (collection unaffected) | - -Not configurable (hardcoded to Lite's defaults until someone needs a knob): the long-running-query read shape (top 5 results; the five noise filters — sp_server_diagnostics, WAITFOR, backups, misc waits, CDC — all on) and the analysis-finding notification policy (notify at severity >= 1.5, 6-hour per-finding cooldown). - -### smtp - -Email delivery is enabled when `host`, `from`, and `to` are all set — there is no separate enable flag. - -| Key | Default | Notes | -|---|---|---| -| `host` | `""` | | -| `port` | `587` | | -| `useSsl` | `true` | | -| `username` | *(none)* | For authenticated relays | -| `encryptedPassword` | *(none)* | Same `--encrypt-password` DPAPI pattern as SQL auth | -| `password` | *(none)* | A literal or an `env:NAME` / `file:/path` reference (#1804) — the non-Windows email path | -| `from` | `""` | | -| `to` | `""` | Comma-separated recipients | -| `emailCooldownMinutes` | `15` | Email/webhook channel cooldown (clamped 1–120) | - -### webhooks - -A channel is enabled by a non-empty URL. - -| Key | Default | Notes | -|---|---|---| -| `teamsUrl` | `""` | Teams incoming webhook | -| `teamsProxy` | `""` | Optional proxy address | -| `slackUrl` | `""` | Slack incoming webhook | -| `slackProxy` | `""` | Optional proxy address | - -### mcp - -The embedded MCP server, over Streamable HTTP bound to `localhost` by default (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan) to reach it — and the store — from the LAN). It exposes the same tool names Lite and the Dashboard expose, plus small Darling-only WRITE surfaces — Custom Views management, alert tuning, and server onboarding (see the last three bullets): - -- **Six diagnostic-analysis tools** — `analyze_server`, `get_analysis_facts`, `compare_analysis`, `audit_config`, `get_analysis_findings`, `mute_analysis_finding`. -- **Five plan-analysis tools** — `analyze_query_plan` (by `query_hash`), `analyze_procedure_plan` (by `sql_handle`), `analyze_query_store_plan` (by `database_name` + `query_id`), `analyze_plan_xml` (raw showplan XML, no fetch), and `get_plan_xml` (raw stored plan XML by `query_hash`). These run the shared execution-plan analyzer over the plan XML the collectors already captured into the store — a stored-plan read, never a live query against the monitored server. `analyze_query_plan`/`get_plan_xml` accept an optional `database_name`, and `analyze_query_store_plan` an optional `plan_id`, to pin the exact stored plan when the caller knows it. -- **Fifteen core data-read tools** — the diagnostic reads an assistant needs to investigate a server, each a stored read of the collected data (never a live query against the monitored server): - - *Resource metrics* — `get_cpu_utilization`, `get_wait_stats`, `get_wait_trend`, `get_wait_types` (the distinct observed wait types, to pick one for `get_wait_trend`), `get_memory_stats`, `get_memory_clerks`, `get_file_io_stats`, `get_tempdb_trend`, `get_perfmon_stats`. - - *Query performance* — `get_top_queries_by_cpu`, `get_top_procedures_by_cpu`, `get_query_store_top` (these hand back the `query_hash` / `sql_handle` / `query_id` + `plan_id` keys the plan-analysis tools consume). - - *Discovery / health* — `list_servers` (with collection-freshness status, and the [declared peer stores](#peers) when the fleet is split across several Darling boxes), `get_collection_health`, `get_server_properties`. - - These are the tools the analysis findings' `next_tools` recommendations point at, so a client following a finding's advice resolves them on this same server. Result shapes match Lite's (the store is Lite's collector schema); where Lite and the Dashboard's shapes diverge, Darling follows Lite — the shape its collector-mirror store can serve faithfully. -- **Twenty diagnostic-depth data-read tools** — deeper reads for a blocking / deadlock / session / configuration / storage investigation, each a stored read: - - *Blocking / deadlocks* — `get_blocking` (blocked/blocking pairs from the blocked-process-report XE + the always-on DMV fallback), `get_deadlocks`, `get_deadlock_detail` (raw graph XML), `get_blocked_process_xml` (raw report XML), and the per-minute count series `get_blocking_trend` / `get_deadlock_trend`. - - *Sessions* — `get_session_stats` (latest per-application connection counts), `get_active_queries` (captured running-query snapshots), `get_waiting_tasks`. - - *Config* — the change history `get_server_config_changes`, `get_database_config_changes`, `get_trace_flag_changes`, plus the latest-snapshot pair `get_database_scoped_config` / `get_query_store_health` (per-database Query Store health: actual vs desired state, readonly_reason decoded, storage vs cap) and the current-config snapshots `get_server_config` / `get_database_config` / `get_trace_flags` (what sp_configure / sys.databases / the active trace flags are set to **right now** — the companion to the `*_changes` diffs, which are empty on a stable server). - - *Index / object* — `get_table_index_sizes` (size + growth), `get_index_usage` (Unused / Write-only / Active), `get_object_locking` (lock/latch contention), `get_database_sizes`. - - The three config-change tools diff the store's config snapshots. This edition captures configuration **when the service connects** to a server (not on a fixed schedule), so a change is detected between two connect snapshots and at least two are needed — a stable, always-connected deployment may show no changes until the next connect. They emit only the values the collectors capture; the Dashboard's `requires_restart` / setting `description` / `setting_type` / generated change-narrative enrichment is not collected here and is omitted. The Dashboard's `get_blocking_deadlock_stats` aggregate is **not** hosted (Darling has no blocking/deadlock rollup table — use `get_blocking` / `get_deadlocks` for the raw events). - -- **Eight resource-contention + jobs data-read tools** — deeper reads for an internal-contention / worker-thread / plan-cache / SQL Agent investigation, each a stored read of the latest collected snapshot: - - *Latch / spinlock* — `get_latch_stats` (top latch classes by wait time, per-second rates), `get_spinlock_stats` (top spinlocks by collisions). - - *Memory grants* — `get_resource_semaphore` (workspace-memory target / max-target ceiling vs granted / used), `get_memory_grants` (per-pool grant detail), `get_memory_pressure_events` (RING_BUFFER_RESOURCE_MONITOR notifications — the process/system pressure indicators, not on Azure SQL DB). - - *Plan cache / scheduler* — `get_plan_cache_bloat` (single-use vs multi-use + bloat level), `get_cpu_scheduler_pressure` (runnable queue, worker utilization, pressure level). - - *Jobs* — `get_running_jobs` (running SQL Agent jobs vs historical average / p95). - - The Dashboard's per-class latch `severity` / `description` / `recommendation`, spinlock `description`, plan-cache `bloat_level`, and CPU-scheduler `pressure_level` / `recommendation` are the Dashboard / reporting-view CASE derivations (not collected columns), reproduced service-side so the full result shape is served. Darling's delta collectors store no `sample_interval_seconds`, so per-second latch/spinlock rates are derived from the collection interval, and the Dashboard's `get_resource_semaphore` `sample_interval_seconds` is not emitted for the same reason (`max_target_memory_mb`, the workspace-memory ceiling, is added since the store carries it). - -- **Seven PostgreSQL data-read tools** — the read surface for a PostgreSQL target's collectors, each a stored read (see [PostgreSQL targets](#postgresql-targets)): - - *Waits and queries* — `get_pg_wait_stats` (top wait events in the window, decoded to type + event name), `get_pg_top_queries` (query shapes by total execution time, carrying Aurora's storage-vs-cache I/O split and per-statement peak memory). - - *Outage predictors* — `get_pg_wraparound_risk` (XID and MultiXact freeze headroom per database), `get_pg_xmin_horizon` (why vacuum is reclaiming nothing, attributed to the specific holder), `get_pg_replication_slots` (slot health, including whether retained WAL is still growing). - - *Maintenance* — `get_pg_autovacuum_health` (tables behind on vacuum or analyze, ranked by how far past each table's OWN trigger threshold it is — the ratio, not the dead-tuple count, because the same count is routine on a large table and urgent on a small one). - - *I/O attribution* — `get_pg_io_stats` (reads, hits, extends and evictions by backend type, object and context). The context dimension is the one with no SQL Server counterpart and the one that changes the remedy: it separates ordinary buffer-pool misses, where more `shared_buffers` or a better index helps, from sequential scans that deliberately bypass the pool through a small ring buffer, where neither will. - - These are separate tools rather than widened SQL Server ones. PostgreSQL's waits are a two-level type/event taxonomy with no signal-wait concept reported in microseconds, and the wraparound / horizon / slot signals have no SQL Server counterpart at all — sharing a result shape would mean lying about a unit or emitting mostly-null columns. The three outage predictors are the ones worth wiring to a pager: each names a condition that stops the server outright, and each is silent until it is nearly too late. - -- **Five trend data-read tools** — windowed time-series siblings of the core reads, each a stored read of the collected series over the window (BOTH-sides, naive-UTC): - - `get_memory_trend` (total / target server memory, buffer pool, plan cache over time), `get_perfmon_trend` (a single counter's value + delta, `counter_name` required), `get_file_io_trend` (per-database read/write latency, top-10 busiest files), `get_query_trend` (one query's per-collection history by `query_hash` + `database_name`), `get_query_duration_trend` (overall elapsed-ms/sec + executions/sec). - - Each mirrors the viewer's proven chart read (byte-identical Postgres SQL); the shape follows Lite where the SKUs diverge. `get_perfmon_trend` reproduces Lite's miss vocabulary (Page Life Expectancy is intentionally not collected; an unknown counter hands back the collected names). `get_memory_trend` carries a `total_granted_mb` field for field-for-field parity with Lite, where its memory_stats-only read leaves it 0 (the grant overlay is a separate chart series). - -- **Eight system-health parse-on-read tools** — the Dashboard's `get_health_parser_*` family, over Darling's raw `system_health_events`: - - `get_health_parser_system_health` (corruption + contention counters), `get_health_parser_severe_errors` (severity ≥ 19, with `database_id` resolved to a name), `get_health_parser_scheduler_issues`, `get_health_parser_memory_conditions`, `get_health_parser_memory_broker`, `get_health_parser_memory_node_oom`, `get_health_parser_cpu_tasks`, `get_health_parser_io_issues`. - - Where the Dashboard reads its server-side-parsed `collect.HealthParser_*` tables, these shred the raw extended-event XML **on read** with the shared `SystemHealthParser` (the same parser the viewer's System Events tab uses) and gate with the service-side twin of the viewer's `SystemEventSignificance` — returning the same SIGNIFICANT warning set the Dashboard surfaces (sp_HealthParser at `@warnings_only = 1`). `get_health_parser_system_health` is the one UNGATED category (its counter series plots every snapshot). Each row carries the full sp_HealthParser column set keyed on the event's `event_time`; the tools window on `event_time` (the event's real time), so "last 24 hours" means events that happened in the last 24 hours. - -- **Five alert + health-overview tools** — the fleet-triage reads the fleet edition previously lacked, each a stored read over the monitoring store (no live hit): - - *Alerts* — `get_alert_history` (what fired, value vs threshold, delivery success/failure, muted — fleet-wide by default, or scoped to a server), `get_alert_settings` (the current alert config the service is using — per-alert enable/thresholds, cooldown, excluded databases, delivery mode, analysis cadence), `get_mute_rules` (the alert mute rules in force, so a suppressed server is distinguishable from a healthy-quiet one). - - *Health overview* — `get_server_summary` (one-shot per-server CPU / memory / recent blocking / recent deadlocks), `get_daily_summary` (a day's composite health band — Healthy / Warning / Critical — folded through the shared `DailyHealthBandCalculator`, plus the signals behind it). - -- **Eight Custom Views tools (Darling-only)** — discover, create, and manage the saved dashboards/notebooks a user composes from the curated measure catalog (the same views the web viewer's editor builds), stored in `config.custom_views`. None touches a monitored SQL Server or the collected performance data — the write tools write only view definitions to the monitoring store. - - *Discover* — `describe_custom_view_catalog` (the compose vocabulary — measures with their source/kind/valid-aggregates/allowed-dimensions/units/per-server-type availability, dimensions, unit families, aggregates, time buckets, filter ops, and viz types). An MCP client calls this FIRST so a composed panel uses only legal identifiers instead of guessing at names; it returns the SAME `/api/catalog` vocabulary the web composer's picker binds to. Read-only static reference — no store, no server. - - *Read* — `list_custom_views` (summaries: id, name, description, kind, version), `get_custom_view` (one view's full definition + version). - - *Author* — `validate_custom_view` (dry-run a definition against the catalog + composer rules, no save), `create_custom_view` (validate then save), `update_custom_view` (validate then replace in place, optimistic-concurrency on `version`), `delete_custom_view`. - - *Self-test* — `run_custom_view_panel` (compile + run a single composed panel and return `{sql, rows, annotations}` — the composer's live preview, for checking a generated panel's data before saving). - - The create/update/delete tools are the one view-authoring **write** surface; create/update run the SAME `ValidateDefinition` authority as `validate_custom_view`, so an invalid definition is rejected before it stores; every tool routes through the SAME store + validator + compile-and-run + catalog the web viewer's editor uses (no divergent second implementation). This write surface is part of what the MCP token gates — see [What a token can reach](#opt-in-network-endpoints-lan) below. - -- **Three alert-tuning write tools (Darling-only)** — `update_alert_settings`, `create_mute_rule`, and `delete_mute_rule` let an MCP client TUNE the alert engine the fleet shares — the SAME config `get_alert_settings` / `get_mute_rules` read and the Viewer's Settings window writes. `update_alert_settings` is a PARTIAL update of the single global settings row: read via `get_alert_settings`, change fields, and send only those back in the same nested shape; every field is validated against the SAME ranges/enums the Settings window enforces BEFORE any write, an out-of-range or unknown field returns `{status:"invalid"}` and writes nothing, and the write self-bumps `config_version` so the running service hot-reloads within one collection sweep. `create_mute_rule` / `delete_mute_rule` reuse the SAME `PgMuteRuleStore` `get_mute_rules` reads through (and the same GUID id-generation the Viewer's mute-create path uses). None touches a monitored SQL Server or the collected data — only the shared alert configuration; SMTP/webhook delivery credentials are out of scope (the `mcp` role cannot read or write the secret columns). It is part of what the MCP token gates — see [What a token can reach](#opt-in-network-endpoints-lan) below. - -- **Two server-onboarding write tools (Darling-only)** — `add_servers` (BULK) and `remove_server` let an MCP client stand up or tear down FLEET monitoring conversationally ("monitor these twenty servers with this login"), the service-side twin of the Viewer's Add / Manage Servers dialogs. `add_servers` takes a JSON **array** of server objects (`host` required; optional `display_name` / `database` / `read_only_intent` / `multi_subnet_failover`; `auth` `Windows`/`SQL` with `username`+`password` for SQL; and the exposed TLS options `encrypt_mode` `Optional`/`Mandatory`/`Strict` + `trust_server_certificate`) and processes them **in order**: it validates each entry, PROBES the connection in-process (reusing the same `DarlingServerConnector.ProbeAsync` the `--test-connection` verb runs — the service holds the network path + credentials, so no `test_connect` command plane is needed), skips a case-folded duplicate (`duplicate`) of an already-monitored server or an earlier entry, DPAPI-encrypts the SQL password (the service identity, so it round-trips at collection time), and INSERTs the row mirroring the service's own seed shape. A server that fails to connect is `connection_failed` and the batch continues; Entra/MFA/Service-Principal/Managed-Identity auth is `invalid` (the service connects with Windows or SQL only). `remove_server` DELETEs a monitored server by name (resolved the same way every `server_name` is) — already-collected history is kept. Both write only the monitoring store's `config.config_monitored_servers` registry; neither runs anything on a monitored server beyond the one-time probe. **The SQL password travels to the endpoint inside `add_servers`' request** and is DPAPI-encrypted at rest (never returned) — it is part of what the MCP token gates, and it puts a credential on the wire; see [What a token can reach](#opt-in-network-endpoints-lan) below. - -| Key | Default | Notes | -|---|---|---| -| `enabled` | `false` | **Off by default** — a headless service does not open a local port unless you ask | -| `port` | `5152` | Chosen so all three editions coexist on one machine (Dashboard 5150, Lite 5151) | - -Register with Claude Code: - -``` -claude mcp add --transport http --scope user sql-monitor-darling http://localhost:5152/ -``` - -If the port is already in use at startup, the MCP server logs an error and does not start; collection is unaffected. - -### web - -The embedded read-only **web dashboard** — a browser view of the monitoring store, served over HTTP on its OWN port (default **5153**), separate from the MCP server. It is a distinct surface from [`### mcp`](#mcp): its own enable flag, port, token, and exposure block, because the two gate different blast radii (the MCP token guards `analyze_server`'s **live outbound** connections to your monitored SQL Servers; the web dashboard is **read-only over the collected store**). It connects to the store as the least-privilege `viewer` role. Loopback-only by default; see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan) to reach it from the LAN. - -| Key | Default | Notes | -|---|---|---| -| `enabled` | `false` | **Off by default** — a headless service does not open a local port unless you ask | -| `port` | `5153` | Chosen so all four local surfaces coexist on one machine (Dashboard 5150, Lite 5151, Darling MCP 5152) | - -Once enabled, open `http://localhost:5153/` in a browser on the service host. Like the MCP server, `enabled`/`port` here are the file SEED; after first start they live in the control plane and the Viewer's Settings toggles them LIVE (the service starts/stops/rebinds the dashboard within seconds — no restart). If the port is already in use at startup, the web host logs an error and retries on a calm cadence; collection is unaffected. - -**What you see.** The dashboard opens on a **Fleet Overview**: a card per enabled server with a status dot, six per-metric health bands (CPU, threads, memory, blocking, deadlocks, collectors), and its last collection time — all banded server-side, so the browser only renders (a server that has never reported shows an amber "Awaiting first collection", never a red offline). Above the cards a worst-first "Needs attention" list surfaces the servers to look at, or an all-healthy line when there is nothing to chase. Click a card to **drill into one server**: an overview, wait stats with a trend for the heaviest wait, active queries, a CPU chart, memory and file-I/O trends, and collection health — the same collected data the viewer shows, over inline charts. A fleet-wide **Alert History** page (with a server filter box) rounds out phase 1. It is a read-only view — no settings, no write paths, no live-server queries — and refreshes every 60 seconds (pausing while the tab is hidden). The frontend ships fully self-contained (no CDN, no fonts, no remote anything), so it works on an air-gapped host with no internet access. - -### peers - -**Declared peer stores** — optional, and only relevant when the fleet is split across **several Darling boxes**, one store each (SQL Server primaries on one box, their readable replicas on another, PostgreSQL on a third). Each box's MCP server answers over **its own** store only, so a server monitored by a sibling resolves as not-found — which an agent cannot tell apart from *"nobody monitors this server."* Declaring the siblings fixes that at the three places an agent forms its picture of the fleet. - -**Disclosure only.** There is no address and no credential in this block, and nothing behind it: the service never contacts a peer, cannot read a peer's data, and cannot tell whether a peer is even running. A peer is a **name** plus a **sentence**, so an agent (or its human) can pick the right endpoint. Everything here is sent verbatim to every connected MCP client, so the service **refuses to start** if any peer text looks like a connection string or credential. - -| Key | Default | Notes | -|---|---|---| -| `thisStoreCovers` | `""` | One sentence naming what THIS store monitors — the anchor the peer list is relative to | -| `stores[].name` | — | **Required.** Whatever an operator would recognize (the box name, "the use1 store") | -| `stores[].covers` | `""` | A short sentence naming what that store monitors. Human prose — never parsed, only shown | -| `stores[].matches` | `[]` | Optional server-name **substrings** that store monitors, case-insensitive. The only machine-checked field | - -```jsonc -"peers": { - "thisStoreCovers": "the 42 us-east-1 SQL Server primaries", - "stores": [ - { - "name": "prod-pos-use2-monitor-01", - "covers": "the readable replicas of those same 42 primaries, in-region from us-east-2", - "matches": ["use2"] - }, - { "name": "prod-pos-pg-monitor-01", "covers": "the Aurora PostgreSQL clusters", "matches": ["-aurora-"] } - ] -} -``` - -What it changes, with peers declared: - -- **The MCP instructions** gain a Fleet Coverage section, high enough that an agent reads which store it is talking to before it reads the tool census. -- **`list_servers`** gains `this_store_covers`, a `peer_fleets` array, and a `peer_note`. Both are always present: an *empty* `peer_fleets` has two very different meanings (this really is the only store, or nobody declared the siblings) and the service cannot tell them apart, so `peer_note` says exactly that rather than letting an empty array read as "this is the whole fleet." An **empty registry** answers in prose rather than JSON, and carries the peer list too — a store with nothing registered is a fresh or just-restarted box, which is the worst place to drop the disclosure. -- **The server-resolution miss** appends the disclosure to the existing "Could not resolve server. Available servers:" listing, naming the peer whose declared coverage matches — so *not monitored here* stops looking like *not monitored anywhere*. - -`matches` is deliberately plain substrings, no globbing and no regex: it exists to answer "which region/role prefix is this name?", and a pattern language would be a config surface with its own failure modes. Blank entries are dropped — an empty substring matches every name, which would make one peer claim the whole fleet. A peer with no `matches` is still disclosed everywhere; it just cannot be singled out on a miss, and the miss message says so instead of implying the server is unmonitored. - -A **file-only** block (not seeded into the control plane): it describes the deployment topology of *this* box, which must not be editable from a peer's Viewer. An edit takes effect on the next service restart. There is deliberately **no cross-store connectivity** here — actual federated reads (auth between stores, latency, partial failures) are a much larger surface, and may never be worth building if disclosure alone makes the split legible. - -**Declaring nothing changes nothing, with one exception worth knowing about on upgrade.** The instructions, the resolution-miss message, and `list_servers`' empty-registry sentence are byte-for-byte what they were. But `list_servers`' JSON envelope carries `this_store_covers`, `peer_fleets` and `peer_note` on *every* response, declared or not — so a script comparing that tool's exact shape sees three new keys even if you never write a `peers` block. That is deliberate: an empty `peer_fleets` means *either* "this is the only store" *or* "nobody declared the siblings", and a note that only appeared when peers were declared would say nothing in precisely the case that produces the wrong conclusion. - -**A `peers` block that fails validation is refused whole, and nothing is disclosed** — not the valid subset. An unfinished block that asserts coverage which may be wrong is worse than no block, and the service logs each problem at Critical. The check runs inside the publish rather than only in config validation, because the MCP host loads its own config and deliberately never validates it (its fail-closed checks are host-local), so validation alone would leave the one path that actually broadcasts uncovered. - -### No Schedule Knobs, by Design - -There are deliberately **no collection-schedule or retention settings** in `darling.json`. The service consumes the shared per-collector defaults (`CollectorScheduleDefaults`) — the same cadences and retention horizons a fresh Lite install uses, identity-pinned by tests so the two editions cannot drift. If a schedule knob is ever genuinely needed, it will be added then, not speculatively. - ---- - -## PostgreSQL Targets - -Darling monitors PostgreSQL alongside SQL Server. For the ordered procedure with a proof point at each step, see [**the first-target runbook**](../docs/postgres-first-target-runbook.md); this section is the reference for what each piece does. - -Add `"engine": "postgres"` to a `servers` entry and that target is collected by the PostgreSQL collectors instead of the T-SQL ones: - -```json -{ - "name": "orders-prod", - "engine": "postgres", - "host": "orders-prod.cluster-abc123.us-east-1.rds.amazonaws.com", - "auth": "sql", - "username": "darling_monitor", - "encryptedPassword": "" -} -``` - -**Which path registers a target depends on the store, not the file.** `darling.json` seeds `config.config_monitored_servers` once, when it is empty; after that the registry is authoritative and a darling.json edit adds nothing. So a fresh install declares its PostgreSQL targets in the file, and an existing one adds them with the [`add_servers`](#mcp) tool (or the Viewer's Add Server dialog), which takes effect within one collection sweep without a restart. The registry carries `engine` and `port` per row, so a target keeps its engine across restarts and reloads. - -`auth` must be `"sql"` — PostgreSQL has no integrated-authentication path here, and an entry asking for it fails [`--test-connection`](#validate-the-config-pre-flight) rather than waiting to fail at first connect. `add_servers` enforces the same rule, and unlike the file parser it REFUSES an unrecognized `engine` rather than resolving it to SQL Server: the file's leniency keeps one bad line from stopping the whole fleet at startup, while onboarding is a single deliberate act where a silent fallback would surface as a connection failure against the wrong port. Password handling is identical to a SQL Server entry: `--encrypt-password` produces the DPAPI blob, and the `env:NAME` / `file:/path` references work the same way. TLS defaults to full certificate verification (`SslMode=VerifyFull`); `trustServerCertificate` relaxes it to `Require`, which is the setting Aurora usually needs since it presents an RDS CA a stock trust store does not know, and `"encryptMode": "Optional"` relaxes it further to `Prefer`. - -**One store, both engines.** The PostgreSQL collectors write to the same store as the SQL Server ones, into their own tables, on the same naive-UTC contract and the same `server_id` identity. Nothing is partitioned by engine — a mixed fleet is one store, one viewer, one MCP endpoint. - -**A collector table must never be named after a `pg_catalog` object.** `pg_catalog` is searched implicitly and *first*, ahead of every entry in `search_path`, so an unqualified reference to such a name resolves to the system object no matter what the store holds. It fails loudly in one place — `CREATE INDEX` on a view is 42809, which aborts the migration and leaves the store unusable — and silently everywhere else: a reader's `FROM ` would return the monitoring store's own system view instead of collected history, so the tool reports nothing and any alert behind it never fires. That is why the slot collector stores into `pg_replication_slot_stats` while still being *named* `pg_replication_slots` after the view it reads (the same split as `query_store` → `query_store_stats`). A live-store test asserts no collector table shadows a catalog object, against the real catalog rather than a hardcoded reserved list. - -**A collector never runs against the wrong engine.** Every definition declares its `TargetEngine`, and both SKUs check it before dispatch, so a PostgreSQL target is never sent T-SQL and a SQL Server target never sees `pg_stat_statements`. A store monitoring only SQL Server still carries the five PostgreSQL tables, empty; nothing else about it changes. - -### Permissions on a PostgreSQL target - -One role covers every collector: - -```sql -CREATE ROLE darling_monitor WITH LOGIN PASSWORD ''; -GRANT pg_monitor TO darling_monitor; -``` - -`pg_monitor` is the standard PostgreSQL monitoring role — it bundles `pg_read_all_stats`, `pg_read_all_settings`, and `pg_stat_scan_tables`. Without it the statistics views still return rows, but only for the connecting user's own backends, which silently turns fleet monitoring into self-monitoring. On Amazon Aurora and RDS the same grant works: `GRANT pg_monitor TO darling_monitor;` as an `rds_superuser`. No superuser is needed, and nothing is created on the monitored server — unlike a SQL Server target, there are no Extended Events sessions to provision and no server setting to bootstrap. - -`pg_stat_statements` must be present for `pg_statement_stats`, which means the extension in `shared_preload_libraries` (a restart, or a parameter-group change plus reboot on Aurora/RDS) and `CREATE EXTENSION pg_stat_statements;` in the database Darling connects to. The extension tracks **all** databases in the cluster keyed by `dbid`, so one installation in the connect database covers the whole instance. The other six collectors need nothing installed — they read core catalogs and Aurora's built-in functions. - -### What gets collected - -| Collector | Source | Cadence / retention | Why it exists | -|---|---|---|---| -| `pg_wait_stats` | `aurora_stat_system_waits()` | 1 min / 30 d | **Aurora only.** Core PostgreSQL has no cumulative wait counters at all — `pg_stat_activity.wait_event` is an instantaneous sample — so there is no equivalent to `sys.dm_os_wait_stats` to read on a non-Aurora target | -| `pg_statement_stats` | `aurora_stat_statements()` | 1 min / 30 d | **Aurora only.** Per-query-shape totals, matching `query_stats`' cadence. Aurora's function adds the storage-vs-cache I/O split (`storage_blks_read` / `orcache_blks_hit`) and per-statement peak memory, neither of which core PostgreSQL exposes | -| `pg_wraparound_stats` | `pg_database`, `pg_class` | 5 min / 90 d | XID and MultiXact freeze headroom per database. The highest-consequence signal PostgreSQL has and one with no SQL Server counterpart: run out of transaction IDs and the server stops accepting writes. Freeze headroom moves in autovacuum-sized steps rather than continuously, so 5 minutes is ample and 90 days shows the age trend against the actual freeze threshold | -| `pg_xmin_horizon` | `pg_stat_activity`, `pg_replication_slots`, `pg_stat_replication`, `pg_prepared_xacts` | 1 min / 30 d | Why vacuum is reclaiming nothing. Four unrelated causes produce an identical symptom and need completely different fixes, so this attributes the specific holder instead of reporting the number. Per-minute because a holder is the fast-moving leading indicator — the useful answer is which session or slot appeared minutes ago | -| `pg_replication_slots` → `collect.pg_replication_slot_stats` | `pg_replication_slots` | 1 min / 90 d | Slot health and retained WAL. An abandoned slot retains WAL without bound by default, filling the volume and stopping the server, and it grows at whatever rate the server writes WAL — hours on a busy writer, not days | -| `pg_autovacuum_stats` | `pg_stat_user_tables`, `pg_class` | 60 min / 90 d | **Writers only**, and per database. Per-table autovacuum state. Stores each table's own computed trigger threshold beside its dead-tuple count, honouring per-table `reloptions` overrides rather than only the GUCs — without the threshold a dead-tuple count is not actionable, since the same count is routine on a large table and urgent on a small one | -| `pg_io_stats` | `pg_stat_io` | 1 min / 30 d | I/O attributed to a `(backend_type, object, context)` triple rather than to a file — who did it, to what, and why. PostgreSQL 16+; valid on a standby. Every counter column is nullable on purpose: PostgreSQL uses NULL for "does not apply to this combination", and on Aurora the whole write side is NULL because backends there do not write data files. The read reports whether write counters are TRACKED, so absent writes cannot be misread as zero writes | -| `pg_blocking` → `collect.pg_blocking_edges` | `pg_stat_activity`, `pg_blocking_pids()` | 1 min / 30 d | Who is blocked, by whom, and what state each side was in — stored as an edge list, one row per (blocked, blocking) pair, so the read layer can assemble chains and name the root. **This is a SAMPLE, not an event log**, and that is the one thing to carry away: SQL Server's blocked-process report is written by the engine when blocking crosses a threshold, whereas PostgreSQL records nothing unless something asks, so blocking shorter than the interval is never seen. Valid on a standby, where recovery conflicts are blocking that happens nowhere else. One minute is the floor worth paying for: `pg_blocking_pids()` takes ShareLock on the lock manager partitions per call, so it is evaluated only for backends already waiting on a lock | - -The two Aurora-only collectors are gated on Aurora detection, not on configuration: the connect probe looks for `aurora_version()` and the gate follows what it finds. Point Darling at self-managed PostgreSQL and the four core-catalog collectors run while those two sit out. - -`pg_autovacuum_stats` additionally gates OFF on a standby (`pg_is_in_recovery()`), and the reason is worth knowing because it is not a permissions or availability problem. `pg_stat_user_tables` reads fine on a replica and reports **all zeros**: measured on Aurora 17.7, the same cluster, database and 15 tables, the writer reported 13,654,458 dead tuples and 150,790,506 live tuples while the reader reported 0 for every tuple counter. Those are the writer's stats-collector numbers and they are not replicated. Ungated, a replica target would return no rows, the activity filter would read that as "nothing has pending work", and you would get a confident report of perfect autovacuum health for a cluster 13 million dead tuples behind. For the same reason, treat an empty `get_pg_replication_slots` result from a replica as per-instance rather than cluster-wide — slots live on the writer. - -Cadences and retention are the shared defaults, with no knobs, exactly as for SQL Server. - -Three of the seven are outage predictors rather than performance metrics, which is deliberate: PostgreSQL's most damaging failures are quiet, slow, and fully predictable days ahead, and nothing in the engine raises its hand about them. The read surface for each is an MCP tool — see [the tool list](#mcp). - -### What it does not do yet - -Plan capture and the blocking-chain reads have no PostgreSQL equivalent in the store yet. Alerting and scheduled analysis are still SQL-Server-shaped, so a PostgreSQL target collects and is readable through MCP and the viewer, but does not yet raise alerts or produce analysis findings. **The three Tier 0 outage predictors DO alert.** Wraparound risk, a blocked vacuum horizon and replication-slot retention are evaluated on the alert cadence and delivered through the same deliverer, history and mute rules as every SQL Server alert, so they land in the same places and obey the same suppression. They ride alongside the shared engine rather than inside it, via a separate `IPostgresAlertReadAdapter` consulted only for PostgreSQL targets — Lite has no PostgreSQL target, and extending the shared adapter would have left it implementing three methods that can only return empty. Thresholds are derived from the server's own settings (wraparound grades against that cluster's `autovacuum_freeze_max_age`, not a constant) and are not yet configurable; see [`docs/postgres-alerting-design-note.md`](../docs/postgres-alerting-design-note.md). - -Scheduled ANALYSIS is still SQL-Server-shaped, so a PostgreSQL target does not yet produce analysis findings. - -Collector FAILURES are classified, though. A PostgreSQL fault is routed through the same `ITargetProvider.Classify` the engine seam exposes, so a persistent, operator-actionable condition records as a non-fatal skip with an explanation instead of logging `ERROR` every cycle — a `pg_stat_statements` view that was never created (SQLSTATE 42P01), a source Aurora does not implement (0A000), a feature switched off in the parameter group (55006). The message says which kind it is, because the store's non-fatal bucket is named PERMISSIONS and none of those is a missing grant. Connection-level failures (the 08 class, 57P0x) still force a reconnect and reprobe; a `statement_timeout` (57014) deliberately does not, since dropping the connection over a slow query would turn a tuning problem into a reconnect storm. - -The per-database fan-out itself is done — `pg_autovacuum_stats` is the collector that exercises it. Worth knowing what it costs: a SQL Server collector can reach another database without reconnecting (`EXECUTE [db].sys.sp_executesql`), while a PostgreSQL connection is bound to one database for its lifetime, so a per-database PostgreSQL collector is necessarily one connection per database per cycle. That is why its cadence is hourly and why new per-database collectors should be added deliberately rather than by default. - ---- - -## Operations - -### The Store - -The service migrates the store itself at startup — plain versioned SQL scripts, each applied once inside its own transaction, tracked in `darling_schema_version`, safe under concurrent starters (advisory-locked). Current schema is **v73** — `StorageVersion.SchemaVersion` is the source of truth and a test pins it to the highest rung in the ladder. - -The notable rungs are below. For the **complete** current schema, read `Darling/Darling.Tests/Fixtures/migration-ladder-*.sql` — the whole ladder as resolved SQL, regenerated per release; it is generated, so don't hand-edit it. - -| Version | Contents | -|---|---| -| **V1** — collector tables | One table per collector, all 49, generated from the shared collector definitions (column-for-column identical to Lite's DuckDB schema): `wait_stats`, `latch_stats`, `spinlock_stats`, `query_stats`, `procedure_stats`, `query_store_stats`, `query_snapshots`, `plan_cache_stats`, `cpu_utilization_stats`, `cpu_scheduler_stats`, `file_io_stats`, `memory_stats`, `memory_clerks`, `memory_pressure_events`, `tempdb_stats`, `perfmon_stats`, `deadlocks`, `blocked_process_reports`, `dmv_blocking_snapshots`, `memory_grant_stats`, `waiting_tasks`, `session_stats`, `session_summary_stats`, `running_jobs`, `database_size_stats`, `index_object_stats`, `server_properties`, `system_health_events`, the four config snapshots (`server_config`, `database_config`, `database_scoped_config`, `trace_flags`), and the eight PostgreSQL tables listed under V63–V69 and V71 below | -| **V2** — observability | `servers` (registry, upserted on every successful connect: identity, display name, engine edition, major version) and `collection_log` (one row per collector run: SUCCESS / PERMISSIONS / ERROR, row count, SQL-phase and storage-phase timings) | -| **V3** — alerting | `config_alert_log` (one history row per fired alert), `config_edge_trigger_watermarks` (restart-surviving edge-trigger and failed-job watermarks), `config_mute_rules` (alert mute rules; starts empty) | -| **V4** — analysis | `analysis_findings` (persisted findings incl. the stored remediation action), `analysis_muted` (muted finding patterns), and 17 `v_
` passthrough views so the shared analysis SQL runs verbatim against this store | -| **V5** — viewer passthrough views | The five remaining `v_*` passthrough views (`v_running_jobs`, `v_server_config`, `v_database_scoped_config`, `v_trace_flags`, `v_collection_log`) that complete the viewer's read layer | -| **V6** — memory passthrough views | `v_memory_clerks` and `v_memory_pressure_events`, the two views the Memory tab reads | -| **V7** — plan-capture columns | Nullable plan-XML columns for the viewer's View Plan surfaces: `procedure_stats.query_plan_xml`, `blocked_process_reports.blocked_query_plan_xml` / `blocking_query_plan_xml`, `deadlocks.victim_query_plan_xml` | -| **V8** — schema split (collect/config) | Moves the tables into the `collect` and `config` schemas (least-privilege security split); the shared SQL keeps using bare names, resolved via `search_path = collect, config, public` | -| **V9** — inventory + cost fields | `server_properties` inventory columns (`sqlserver_start_time`, `host_os_version`, `ag_replica_role`) and `servers.monthly_cost_usd` (the FinOps per-server budget) | -| **V10** — latch + spinlock collectors | `latch_stats` and `spinlock_stats` tables plus their `v_*` views | -| **V11** — CPU scheduler + plan cache collectors | `cpu_scheduler_stats` and `plan_cache_stats` tables plus their `v_*` views | -| **V12** — session summary collector | `session_summary_stats` (server-wide connection-leak / idle signal) table plus its `v_*` view | -| **V13** — system health events collector | `system_health_events` (raw `system_health` Extended Events capture) table plus its `v_*` view | -| **V14** — refresh passthrough views | `CREATE OR REPLACE` on every `v_*` view so a store upgraded across a column-adding migration picks up the new columns (Postgres freezes a view's `SELECT *` expansion at create time) | -| **V15** — index metadata columns | Per-index definition columns on `index_object_stats` (ordered key/included column lists, filter, uniqueness/constraint/FK flags, `is_disabled`, and the reconstruct-a-CREATE options — compression, fill factor, page/row locks, etc.) for monitor-side UNUSED/DUPLICATE index analysis, and refreshes `v_index_object_stats` | -| **V16** — server UTC offset | Nullable UTC-offset column on `server_properties` so the viewer can render timestamps in the monitored server's own local time (the Server-time display mode ported from Lite; Server-time = stored naive-UTC + this offset) | -| **V17** — config control plane | The viewer-writable DESIRED-state tables (`config_service`, `config_monitored_servers`, `config_alert_settings`, `config_collector_schedules`) plus a `config_version` reload beacon — statement-level bump triggers increment it on any write, and the service polls that one integer each sweep and reloads only when it changes. Server secrets are DPAPI blobs, never plaintext | -| **V18** — alert delivery mode | Global `delivery_mode` (Summary / PerEvent) + `per_event_max` on `config_alert_settings`, plus a nullable per-server `alert_delivery_mode_override` on `config_monitored_servers` (null = inherit the global), resolved through the shared `AlertDeliveryModeResolver` (#1236 / #1141) | -| **V19** — analysis state marker | `collect.analysis_state` — the service-produced per-server "insufficient data" marker (with message + time) the viewer reads, so a not-enough-history analysis pass surfaces a reason instead of a blank | -| **V20** — alert tuning knobs | The previously-hardcoded alert tuning the viewer now customizes on `config_alert_settings`: the long-running-query read shape (`long_running_query_max_results` + five noise-filter opt-outs the shared `AlertEngine` forwards) and `notify_connection_changes` (the Server-Unreachable / Restored connect-edge gate) | -| **V21** — default trace events collector | `default_trace_events` table + its `v_*` view — the significant Default Trace events (file growth, ErrorLog, security audit, optional Object DDL) the viewer's System Events tab reads | -| **V22** — index-object latest index | The engine-agnostic `idx_index_object_stats_latest` partial index backing the latest-capture-per-index reads | -| **V23** — collection-log hypertable | Converts `collection_log` to a TimescaleDB hypertable (an object-invisible no-op on plain PostgreSQL) | -| **V24** — job history collector | `job_history` table + its `v_*` view — the SQL Agent Job History surface (#1433) | -| **V25** — agent status collector | `agent_status` table + its `v_*` view — SQL Agent up/down status (#1433) | -| **V26** — generic webhook channel | The generic-webhook columns on `config_notification` (`generic_url`, `generic_headers`, `generic_body_template`, `generic_proxy`) for POSTing alerts to any endpoint (#1506) | -| **V27** — deadlocks database name | `deadlocks.database_name` (the Azure SQL DB per-database deadlock-capture watermark key, #1535) and a refreshed `v_deadlocks` | -| **V28** — Query Store replica role | `query_store_stats.replica_role` (SQL Server 2022+ AG secondary-replica attribution, #1546) and a refreshed `v_query_store_stats` | -| **V29** — long-query completions collector | `collect.long_query_completions` + its index — the opt-in long-running-query completion trace's store table (#1496) | -| **V30** — web dashboard config | `config_service.web_enabled` + `web_port` — the read-only web dashboard's live enable/port toggle, the twin of `mcp_enabled`/`mcp_port` (#1562) | -| **V46** — automatic plan correction | `collect.plan_correction` + its index — the #1952 collector's store table (FORCE_LAST_GOOD_PLAN enablement plus the engine's live recommendation set). Additive and view-less, so a fresh store gets it from V1's generated schema and V46 is what an already-existing store gets | -| **V47** — ADR persistent version store | `collect.pvs_stats` + its index + the `v_pvs_stats` passthrough view — the #1951 ADR version-store collector's store table. A fresh store gets the table from V1's generated schema; V47 is what an already-existing store gets, and the view is what keeps the Darling viewer's FinOps read byte-identical to Lite's | -| **V61** — per-fingerprint occurrence counters | `config.config_incident_occurrences` — the accumulator’s memory for the monotonic count behind an alert incident (#2216). The count that rides on an incident is a GAUGE (it falls as events age out of the read window), so a consumer seeing only throttled deliveries cannot recover how many events happened between two of them. A NEW table, not columns on `config_edge_trigger_watermarks`: the key is wrong (per (server, metric) vs per (server, metric, fingerprint)) and Lite writes that row with a PARTIAL `INSERT OR REPLACE` column list, so an added column would zero itself every time an alert fired | -| **V62** — plan-XML codec knob | `config.config_service.plan_xml_compression` (#2171). `gzip` (default) keeps today’s write path; `none` stores plain text in `query_plan_xml` so direct-SQL readers get plans back — PostgreSQL exposes no inflate, so gzip bytes are unreadable without an untrusted-language UDF. Rides `config_service` like V58/V59 so the `config_version` trigger makes a flip visible to the next reload poll | -| **V63–V69** — PostgreSQL collector tables | `collect.pg_wait_stats`, `collect.pg_statement_stats`, `collect.pg_wraparound_stats`, `collect.pg_xmin_horizon`, `collect.pg_replication_slot_stats`, `collect.pg_autovacuum_stats`, and `collect.pg_io_stats`, each with its time index — one rung per PostgreSQL collector. Additive and view-less, exactly like V46/V47: a fresh store gets all seven from V1's generated schema, and these rungs are what an already-existing store gets. They add tables only, so a store that monitors no PostgreSQL target carries seven empty tables and nothing else changes | -| **V70** — monitored-server engine + port | `config.config_monitored_servers.engine` (`NOT NULL DEFAULT 'sqlserver'`) and `.port` (`NOT NULL DEFAULT 0` = the driver's default). The registry is authoritative for the server list once seeded, and these were the two `MonitoredServer` fields with no column — so a PostgreSQL target round-tripped as a SQL Server one and was connected to with `SqlConnection`. Every existing row means exactly what it meant before, and the SQL-Server-only writers keep inserting without naming either column | -| **V71** — PostgreSQL blocking edges | `collect.pg_blocking_edges` + its time index — the eighth PostgreSQL collector's store table. One row per (blocked, blocking) pair rather than a rendered tree, which is what lets the read layer compute root blocker, chain depth and fan-out in SQL instead of parsing a string. Additive and view-less exactly like V63–V69. **Sparse by design**: empty on a healthy instance, and because PostgreSQL has no engine-side blocked-process recorder, a gap means "not sampled" rather than "not blocked" — a count over this table measures how often blocking was *caught* | -| **V72** — Query Store plan map | `collect.query_store_plan_map` — `(server_id, database_name, plan_id)` → digest, so Query Store facts can reference plan XML they no longer carry once that content moves into the shared `query_plan_dim`. Plan XML was stored INLINE on `query_store_stats` at roughly 5x redundancy. Not a hypertable: one row per distinct plan per database, so it is dimension-shaped and pruned on `last_seen` rather than by `drop_chunks`. Its `last_seen` is load-bearing — the dimension GC sweeps on timestamps rather than counting references, so ending the re-shipping also ends the liveness signal that used to keep those dim rows alive | -| **V73** — PostgreSQL statement text | `collect.pg_statement_text` — `(server_id, queryid)` → statement text, refreshed hourly, so `get_pg_top_queries` returns something readable (#2219). `pg_statement_stats` stores no text because `showtext` is a real per-collection cost and normalized text is highly repetitive; but `queryid` is NOT stable across a major version upgrade, so without this the stored history joins to nothing after one — a list of integers that used to be your slowest queries, unrecoverable because the live view no longer holds the old ids. Text is INLINE rather than a `query_text_dim` digest: the dimension route needs the GC liveness interlock whose failure mode is silently missing text, and inline cannot dangle. Not a hypertable and not a collector table, exactly like V72 — a bespoke upsert path, pruned on `last_seen` with a margin that makes text OUTLIVE the statistics referencing it | -| **V76** — Query Store health | `collect.query_store_health` + its index + the `v_query_store_health` passthrough view — the #2319 per-database `sys.database_query_store_options` collector's store table: actual vs desired state (the cap-hit READ_ONLY transition and its readonly_reason), current vs max storage, cleanup thresholds, and the runtime-stats interval length. A fresh store gets the table from V1's generated schema; V76 is what an already-existing store gets | -| **V77** — Activity-driven plan fetch | Three strokes behind #2312's reshape of the Query Store plan/text fetch: `query_store_plan_map.digest` goes **nullable** (a plan whose XML the engine cannot persist gets a NULL-digest map row — the content-less marker that stops the probe re-selecting it forever), `query_store_text` gains `query_hash` (the Query Store reset detector: an id whose stored hash differs from the live one names a DIFFERENT statement now and its text refetches within one cycle), and the retired `planwm:`/`textwm:` watermark state rows are deleted wholesale. The fetch itself no longer walks the plan catalog by watermark — the cycle's collected rows name their plans, the store answers which are missing, and only those are fetched | - -All timestamps in the store are **naive-UTC** `timestamp` columns — the product-wide cross-store contract (Lite's DuckDB does the same). - -### Reading the store directly (plan XML is compressed) - -The store is deliberately queryable — it is documented PostgreSQL with named tables, and people build -panels and reports straight off it. One thing will surprise you if you do that: **execution-plan XML is -stored gzip-compressed**, and has been since v3.4.0. - -`collect.query_plan_dim` holds plan content once, keyed by a content digest, in one of two columns: - -| Column | Meaning | -|---|---| -| `query_plan_gz` (`bytea`) | The plan XML, **gzip-compressed** (magic bytes `1f 8b`). This is where new plans go. | -| `query_plan_xml` (`text`) | Uncompressed plan XML. Nullable since v3.4.0; only rows written by older builds still carry it. | - -So a consumer that reads only `query_plan_xml` silently returns nothing for anything collected by a -current build. **`query_plan_xml IS NULL` does not mean "no plan" — it means look at `query_plan_gz`.** - -Both apps and every MCP tool decompress client-side, so nothing in the product is affected; this note -exists because the change altered the contract for direct SQL consumers and the v3.4.0 release notes did -not say so. That omission is on us. - -**Getting the XML back.** PostgreSQL has no built-in gunzip for arbitrary `bytea`, so a plain-SQL -consumer cannot decompress in the database without an extension. Practical options, in the order most -people should try them: - -1. **Ask the product for the plan** rather than the store — `get_plan_xml` over MCP, or the Viewer's - plan surfaces. Both hand back decompressed XML and neither cares how it is stored. -2. **Decompress in your client.** Any language's gzip library reads the bytes directly. Python: - `gzip.decompress(row['query_plan_gz']).decode('utf-8')`. PowerShell: a `GZipStream` over a - `MemoryStream` of the bytes. C#: the same, which is exactly what the apps do. -3. **Ship a UDF into your own store** if your tooling is SQL-only (Grafana, a reporting view). A - `plpython3u` function works and has been used in the field, at the cost of an untrusted-language - extension in a monitoring database — weigh that against how much you need it. - -Why compressed at all: plan XML dominates store size, and gzip took a production dim table from 885 GB -of raw text to 64 GB — a 14x reduction. That is the tradeoff being made on your behalf. - -### TimescaleDB (Optional, Auto-Adopted) - -At startup, right after migration, the service attempts `CREATE EXTENSION IF NOT EXISTS timescaledb` and checks `pg_extension`: - -- **Present** — every collector table is converted to a hypertable (partitioned on its own time column into **1-day chunks**, existing rows migrated) and gets a compression policy: chunks older than **1 day** compress automatically (segmented by `server_id`), checked **hourly**. The hourly tick is passed explicitly because TimescaleDB's own default is **12 hours** for 1-day chunks — that is a second, separate wait *after* a chunk is already eligible, and on a field store it left the newest closed chunk (always the least-compressed data on disk) uncompressed for most of a day. Stores created before this shipped are retuned automatically on the next service start. The short intervals matter at the 1-minute collection cadence — a chunk cannot compress until it closes and then ages, so TimescaleDB's 7-day default left the store fully uncompressed for ~2 weeks (a near-idle 5-server fleet still reached ~1 GB in a couple of days); 1-day chunks + 1-day compress keep it compact (measured ~16.7x on perfmon, ~6.4x on the plan-XML-heavy query_stats). Compressed chunks stay fully queryable — this is Darling's archival tier, the centralized-store answer to Lite's Parquet archive. Everything is idempotent and re-converges on every service start; a table that fails conversion stays a plain table and keeps working. -- **Absent** — the service logs one Information line and runs in plain-PostgreSQL mode, which is a fully supported configuration, not a degraded one. - -`IF NOT EXISTS` short-circuits before privilege checks, so a store whose administrator pre-created the extension works for a service login that could never create it. - -### Background workers: sizing an unmanaged store, and what happens if you don't - -**This section is for bring-your-own PostgreSQL only.** In managed mode the service sizes these itself on every start and there is nothing to do. - -Every TimescaleDB policy — compression, retention, continuous-aggregate refresh — runs in a **background worker**, and a policy that cannot get a worker does not run. PostgreSQL's stock `max_worker_processes = 8` is far below what this store needs, so an unmanaged store left at the defaults silently does very little compressing. - -Managed mode derives the two settings from the live hypertable count, and an unmanaged store wants the same numbers: - -``` -timescaledb.max_background_workers = + 2 -max_worker_processes = 3 + timescaledb.max_background_workers + 8 -``` - -Today that is **51** and **62** for 49 hypertables (the 48 collector tables plus `collection_log`). The `+ 2` is not slack — it is exactly TimescaleDB's own two built-in jobs, `policy_telemetry` and `policy_job_stat_history_retention`, so a fully migrated store holds precisely one job per worker: - -```sql -SELECT proc_name, count(*) FROM timescaledb_information.jobs GROUP BY proc_name; -``` - -Both settings need a **server restart** (`max_worker_processes` is restart-only — a reload leaves the old value serving), and the hypertable count grows as collectors are added, so re-check it after a major upgrade rather than pinning 52/63 forever. - -**One store per cluster is the assumption.** `timescaledb.max_background_workers` is a **cluster-wide** pool shared by every database, while the derivation above is **per-store**. Managed mode puts one store on one cluster so the two coincide, but if you run **N Darling stores on one PostgreSQL cluster** — or share the cluster with any other TimescaleDB database — multiply both numbers by N. Each database with the extension loaded also permanently holds a scheduler slot out of that same pool, so the sharing starts before any policy fires. - -**What under-provisioning looks like.** The postmaster log (`pg.log`, or wherever your cluster logs) is where it shows up, in one of two shapes: - -``` -WARNING: failed to launch job 1042 "Columnstore Policy [1042]": out of background workers -WARNING: ... failed to start a background worker -``` - -The first means TimescaleDB's own pool is full; the second means PostgreSQL's is. Neither is fatal and neither corrupts anything — the job is skipped and retried on its next schedule, so **light contention is benign** and you may see a couple of these without any consequence. It matters at scale: when the shortfall is persistent rather than momentary, compression falls behind the 1-day policy and the store grows at its uncompressed rate (measured compression is ~16.7x on perfmon and ~6.4x on the plan-XML-heavy `query_stats`, so the gap is large), retention stops reclaiming chunks, and the jobs that keep losing the race are the ones whose backlog is worst. `timescaledb_information.job_stats` is the check that settles it — a healthy store shows successes with no failures: - -```sql -SELECT sum(total_runs), sum(total_successes), sum(total_failures) FROM timescaledb_information.job_stats; -``` - -### Retention - -A purge runs on the first sweep after startup and then daily, driven by the same shared per-collector horizons Lite uses: - -| Horizon | Tables | -|---|---| -| 7 days | `query_snapshots`, `waiting_tasks`, `running_jobs` | -| 30 days | Most collector tables (wait/query/procedure/Query Store stats, CPU, memory, file I/O, tempdb, perfmon, deadlocks, blocking, sessions, config snapshots), plus `collection_log` and `analysis_findings` | -| 90 days | `database_size_stats`, `index_object_stats`, `pvs_stats` | -| 365 days | `server_properties` | - -On plain PostgreSQL the purge is DELETE-based. With TimescaleDB it switches to `drop_chunks` — a metadata-only detach of whole expired chunks (rows inside a partially-expired chunk survive until the whole chunk ages out; up to ~1 day of grace at the 1-day chunk width), with a per-table DELETE fallback for any table that is not a hypertable. Failure-isolated per table: one stuck purge is logged and retried the next day without stopping the sweep. - -#### The rollup tiers, on a TimescaleDB store - -The table above is the **collector** horizon, and for three tables it is not the binding one. `query_stats`, `procedure_stats` and `query_store_stats` are rolled up into hourly and daily continuous aggregates, and a separate tiered policy drops their raw chunks at **4 days** — the aggregates hold the history past that point, and a read is routed to whichever tier covers the window it asks for. On a store without TimescaleDB none of this exists and the collector horizons above are the whole story. - -| Tier | Horizon | -|---|---| -| Raw `query_stats`, `procedure_stats`, `query_store_stats` | 4 days | -| Hourly **history** rollups | 90 days | -| Daily **history** rollups | kept indefinitely (no policy) | -| Baseline aggregates | 35 days | -| `query_store_stats_interval_hourly`, `query_store_stats_interval_daily` | 7 days, 10 days | - -Every one of these is visible in `timescaledb_information.jobs`, and the last row is the one worth knowing before you look: those two are **internal dedup plumbing, not history**. The corrected Query Store rollups are built from them, nothing reads them directly, and each horizon is sized only to outlive whatever gates on it — 7 days has to exceed raw's 4, and the 10-day layer has to outlive the 7-day one it consumes. So a horizon SHORTER than the tier above it is correct there and costs no history, which is the opposite of how it reads at a glance. The service's startup summary line names all of these for the same reason. - -No raw tier is ever dropped before the aggregate that preserves it has caught up: each policy is created paused, and arms itself only once its rollup demonstrably covers what the tier below holds. - -### Logs - -The service's PRIMARY log is a **rolling file** under `%ProgramData%\PerformanceMonitorDarling\logs\darling-service_yyyyMMdd.log` — every collector run line, connect edge, reload notice, warning, and error lands there (buffered writes, one file per day, 14-day retention, and a logging failure can never crash the service). Console runs write the same file plus console output. - -Warnings and errors also go to the **Windows Application event log** (source `PerformanceMonitor Darling`) — but only if that event source exists. Registering an event source requires elevation, and the recommended `NT SERVICE` virtual account cannot do it, so run the `New-EventLog` line in the install steps above (or any elevated run of the exe) once; without it, Windows silently drops the events and the file log is your only surface. Collection outcomes are also queryable in the store itself — `collection_log` records every collector run per server with status and timings, and the viewer's Collection Health tab renders exactly that. - -### The Viewer - -`PerformanceMonitor.Darling.Viewer.exe` is a WPF app that talks **only to the PostgreSQL store** — it never connects to your monitored SQL Servers. It reads the same `darling.json` the service uses, but only the `postgres` section, resolved in the same order (explicit path, then `DARLING_CONFIG`, then `darling.json` next to the binary) plus one viewer-only fallback: the parent directory, so the release zip's layout — viewer in a `viewer\` subfolder, `darling.json` beside the service exe — works with no setup. A viewer seat on **another machine** is set up by exporting that config folder from the service host — see [Connect a Remote Viewer](#connect-a-remote-viewer). If the file is missing it shows a hint instead of crashing. - -At startup the viewer writes **which of those rules won**, the absolute path it produced, and whether that file exists to `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log` — before it tries to read the file, so a missing or malformed one still says where it looked. Once the file loads it adds a non-secret summary of what it parsed (host, port, username, database, SSL mode, search path, whether the connection string was read verbatim or derived from `postgres.managed`, and the certificate — the value as written, the absolute path it resolves to, the folder a relative one was anchored to, and whether that file exists). Credentials are never written. The same block appears in the connection-failure window with a **Copy details** button — see [Troubleshooting](#troubleshooting). - -The layout mirrors the Lite desktop app: a left sidebar lists the servers from the `servers` registry the service maintains, and the top tab strip holds three fixed **aggregate tabs** — Overview, Recommendations, and Alerts — alongside a closable **per-server tab** for each server you open. Overview (the all-servers server-cards grid) and Alerts (the all-servers alert history) span every server; Recommendations has its own server selector, independent of the sidebar. **Double-click a server** in the sidebar — or **double-click its Overview card** — to open (or focus) its tab, and close it with the × on the tab header; an empty-state panel is shown until the store has at least one server. - -Each per-server tab has fourteen inner tabs: - -| Inner tab | Contents | -|---|---| -| **Overview** | Five correlated, X-axis-synced timeline lanes over the last 24 hours — CPU % (SQL Server vs SQL+other Total), total wait ms/sec, blocking + deadlocking, buffer pool MB, and file-I/O latency — each with a ±2σ baseline band and anomaly markers, all sharing one crosshair so a spike in one lane lines up against the others | -| **Wait Stats** | A searchable wait-type picker (poison + usual-suspect + `PAGELATCH_` defaults, checked-to-top, a 30-type selection guide) beside a per-**type** trend chart for the checked types over the last 24 hours, with a Wait Time (ms/sec) ↔ Avg Wait Time (ms/wait) metric toggle — the per-type companion to the Overview's single total-wait lane | -| **Queries** | Six sub-tabs over the last 24 hours — **Performance Trends** (a 2×2 of per-second trend charts: query duration, procedure duration, Query Store duration, execution count), **Active Queries** (the ~26-column filterable snapshot grid of captured running queries with a time-range slicer, a **Latest Snapshot** button that re-reads the newest stored capture, and per-row Estimated / Actual plan buttons that open the stored plan in the Plan Viewer), **Top Queries by Duration** (the full query-stats grid with in-grid bar cells for executions/CPU/duration/reads and a CPU-by-database breakdown), **Top Procedures by Duration**, **Query Store by Duration**, and **Query Heatmap** (query counts per 5-minute bin × per-execution magnitude bucket, by a chosen metric; right-click a cell to drill into Active Queries for that window) — the three grids each carry a time-range slicer (drag to narrow the window) and a shared **Compare** control that overlays the current window against a baseline period (yesterday, last week, or same day last week), flagging new and vanished queries | -| **Plan Viewer** | Hosts execution plans as closable sub-tabs (the shared plan-viewer control, the same one Lite and the Dashboard use). Right-click a **Top Queries** or **Query Store** row and choose **View Plan** to open the plan the service captured for it (`query_stats.query_plan_xml` / `query_store_stats.query_plan_text`); Top Queries rows also carry a **Query Plan** column whose Download button saves the stored plan as a `.sqlplan` file (enabled only when a plan was captured). Top Procedures and the blocking / deadlock reports deliberately do **not** surface a plan here — procedure plans aren't stored, and blocked-process / deadlock rows carry only a `sql_handle` (not plan XML); resolving either to a plan needs a live SQL connection the viewer never makes. "Get Actual Plan" (a live re-execution) is likewise out | -| **CPU** | Raw per-sample CPU utilization (SQL Server vs other processes) over the last 24 hours — every ring-buffer sample, full-bleed as two series; the Overview's CPU lane plots the same raw samples compactly (SQL vs SQL+other Total) with a baseline | -| **Memory** | Four sub-tabs over the last 24 hours — **Overview** (a summary strip of physical / SQL Server / target / buffer pool / plan cache / page-file memory plus the system memory state and model, over a Total-vs-Target-vs-Buffer-Pool memory trend with a memory-grants overlay), **Memory Clerks** (a searchable clerk-type picker — top-5 default, checked-to-top, clear-only-the-filtered — beside a per-clerk memory trend for the checked clerks with a non-buffer-pool total and top-clerk summary), **Memory Grants** (per-resource-pool grant sizing — available / granted / used MB — and activity — grantees / waiters / timeouts / forced grants), and **Memory Pressure Events** (hour-bucketed stacked bars of `RING_BUFFER_RESOURCE_MONITOR` pressure, SQL Server vs OS, medium vs severe) | -| **File I/O** | Two sub-tabs over the last 24 hours — **Latency** (per-file read and write latency, with a dashed queued-I/O overlay) and **Throughput** (per-file read and write MB/s) — the top 10 files by activity | -| **tempdb** | Three stacked charts over the last 24 hours — space usage (user / internal objects / version store), total allocated size, and per-file I/O latency | -| **Blocking** | Four sub-tabs over the last 24 hours — **Trends** (lock-wait rate, blocking incidents, deadlocks), **Current Waits** (waiting-task duration by wait type, blocked sessions by database), **Blocked Process Reports** (the full ~25-column filterable grid — XE reports preferred with the always-on DMV blocking snapshot merged in as fallback, each row badged with its source, a time-range slicer, per-row report-XML save, and long-block highlighting; double-click or right-click **View Block Chain** to reconstruct and draw the blocking chain the row belongs to), and **Deadlocks** (one filterable row per process parsed from each deadlock graph, a slicer, per-row graph-XML save; double-click or right-click **View Deadlock Graph** to draw the deadlock graph) | -| **Perfmon** | A searchable counter picker with the shared counter packs (General Throughput, Memory Pressure, CPU / Compilation, I/O Pressure, TempDB Pressure, Lock / Blocking) beside a per-counter delta trend for the checked counters (up to 12) over the last 24 hours | -| **Running Jobs** | Latest snapshot of currently-running SQL Agent jobs — start time, current vs average vs p95 duration, % of average, and a highlighted row when a job is running past its p95 (a store-derived banner appears when the service's login lacks msdb access) | -| **Configuration** | Four column-filterable snapshot grids of the server's latest capture — server configuration (`sys.configurations`), database configuration (28 columns of `sys.databases`), database-scoped configuration, and trace flags | -| **Daily Summary** | A one-row roll-up of the selected day (default today, UTC, with a date picker) — total wait time, the top wait type, distinct query count, deadlock / blocking-event / high-CPU-sample counts, collector errors, and an overall health band | -| **Collection Health** | Three sub-tabs — **Health Summary** (a 7-day per-collector roll-up: run / success / error counts, failure rate, average duration, last success / run / error, and a health band of HEALTHY / WARNING / STALE / FAILING / NEVER_RUN / NO_PERMISSIONS — double-click a collector to open its full run history), **Collection Log** (the recent run log with per-run SQL and store-write timings and row counts), and **Duration Trends** (a per-collector success-duration scatter) | - -The three aggregate tabs — **Overview** and **Alerts** span every server; **Recommendations** has its own server selector, independent of the sidebar: - -| Tab | Contents | -|---|---| -| **Overview** | A card per registered server (all servers, not the sidebar selection): server name + status dot, CPU (total non-idle with the SQL-only number alongside), memory, blocking and deadlock counts over the last hour, and last-collection time, each colour-banded (CPU ≥ 80% red / ≥ 50% amber / green; blocking and deadlocks red-or-amber when present) with a red **Offline** overlay. Status is derived from **collection freshness** — the newest `collection_log` age — rather than a live ping (the viewer never connects to the monitored servers): fresh is Online, older than twice the fastest collector's one-minute cadence is a Warning, and no recent collection is Offline. **Double-click a card** to open that server's tab. Refreshes every 30 seconds | -| **Recommendations** | The latest analysis run's findings for the tab's **own selected server** — a server selector independent of the sidebar, a Refresh button, and a status line showing the last analysis time — re-skinned to Lite's advise-only **card** design: a scrollable list of collapsible **incident** sections, each holding severity-banded cards (a severity badge, the affected `[database]`, the title, and the advice). Every card offers **Ask AI** (copies an MCP investigation prompt referencing `analyze_server` / `get_analysis_findings`); a card whose stored remediation carries a copy-paste statement also offers **Copy fix** (copies the suggested T-SQL). Advise-only — the viewer never applies anything, and there is no mute affordance here (alert muting lives on the Alerts surface). There is no in-app "Generate now": the service runs analysis on its own 30-minute cadence, so the status line surfaces the last analysis time instead | -| **Alerts** | The full alert history from `config_alert_log` across **all servers** (newest first, selectable time range), with a Server column and a Server filter. Double-click a row (or **View Details**) for a modal detail window showing the alert's stored detail and structured advice / remediation / drill-down from its dedup-fingerprint context. **Dismiss Selected / Dismiss All** hide alerts from the view (a durable `dismissed` flag on `config_alert_log`); column filters, Copy Cell/Row/All, and Export to CSV match Lite's grid. Right-click to **Mute This Alert** or **Mute Similar** (metric-only), and a **Manage Mute Rules** button opens the mute-rule editor | - -Only the visible tab loads (Lite's visible-only rule). The Alerts tab and the visible server tab's active inner tab refresh every 60 seconds; the Overview refreshes on its own faster 30-second timer (Lite's Overview cadence); and **Recommendations** refreshes on tab activation, its Refresh button, and its own server-selector change only, never on the timer — its findings change on the service's 30-minute analysis cadence, so a 60-second auto-refresh would be pointless churn (and would reset the incident expanders under the reader), matching Lite. - -The viewer is read-only over collected data, but it does perform a small set of **user-initiated writes** — and those go straight to the PostgreSQL store, which is the coordination point (the service honors them on its next read; there is no viewer-to-service channel). From the Alerts tab, creating a mute rule from an alert (**Mute This Alert** / **Mute Similar**) or adding, editing, toggling, deleting, or purging one via **Manage Mute Rules** writes `config_mute_rules` (a rule scopes to a server by name, exactly as Lite's mute rules do); and **dismissing alerts** sets the `dismissed` flag on `config_alert_log` so they drop out of the Alert History view (a single atomic UPDATE — Darling has no parquet archive tier, so there is no dismissed-archive sidecar). The viewer never writes collector data. - -### Restart Semantics - -The service is built to restart cleanly, any time: - -- **Delta continuity** — delta-based collectors (wait stats, file I/O, perfmon, memory grants) re-seed their baselines from the store at startup, so the first cycle after a restart produces real deltas instead of zeroes. -- **Alert no-re-fire** — edge-trigger watermarks and the failed-job watermark persist in `config_edge_trigger_watermarks`, and per-alert cooldowns re-seed from `config_alert_log`, so a restart does not replay alerts you already received. -- **Idempotent store setup** — migrations are versioned and skip what is already applied; TimescaleDB conversion and compression policies re-converge as no-ops. -- **Per-connect snapshots** — the on-connect config snapshot collectors run once per (re)connect, mirroring Lite's server-open behavior. -- Mute rules (`config_mute_rules`) load once at service startup — restart the service after adding rows. - -A monitored server that is down is retried every 60 seconds forever; a collector that errors is logged and retried at its next scheduled time; a mid-cycle connection-level failure forces a clean reconnect and re-probe. The loop never dies for one bad cycle. - ---- - -## Connect a Remote Viewer - -For the person sitting at a machine with **nothing installed on it**, whose only goal is looking at a Darling service that already runs somewhere else. Three steps, nothing to hand-edit. - -**The one service-side prerequisite.** The store has to be reachable from your LAN — a `postgres.network` block on the service host, which the `--configure-network` wizard writes for you. A store still on its loopback default accepts no remote viewer at all, and no amount of viewer-side configuration changes that. See [Store endpoint (viewer over the LAN)](#store-endpoint-viewer-over-the-lan) for that side; everything below assumes it is done. - -### 1. Export the handoff folder (on the service host) - -``` -PerformanceMonitor.Darling.Service.exe --export-viewer-config -``` - -It writes the viewer machine's **whole configuration folder** — connection string resolved, certificate copied, every field documented in place: - -``` -viewer-config\darling.json the complete viewer config: the resolved connection string and - "managed": false already set, every field explained in comments - IN the file -viewer-config\server.crt the store's TLS certificate, the file the connection pins -viewer-config\README.txt the same field reference in plain text, including the valid - "Root Certificate=" values and the one-line install instruction -``` - -The folder lands beside the service's own `darling.json` by default. Pass a directory to put it elsewhere (`--export-viewer-config D:\handoff`), and `--config ` if `darling.json` is not where the service would resolve it. - -**The exported `darling.json` contains a live database password** — that is what the viewer authenticates with. The verb says so before it writes, ACLs the file to SYSTEM + Administrators + the account running it + INTERACTIVE (the Viewer reads it interactively, the same posture as the admin/viewer credentials), and confirms the ACL took: if the secret is still readable by ordinary users it says so and exits non-zero. Copy the folder over a channel you trust and keep it ACL'd on the viewer machine. - -The verb refuses rather than clobbers: it will not export into the **service's own config directory** (that would overwrite the service's `darling.json` with the viewer's, destroying its servers, encrypted passwords and tokens), will not overwrite a file it did not write, and will not follow a junction or symlink. A destination it cannot use is named in the refusal. - -### 2. Copy the folder to the viewer machine - -Put the three files **next to `PerformanceMonitor.Darling.Viewer.exe`** — that works with nothing edited. (The Viewer ships in the same release zip as the service, in its `viewer\` subfolder; from source it is `dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release`.) - -To keep the folder somewhere else instead, point the `DARLING_CONFIG` environment variable at the exported `darling.json`. That works unedited too: a bare or relative `Root Certificate` resolves against **the folder holding `darling.json`**, so the `server.crt` beside it is found wherever you keep the folder ([#1970](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1970)). Keep the three files together and the folder can live anywhere. - -### 3. Start the Viewer - -That is the whole setup. Re-run the export after a credential or certificate rotation — the store's certificate regenerates when its bind IP changes — and copy the folder over again; it replaces its own previous output without ceremony. - -### If it does not connect - -The failure window carries a **Configuration this viewer used** block naming the `darling.json` it actually read, which rule picked it, and the host, port, username, database, SSL mode, search path and certificate path it parsed — with a **Copy details** button, and the same lines in `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log`. Read it before changing anything: it separates *the viewer read a different file than you edited* from *it read your file and a value in it is wrong*. It never contains a password. See [Troubleshooting](#troubleshooting) for the individual failures. - -### Manual configuration (fallback) - -Only for the case where you want the connection string itself — to paste into a config that already exists, or to check what the viewer will dial. The export above is the supported path; this one is the same values, assembled by hand. - -``` -PerformanceMonitor.Darling.Service.exe --print-viewer-connection -``` - -It decrypts the `network.role` credential and prints a paste-ready connection string plus the server certificate PEM. Every warning is printed **before** the payload, but the payload is still a **live database password on STDOUT** — redirect it to an ACL'd file or pipe it to the clipboard (`... --print-viewer-connection | clip`); do not leave it in shell scrollback, CI logs, or a screenshare. The minimal viewer `darling.json` it targets is bring-your-own mode with the string pasted in verbatim (the string is consumed as-is), and the emitted PEM saved where `Root Certificate` points: - -```json -{ - "postgres": { - "managed": false, - "connectionString": "Host=192.168.1.205;Port=5641;Username=viewer;Password=...;Database=darling;Search Path=collect,config,public;SSL Mode=VerifyFull;Root Certificate=server.crt" - } -} -``` - -`"managed": false` is not a typo next to the service's `"managed": true`: the flag says who **owns** the PostgreSQL, not who is connecting. A viewer left on `true` goes looking for a bundled local PostgreSQL that is not there. (The export sets it for you, which is the point.) - -**`Root Certificate=` — what the field accepts.** It is a path to the PEM the connection validates the store's certificate against, and under `SSL Mode=VerifyFull` it is what makes the check meaningful. A relative value anchors to **the folder holding the `darling.json` the viewer read**, never the process working directory, so how the Viewer was launched cannot change the answer: - -| Value | Resolves to | -|---|---| -| `server.crt` | that name in the folder holding `darling.json` — the exported layout, correct wherever the folder lives | -| `certs\server.crt` | same anchor, one level down | -| `C:\Darling\server.crt` | an absolute path, used exactly as written, for a certificate kept somewhere else | -| omitted | nothing viewer-side to pin against: the store's certificate must already chain to a root the machine trusts. A managed store's certificate is **self-signed**, so it never does — omitting the field there fails `VerifyFull` | - -**Where the certificate comes from.** In managed mode the service generates `server.crt` / `server.key` **beside the data directory** (`%ProgramData%\PerformanceMonitorDarling\pg\` unless you set `postgres.dataDirectory`), with an IP SAN for the `network.listen` address and a DNS SAN for the machine hostname. It **auto-regenerates if the bind IP changes**, so verify-full keeps working after a `listen` change — and every viewer must then re-copy the new certificate, because an old copy stops matching. To rotate on demand, delete the pair beside the data directory; the service regenerates it on its next start. - -**Bring-your-own PostgreSQL.** Darling generates no certificate — your PostgreSQL's TLS is yours to configure — so `Root Certificate` points at the PEM that signed **your** server's certificate (the CA certificate, or the server's own certificate if it is self-signed), exactly the file you would hand `psql` as `sslrootcert`. The same relative-path anchoring applies, so keeping it beside `darling.json` is still the simplest layout. - -**Plaintext at rest on the viewer machine.** However you get there, the connection string holds the role password in cleartext in that machine's `darling.json` (there is no client-side secret store yet). That is acceptable for the read-only `viewer` credential on a single-operator, ACL'd profile; if you use `role: "admin"`, treat that file as a secret and NTFS-ACL it to your account. DPAPI-encrypting the viewer's BYO connection string is future hardening, out of scope today. - ---- - -## Troubleshooting - -**"Cannot load configuration"** (critical, service idles) — no `darling.json` was found at the resolved path. The message names the path it tried; copy `darling.sample.json` there or point `DARLING_CONFIG` at your file. - -**"Configuration problem: ..."** (critical, service idles) — validation failed. The messages are literal and per-field, e.g. `postgres.connectionString is required.`, `servers must contain at least one entry.`, `server 'X': host is required.`, `server 'X': sql auth requires username.`, `server 'X': sql auth requires encryptedPassword (preferred; see --encrypt-password) or password.`, `server 'X': auth must be 'integrated' or 'sql'`. Fix the file and restart the service. - -**"Cannot reach or migrate the Postgres store"** (critical, service idles) — the store connection string is wrong, PostgreSQL is down/unreachable, or the login cannot create tables. Collection does not start until this succeeds; fix and restart. - -**"uses a plaintext password in darling.json"** (warning, every connect) — you set `"password"` instead of `"encryptedPassword"`. It works, but run `--encrypt-password` on the service machine and switch. - -**DPAPI decrypt fails after moving darling.json** — `encryptedPassword` blobs are machine-bound (DPAPI LocalMachine). Re-run `--encrypt-password` on the new machine. - -**"Failed to ensure XE sessions"** — the login lacks `ALTER ANY EVENT SESSION` (or the database-scoped equivalent on Azure SQL Database). Deadlock and blocked-process collection read zero rows until the sessions exist; grant the permission or have an administrator create/start `PerformanceMonitor_Deadlock` and `PerformanceMonitor_BlockedProcess`. "Already exists / already started" XE errors are logged as benign and mean the sessions are up. - -**Blocked-process reports empty** — the blocked-process threshold may still be 0. On AWS RDS set `blocked process threshold (s)` via a Parameter Group (the `sp_configure` bootstrap cannot run there); on Azure SQL Database the threshold is fixed at 20 seconds. Blocking stays visible either way through the always-on DMV blocking snapshot. - -**`PERMISSIONS` rows in `collection_log`** — that collector's reads were denied (SQL errors 229/297/300). Check the [permissions](#permissions-on-monitored-servers); the collector retries every cycle and recovers as soon as the grant lands. - -**"Skipping recently-failed-job check"** (info) — the login cannot read `msdb.dbo.sysjobs` / `sysjobhistory`, so failed-job alerts are skipped. Expected for minimal-privilege monitoring logins. If you want job alerts, add the direct msdb table `SELECT`s from the [permissions](#permissions-on-monitored-servers) section — **not** `SQLAgentReaderRole`, which gates the `sp_help_job*` procedures this product never calls and leaves the reads failing with error 229. - -**"TimescaleDB setup failed — continuing in plain-PostgreSQL mode"** (warning) — the extension exists but conversion hit a problem. Everything still works (DELETE-based retention, plain tables); conversion is retried on the next service start. - -**"out of background workers" / "failed to start a background worker" in the postmaster log, or the store keeps growing despite compression** — bring-your-own stores only: the cluster has fewer worker slots than the store has policies, so compression and retention jobs are being skipped. An occasional one is benign (the job retries on its next schedule); persistent ones mean the store is effectively uncompressed. Size the two settings and restart the server — see [Background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont), and multiply them if the cluster hosts more than one store. `timescaledb_information.job_stats` tells you whether jobs are actually succeeding. - -**"Why are there 40+ postgres.exe processes?"** — the count is three populations, and only one is client connections: (1) PostgreSQL's own system processes (postmaster, checkpointer, WAL/background writers, autovacuum, stats); (2) **TimescaleDB background workers** — the managed conf sizes `timescaledb.max_background_workers` to the hypertable count + 2 (≈52), and every RUNNING compression/retention policy job is its own process, so the count legitimately surges during checkpoint/compression waves and falls back when they finish; (3) client backends — the service's pools are capped at 24, the co-located viewer's at 10. Decompose it live with: `SELECT backend_type, count(*) FROM pg_stat_activity GROUP BY backend_type ORDER BY 2 DESC;` — and remember Windows charges the shared buffer segment to every attached process's working set, so per-process memory numbers cannot be summed. - -**query_store bursts every ~15 minutes** — two or three near-empty cycles, then one large one, is Query Store's own behavior, not a collector bug: the engine buffers in memory and flushes to its persisted tables on `DATA_FLUSH_INTERVAL_SECONDS` (default 900s), so the collector genuinely sees nothing new between flushes. Narrowing the collection interval will not smooth it. The per-database log lines show which database drove a burst. - -**MCP client cannot connect** — MCP defaults to off. Enable it live from the Viewer's Settings (the checkbox writes the control plane; the service starts the endpoint within seconds, no restart), or set `mcp.enabled: true` in `darling.json` for a file-seeded install. If the log says `Port 5152 is already in use — MCP server not started`, change `mcp.port`. The MCP server binds to `localhost` only unless you opt into a LAN endpoint (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan)); a remote client that gets 401 is missing or mismatching the required bearer token, and one that is refused before any response is outside the configured `allowFrom` CIDR. - -**Recommendations tab says no findings** — analysis runs every 30 minutes per server but only once the store holds at least 24 hours of collected data for that server; a fresh install simply has not earned findings yet. - -**The Viewer will not connect** — the failure window carries a **Configuration this viewer used** block naming the `darling.json` it read (and which rule picked it: an explicit command-line path, `DARLING_CONFIG`, beside the viewer, or the service root), plus the host, port, username, database, SSL mode, search path and certificate path it parsed. Read it before changing anything: the two faults it separates are *the viewer read a different file than you edited* and *it read your file and a value in it is wrong*. **Copy details** puts the whole block on the clipboard for a bug report, and the same lines are in `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log`. It never contains a password. - -**"Root Certificate ... exists: NO"** — with `SSL Mode=VerifyFull`, a **relative** `Root Certificate` path resolves against **the folder holding the `darling.json` the viewer read** — not the working directory, so how the viewer was launched no longer changes the answer. The diagnostics block prints that folder and the absolute path it actually opened; either put `server.crt` beside the config or make the `Root Certificate` value an absolute path. If you no longer have the certificate, re-run `--export-viewer-config` on the store host and copy the folder again (see [Connect a Remote Viewer](#connect-a-remote-viewer)) — it regenerates if the bind IP changes, so an old copy stops matching. - ---- - -## How It Runs (Reference) - -Fixed cadences, hardcoded on purpose: - -| What | Cadence | -|---|---| -| Collector sweep loop | Every 15 seconds (each collector runs when its own shared schedule is due — most every 1 minute, some every 5, sizes hourly, index stats daily) | -| Alert evaluation | Every 30 seconds per connected server (Lite's overview cadence) | -| Scheduled analysis | Every 30 minutes per server, 120-second budget, analyzing the last 4 hours; findings persist to `analysis_findings` and high-severity ones notify through the configured channels | -| Retention purge | First sweep after startup, then daily | -| Reconnect attempts | Every 60 seconds while a server is unreachable | - ---- - -## Managed Bundled PostgreSQL - -With `postgres.managed = true` (the sample's default), the service runs its own bundled PostgreSQL 18 + TimescaleDB and a from-zero install needs no database provisioning at all. Windows only, like every DPAPI surface here. - -```json -{ - "postgres": { - "managed": true, - "port": 5641, - "dataDirectory": null - } -} -``` - -**What first run does.** The service looks for `pg-runtime\pgsql\` beside its binary, extracting it from `pg-runtime.zip` when only the zip is present (deleting the extracted directory is therefore always safe — it self-heals). If the data directory has no cluster, it generates a 32-character random password, protects it with DPAPI LocalMachine into `pg-credential.dpapi` beside the data directory (credential first, so a crash mid-initdb never strands a cluster nobody can log into), then runs `initdb` with `scram-sha-256` auth, data checksums, and UTF8/C locale. A marker-guarded block appended to `postgresql.conf` preloads TimescaleDB, sets the port, and restricts listening to `127.0.0.1`; a second versioned block sizes background workers up for the per-hypertable compression jobs, DERIVED from the live hypertable count so it cannot go stale as collectors are added (`timescaledb.max_background_workers = hypertables + 2`, `max_worker_processes = 3 + that + 8` — today 52 and 63 for 50 hypertables; PostgreSQL's default of 8 workers cannot launch them); a third versioned block sizes memory from the host's physical RAM for the up-to-500-servers case (`shared_buffers = min(25% RAM, 1GB)`, `effective_cache_size = 75% RAM`, `maintenance_work_mem = min(max(5% RAM, 1536MB), 25% RAM, 2048MB)`, and a deliberately-modest per-connection `work_mem = clamp(RAM/512, 16MB, 64MB)` — on an 8 GB box that is `shared_buffers 1024MB` / `work_mem 16MB`; the stock 128 MB / 4 MB defaults are fine at small scale but bottleneck at fleet scale). Later blocks re-state single settings that field measurement moved: a fifth caps `shared_buffers` for the co-located store, a sixth turns on the log-rotation ring, and a seventh carries the `maintenance_work_mem` floor that TimescaleDB's compression sort runs on (measured at ~+70% compression throughput on a 16 GB-class host, plateauing by 1536 MB). `postgresql.conf` takes the LAST assignment of a setting, so these override without rewriting anything. Every append is re-checked on every start, so a crash between initdb and the append heals itself instead of silently degrading — and clusters initialized before a given block existed gain it on their next start (effective at the next PostgreSQL restart). Then `pg_ctl start`, `CREATE DATABASE darling`, and the normal startup path (migrations, TimescaleDB adoption — you should see `N/N collector table(s) are hypertables`, both numbers equal and equal to the collector count; a converted count BELOW the total means some table stayed plain and the line above it says which) continues exactly as in bring-your-own mode. The connection string is derived from the stored credential; the Viewer and the MCP host on the same machine derive it the same way, so nothing needs configuring there either. - -**Why scram and not trust, even loopback-only.** Trust auth would hand superuser to any local code that can open a loopback socket — every other local user, and network-capable-but-not-filesystem-capable attack primitives like SSRF from a co-hosted app. With scram the credential travels on the wire, failed attempts are auditable, and access is confined to what can read the DPAPI-protected credential file. `listen_addresses = '127.0.0.1'` keeps the server unreachable off the machine on top — unless you deliberately opt into a LAN endpoint (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan)), which reconciles `listen_addresses`, a `hostssl` pg_hba rule, and TLS on every start and is otherwise off. - -**Lifecycle.** On shutdown the service stops the server (`pg_ctl stop -m fast`) **only when it started it**. A server that was already running — an operator's own `pg_ctl`, or a postmaster that survived a service crash — is adopted for connections but never stopped: you'll see `already running … will not stop it` in the log, and the service keeps collecting into it. - -**The runtime zip.** `pg-runtime.zip` ships beside the service binary in packaged releases. Building from source, produce it once with `Darling\tools\fetch-pg-runtime.ps1` — it downloads the pinned EDB PostgreSQL 18 binaries and TimescaleDB, verifies their SHA256, prunes what the service doesn't need, and writes the zip to `Darling\artifacts\`; copy it next to the built service exe. - -**Server log.** The bundled server's own log is `pg.log` beside the data directory — that's where PostgreSQL explains a refused start; bootstrap errors in the service log quote its tail. - -## Security & Least-Privilege Roles - -The store is split into two schemas so that no consumer connects with more privilege than it needs: - -- **`collect`** — the collector hypertables (one per collector) plus the service-written, user-read metadata (`servers`, `collection_log`, `analysis_findings`, the `v_*` views). Read-only to everyone but the service. -- **`config`** — exactly the tables a human operator changes through the Viewer or MCP: `config_mute_rules`, `config_alert_log` (alert dismissals), `config_edge_trigger_watermarks`, and `analysis_muted`. - -Table names are unchanged — only their schema moved — and the shared SQL keeps using the bare, unqualified names, resolved through `search_path = collect, config, public` (set as the database default and carried on the managed connection strings). This is deliberate: Darling's SQL is byte-identical to Lite's DuckDB SQL, and re-qualifying it would fork that twin. - -**The roles.** The service still owns the store as the `darling` superuser (it does the DDL — migrations, hypertable conversion, retention). On top of that, **managed mode provisions three least-privilege login roles** (BYO provisions two — see below): - -| Role | Privileges | Used by | -|---|---|---| -| `darling` | superuser / owner | the service (collection, migration, provisioning) | -| `admin` | SELECT on both schemas — **including** the secret columns, which the Settings window reads — plus INSERT/UPDATE/DELETE on `config` only. No statement timeout | the Viewer, by default (`connectAs: "admin"`) | -| `viewer` | SELECT on all of `collect`, and on `config` **minus the secret columns** of `config_monitored_servers` / `config_command` / `config_notification` (carved fail-closed, below) + INSERT/UPDATE/DELETE on `config.custom_views` only (the web composer's saved views). Runs under `statement_timeout = 15s` | a locked-down Viewer (`connectAs: "viewer"`), and the web dashboard | -| `mcp` | `viewer`'s exact read surface + INSERT on `collect.analysis_findings` / `config.analysis_muted` + INSERT/UPDATE/DELETE on `config.custom_views` (the custom-view tools) + the alert-tuning writes (INSERT/UPDATE/DELETE on `config.config_mute_rules`, UPDATE on `config.config_alert_settings`, and the `config_service` reload-beacon columns) + the server-onboarding writes (INSERT/UPDATE/DELETE on `config.config_monitored_servers` — the credential column stays SELECT-carved, so it can WRITE a password blob but never READ one back) | the store identity the opt-in MCP **network** endpoint connects as (managed only); dormant until MCP is exposed on the LAN | - -`admin` cannot `DROP`, alter schema, touch `collect` data, or create objects — it can only do what the Viewer's mute-rule / alert-dismiss surfaces need. The `mcp` role is narrower still: it reads exactly what `viewer` reads (the secret config columns are carved out identically) and its writes are a small, enumerated set — the two analysis-table INSERTs (`analyze_server` + `mute_analysis_finding`), the single-table `config.custom_views` CRUD (the custom-view tools), the alert-tuning writes (`config.config_mute_rules` CRUD + a single-row `config.config_alert_settings` UPDATE, plus the two `config_service` beacon columns so a settings write's self-bump trigger can fire), and the server-onboarding writes (`config.config_monitored_servers` CRUD for `add_servers` / `remove_server` — its `config_monitored_servers` write fires the SAME `config_service` beacon trigger, already covered by that column grant) — so a token-holder on the network MCP endpoint can never reach the `config`-table service-credential pivot, the secret columns, or a service flag like `paused`. Even on `config_monitored_servers`, which it may write, the `encrypted_password` column stays in the fail-closed secret carve, so `mcp` can WRITE a credential blob (onboarding) but can never READ one back. `ALTER DEFAULT PRIVILEGES` means new collector tables auto-inherit SELECT for `admin`/`viewer`, so the model never drifts as collectors are added (every `mcp` write is an explicit single-table/single-column grant, deliberately not schema-wide). - -**Managed mode** provisions all of this automatically on every start (idempotent and self-healing), generating a per-role DPAPI-LocalMachine credential — `pg-admin-credential.dpapi`, `pg-viewer-credential.dpapi`, and `pg-mcp-credential.dpapi` beside the data directory, same posture as the owner's `pg-credential.dpapi`. Nothing to configure beyond `connectAs`. - -**Credential file protection.** DPAPI LocalMachine scope is deliberate (the service writes the credential, a *different* interactive user's Viewer reads it), which means the machine-bound blob is decryptable by anything that can *read* the file. So the credential files are locked down with an NTFS ACL that strips the inherited world-read `%ProgramData%` would give them: - -| File(s) | Readable by | -|---|---| -| `pg-credential.dpapi` (superuser), `pg-mcp-credential.dpapi` (the network MCP role) + the transient init pwfile | SYSTEM, Administrators, the service account — **not** interactive users | -| `pg-admin-credential.dpapi`, `pg-viewer-credential.dpapi` | the above **+ `NT AUTHORITY\INTERACTIVE`** (the operator's Viewer) | - -`pg-mcp-credential.dpapi` sits with the superuser (non-interactive) rather than with the Viewer's credentials because only the in-service MCP host reads it — never an interactive Viewer. - -The principal model assumes the **single-operator VM** this edition targets: `INTERACTIVE` == the operator, so the admin/viewer credentials are readable by the Viewer with zero configuration, while non-interactive local code (other services, sandboxed/SSRF socket primitives, scheduled tasks) and the superuser credential are excluded outright. On a shared machine where untrusted users log on interactively, tighten those two files to the specific operator account by hand. The service also refuses to trust a credential file that isn't owned by SYSTEM/Administrators/itself (closing a pre-plant attack), and regenerates an untrusted role credential. - -**A read-only (`viewer`) Viewer degrades gracefully.** It probes its own privileges on connect (`has_table_privilege`), so the mute-rule Add/Edit/Toggle/Delete/Purge buttons and the alert Dismiss / Dismiss All buttons are hidden or disabled, and any write that still slips through returns a clear "read-only connection" message instead of an error. - -**Bring-your-own PostgreSQL.** The schema split runs everywhere (it's a migration — the service applies it on startup and best-effort sets the database `search_path`; if your collection login can't `ALTER DATABASE`, run that one statement yourself as the owner). Role provisioning is managed-only, so for BYO you create the roles yourself, once, with the shipped script: - -``` -psql -h -U -d darling -f Darling/tools/provision-roles.sql -``` - -Edit the two password placeholders (and the database/owner names if yours differ) first. Then point a read-only Viewer's `connectionString` at the `viewer` role. **That script is the authoritative grant list for a BYO store** — it is what actually runs, the table above is its summary, and an `ALTER DEFAULT PRIVILEGES` in it means a store gaining collectors later needs no re-grant. Re-run it after a schema upgrade to cover new tables. **It creates two login roles — `admin` and `viewer`** — the two the Viewer connects as. Managed mode creates a third, `mcp`, but BYO deliberately does not: the MCP **network** endpoint (the only consumer of the `mcp` role) is managed-mode-only, and a BYO operator governs their own PostgreSQL's network exposure. If you expose MCP through your own reverse proxy against a BYO store, point it at whichever least-privilege role you choose (the `viewer` role covers the read tools; `analyze_server`'s finding persistence and `mute_analysis_finding` need INSERT on `collect.analysis_findings` / `config.analysis_muted`). - -## Opt-in Network Endpoints (LAN) - -By default all three network surfaces bind **loopback only** — the store to `127.0.0.1`, the MCP server and the web dashboard to `localhost` — exactly as they always have. Three optional, independent opt-ins let a remote viewer, MCP client, or browser on your **trusted LAN** reach them. This is a home-lab / trusted-subnet feature: **never expose any of these endpoints to the internet.** All three are **managed-mode only** (in bring-your-own mode your own PostgreSQL / reverse proxy governs exposure, and the config is ignored with a warning), and all three are **fail-closed** — any invalid or incomplete field degrades that endpoint back to loopback and logs a critical line rather than exposing it. Removing the config on the next restart closes the box again. - -### Guided setup (`--configure-network`) - -The fastest path is the interactive wizard — run it on the **service host**: - -``` -PerformanceMonitor.Darling.Service.exe --configure-network -``` - -It shows the current exposure (read from the service's own resolvers), then walks you through the **store**, **MCP**, the **web dashboard**, any comma combination (e.g. `1,3`), or all three at once (or a **disable** that removes all exposure). Every answer is validated **by delegation to the exact checks the running service fail-closes on**, so the wizard can never write a config the service would refuse — it re-prompts with the resolver's own reason. It generates the MCP bearer / web access tokens for you (DPAPI-protected; each plaintext is printed once, so save it then), edits `darling.json` **in place preserving every comment** behind a timestamped `darling.json.bak-` backup, prints the scoped firewall command(s), the `--export-viewer-config` handoff, and the web dashboard's browser login URL (`http://:/?token=...`), and offers to restart the service to apply. `install-darling.ps1 -Network` runs it automatically right after the install reaches Running. The manual field reference below documents exactly what it writes. - -### Firewall rules (`--configure-firewall`) - -The service runs as `NT SERVICE\PerformanceMonitor Darling`, an unprivileged virtual account that **cannot create Windows Firewall rules** — and should not be able to. So the rules are managed from the elevated install instead: `install-darling.ps1` runs `--configure-firewall` for you (before the first start, and again after `-Network`), and `uninstall-darling.ps1` removes them. Run it by hand after any edit to a `network` block: - -``` -PerformanceMonitor.Darling.Service.exe --configure-firewall -``` - -Run **elevated**. It reconciles all three scoped rules — store, MCP, web dashboard — against `darling.json` in one pass: it opens the port for every surface that really is exposed and removes the rule for every surface that is not, so it also cleans up after an exposure you turned back off. It is idempotent (safe on every upgrade) and reads **only** `darling.json`, so it works before the store has ever booted — unlike `--enable-mcp` / `--enable-web`, which write the control-plane store and need the service to have initialized it. - -"Really exposed" is decided by the same resolvers the running service fail-closes on, not by reading `listen` at face value. A `network` block the service would degrade to loopback — an unparseable `listen`, a missing or invalid `allowFrom`, an address family that disagrees, a missing token, BYO mode — gets **no open port**, and the verb tells you why. - -The running service never touches these rules. It **checks** them on start and logs what it finds: nothing at all for the normal loopback-only install, one INFO line when an exposed endpoint's rule is present, and one WARN naming the exact command when an exposed endpoint's rule is missing or when a loopback-only endpoint still has a stale rule open. It states each verdict once, not once per retry. - -### Headless enable/disable + firewall (`--enable-mcp` / `--enable-web`) - -On a box with no Viewer, two things are otherwise awkward: the `enabled` flags in the `mcp` / `web` blocks below are only a **first-run seed** — after the first run the store (`config.config_service.mcp_enabled` / `web_enabled`) is authoritative and is normally flipped only from the Viewer's Settings — and the service account (`NT SERVICE\PerformanceMonitor Darling`) **cannot open the firewall itself**. Four verbs, run on the **service host**, close both in one elevated action: - -``` -PerformanceMonitor.Darling.Service.exe --enable-mcp -PerformanceMonitor.Darling.Service.exe --disable-mcp -PerformanceMonitor.Darling.Service.exe --enable-web -PerformanceMonitor.Darling.Service.exe --disable-web -``` - -Each flips only its endpoint's **live store flag** with a targeted `config_service` write; the service **hot-reloads within one collection sweep — no restart.** If that endpoint's `network` block opts into LAN exposure (a non-loopback `listen`), the verb also reconciles that endpoint's **scoped, idempotent-by-name firewall rule**: **run elevated**, it opens (or, on `--disable-*`, removes) the rule; **run non-elevated**, the store toggle still succeeds and it prints the exact elevated firewall command to run by hand (a loopback-only endpoint needs no rule and says so). Managed-mode only, Windows only. So the headless bring-up is: write the `network` block (the wizard above or the manual reference below), then `--enable-mcp` / `--enable-web` from an **elevated** shell. - -### Verify it's actually reachable (and the two failures that look like bugs) - -Enabling an endpoint is **not** the same as reaching it, and both common failures leave the store flag reading `true`, so "it says enabled" is not proof. After `--enable-mcp` / `--enable-web`, verify on the **service host**: - -1. **The listener is on the LAN address, not loopback.** `Get-NetTCPConnection -State Listen | Where-Object LocalPort -eq 5152` (or `5153` for web) must show the box's LAN IP, e.g. `10.0.0.5:5152` — **not** only `::1` / `127.0.0.1`. *Enabled but still loopback-bound* is the single most common failure: the store flag is on, but the service loaded `darling.json` **before** the `network` block existed. The block is read **once at service start** — the enable toggle stops/starts the endpoint with the already-loaded config and does **not** reload the file. **Restart the service** (`Restart-Service 'PerformanceMonitor Darling'`) so it re-reads the block, then re-check the listener; after the restart run `--configure-firewall` **elevated** if the firewall rule is missing (the service account cannot create it, so the service only tells you it is missing). -2. **The scoped firewall rule exists and covers the client.** `Get-NetFirewallRule -DisplayName 'PerformanceMonitor Darling MCP (port 5152)'` (or `... Web (port 5153)`) should be `Enabled=True, Action=Allow`, scoped to the `network.allowFrom` CIDR. If it is absent, the service's own start-up log already says so and names the command; `--configure-firewall` elevated is the one-step fix. Reading rules needs no elevation, so this check works from any shell. - -Then from the **client** host: - -3. **Connect to the box's LAN IP, never `localhost`.** Use `http://:5152/`. `localhost` / `127.0.0.1` only resolves *on the box itself*, so an off-box MCP client pointed at localhost fails silently — this is the number-one "MCP won't connect" cause. Send `Authorization: Bearer ` (the `network.token`), and do a **fresh** `initialize` + `tools/list` rather than trusting a cached tool list from a previous version. - -**After a reinstall:** the installer replaces binaries but does **not** touch `darling.json` (the zip ships only `darling.sample.json`) or the store, so the `network` block and both live flags survive the upgrade — and the reinstall restarts the service, which re-reads the block. If MCP stops connecting afterward it is almost always failure 3 (the client pointed at `localhost`) or a missing firewall rule, **not** lost config: run `--configure-firewall` **elevated** to re-open the rule if check 2 comes up empty (the installer already does this, so an in-place upgrade normally leaves the rules correct). A stale loopback bind is unlikely after a restart unless the block itself is invalid, in which case the endpoint fail-closes to loopback and logs a critical line saying why — fix the block and restart again. A full `--configure-network` re-run is only needed if the `network` block itself is gone. - -### Store endpoint (viewer over the LAN) - -Add a `network` block to `postgres` (managed mode): - -```json -"postgres": { - "managed": true, - "port": 5641, - "network": { - "listen": "192.168.1.205", - "allowFrom": "192.168.1.0/24", - "role": "viewer" - } -} -``` - -On every start the service reconciles this against the live cluster: it adds the bind IP to `listen_addresses`, generates a self-signed TLS certificate (`server.crt` / `server.key` beside the data directory, with both an IP SAN for `listen` and a DNS SAN for the machine hostname), writes a marked `hostssl darling scram-sha-256` rule into `pg_hba.conf` and reloads, and **checks** (never creates — see [Firewall rules](#firewall-rules---configure-firewall)) that the store's scoped firewall rule matches. - -- **`role`** — the pg_hba login role the rule names: `"viewer"` (default, **read-only** — the secure default, covering a laptop reading every dashboard, chart, and finding) or `"admin"` (full remote **writes**; the service logs a warning because `admin` holds the `config_command` / `config_monitored_servers` / `config_notification` service-credential pivot). Never the superuser. This is **distinct from `postgres.connectAs`** (the *local* VM viewer's loopback role, default `admin`): `network.role` is the *remote* role and defaults to `viewer`, so the two have opposite defaults — the local seat is writable, the remote seat is read-only, unless you say otherwise. -- **TLS is verify-full, not `require`.** Because Darling generates the cert, the client can pin it, so the connection string below uses `SSL Mode=VerifyFull` — which actually defends against an on-path MITM (`require` verifies nothing). The store's network pg_hba line is `hostssl`, so a non-TLS network client is refused. -- **The firewall is defense-in-depth, not the boundary** — pg_hba + TLS are. `--configure-firewall` (elevated) creates the store's scoped rule for you along with the other two; the equivalent by hand is: - - ``` - New-NetFirewallRule -DisplayName "PerformanceMonitor Darling store (port 5641)" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5641 -RemoteAddress 192.168.1.0/24 - ``` - -**That is the service side. The viewer side is [Connect a Remote Viewer](#connect-a-remote-viewer)** — `--export-viewer-config` on this host writes the viewer machine's whole configuration folder (config, certificate, and a plain-text field reference), and that section covers copying it over, the certificate's placement and rotation, and the manual `--print-viewer-connection` fallback. - -### MCP endpoint (assistant over the LAN) - -Add a `network` block to `mcp` (managed mode; `mcp.enabled` must be `true`): - -```json -"mcp": { - "enabled": true, - "port": 5152, - "network": { - "listen": "192.168.1.205", - "allowFrom": "192.168.1.0/24", - "encryptedToken": "" - } -} -``` - -When `listen` is a network address **and** a token is present **and** `allowFrom` is a valid CIDR, the MCP host binds that interface behind two gates: a **required bearer token** (checked first, constant-time, no loopback exemption) and an **in-app CIDR check** on the remote address (loopback is always allowed, so local clients keep working). Any missing precondition keeps MCP loopback-only. Prefer `encryptedToken` (a DPAPI blob from `--encrypt-password`); a plaintext `token` works for dev but is warned. Set the same scoped firewall rule for the MCP port: - -``` -New-NetFirewallRule -DisplayName "Darling MCP" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5152 -RemoteAddress 192.168.1.0/24 -``` - -**What a token-holder can — and cannot — do.** Start with the boundary: **no MCP tool runs SQL an AI client wrote against your monitored servers.** No such tool exists, and a stored custom view cannot become one either — a composed query names only `collect.*` collector tables in the monitoring store. The only live contact with a monitored SQL Server is `analyze_server`'s plan fetch and `add_servers`' one-time connection probe, and both run the product's own fixed, read-only queries under the same least-privilege monitoring login the collectors use — the ceiling on what they can see is the ceiling you granted that login, and it has no write grants to hit. Everything else answers from the monitoring store. - -What the token does gate is the monitor's own configuration and collected data: the entire read surface, `analyze_server`, the Custom Views tools (create / modify / delete the saved dashboards and notebooks in `config.custom_views`), the alert-tuning tools (`update_alert_settings` / `create_mute_rule` / `delete_mute_rule`), and the server-onboarding tools (`add_servers` / `remove_server`), which edit the monitored-server registry in `config.config_monitored_servers` — including storing a SQL-auth credential for a server they add. The store-side identity is still the least-privilege `mcp` role: read, the two analysis-table INSERTs, INSERT/UPDATE/DELETE on the single `config.custom_views` table (the same narrow write the web composer's `viewer` role has), the narrow alert-config writes (`config.config_mute_rules` CRUD + a single-row `config.config_alert_settings` UPDATE, plus the `config_service` reload-beacon columns), and the single-table `config.config_monitored_servers` CRUD. So a token-holder can read everything collected, trigger analysis, author custom views, tune alerting, and onboard/offboard servers — and can never reach the `config_command` service-credential pivot, the carved secret columns (SMTP/webhook credentials, and the monitored-server `encrypted_password` blob it can WRITE during onboarding but never READ back, all included), or a service flag like `paused`. Custom-view JSON and alert config carry no secrets. Guard the token like the keys to your monitoring configuration — that is what it opens; your SQL Servers are not behind it. - -**`add_servers` carries a credential in its request.** A SQL-auth `password` rides the request JSON; the service DPAPI-encrypts it at rest and never returns it, but on the wire it is only as protected as the endpoint — the same plaintext HTTP the token rides. On a segment you do not fully trust, front the MCP port with the TLS reverse proxy below, and prefer Windows/integrated auth for onboarded servers where you can — then no per-server secret crosses the wire at all. - -**MCP has no TLS — the MITM control is a TLS reverse proxy.** A self-signed cert breaks real MCP clients, so the MCP endpoint is plain HTTP and the bearer token travels **cleartext on the segment**; an active on-path attacker (ARP spoof, rogue DHCP, compromised switch) could capture and replay it. The in-app CIDR bounds *who can route to* the port; it does **not** protect the wire. If your segment is not fully trusted, put a **TLS-terminating reverse proxy** in front of the MCP port and point clients at that — the named MITM control for this endpoint. (The store endpoint needs no such proxy: it has verify-full TLS built in.) - -### Web endpoint (browser over the LAN) - -Add a `network` block to `web` (managed mode; `web.enabled` must be `true`): - -```json -"web": { - "enabled": true, - "port": 5153, - "network": { - "listen": "192.168.1.205", - "allowFrom": "192.168.1.0/24", - "encryptedToken": "" - } -} -``` - -When `listen` is a network address **and** a token is present **and** `allowFrom` is a valid CIDR, the web host binds that interface behind two gates: an **in-app CIDR check** on the remote address and an **access token**. A browser presents the token ONCE via `?token=` (open `http://192.168.1.205:5153/`, then paste it into the minimal login form, or append `?token=...` directly); the host validates it constant-time, sets an **HMAC-signed, HttpOnly, SameSite=Strict session cookie**, and 302-redirects to strip the token from the URL so it never lingers in history or a Referer header. Subsequent requests ride the cookie. **Loopback is always allowed tokenless, even while exposed** — unlike MCP, the read-only dashboard has no loopback-token requirement. An out-of-CIDR request is refused with **403**. The cookie signing key is per-process, so a service restart invalidates open sessions (just re-present the token). Prefer `encryptedToken` (a DPAPI blob from `--encrypt-password`); a plaintext `token` works for dev but is warned. Set the same scoped firewall rule for the web port: - -``` -New-NetFirewallRule -DisplayName "Darling Web" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5153 -RemoteAddress 192.168.1.0/24 -``` - -**What a web token can reach.** The web dashboard is **read-only over the collected store** — it connects as the least-privilege `viewer` role and hosts no write paths and no live-server queries (no `analyze_server`, no plan re-execution). A token-holder can view everything collected, change nothing, and reach no monitored server — which is why loopback stays tokenless here while MCP's does not. - -**Web has no TLS either — same reverse-proxy control as MCP.** The token/cookie travels cleartext on the segment, so the in-app CIDR bounds *who can route to* the port but does not protect the wire. On an untrusted segment, put a TLS-terminating reverse proxy in front of the web port. +# Performance Monitor Darling — Headless Edition + +Darling is the headless, centralized edition of Performance Monitor: a 24/7 Windows service that collects from your SQL Servers into a central PostgreSQL (optionally TimescaleDB) store, plus a detached desktop viewer that reads that store. No desktop app has to stay open for collection to happen, and every viewer seat reads the same central data. + +It runs the **same monitoring brain as the Lite edition** — one shared codebase, two storage engines: + +- `PerformanceMonitor.Collectors` owns all 48 collector definitions — 41 for SQL Server and 7 for PostgreSQL: the exact query sent to monitored servers, the result-row mappings, the delta rules, the default cadences and retention horizons, and the ignored-wait-types list. Lite writes those rows to DuckDB; Darling writes the same rows to PostgreSQL via binary COPY. Each definition declares which engine it targets, and a collector never runs against the other one — see [PostgreSQL targets](#postgresql-targets). +- `PerformanceMonitor.Alerting` owns the shared alert engine — the same thresholds, edge-trigger gates, cooldowns, and dedup fingerprints Lite uses. +- The analysis/recommendations pipeline (the same inference engine behind both apps' Recommendations tabs and the `analyze_server` MCP tool) runs on a schedule inside the service. + +A collector, alert, or analysis change lands once in the shared libraries and both editions get it. A Darling install monitoring a server even derives the **same `server_id`** Lite would for that server, because the identity rule (`host[:database][:RO]`, hashed) is shared too. + +> **Status: in development.** Darling builds and runs from source (it is wired into the solution and CI), but is not yet packaged into the signed release artifacts. Expect the surface documented here to grow. + +--- + +## When to Choose Darling vs. Lite + +| | **Lite** | **Darling** | +|---|---|---| +| Collection runs | While the desktop app is open (or in the tray) | 24/7 as a Windows service | +| Data lives | Locally per seat (DuckDB + Parquet) | Centrally (PostgreSQL / TimescaleDB) | +| Execution plans | Not stored (fetched live when you view a query) | Captured and stored, TOAST-compressed (`capturePlans`, default on) | +| Viewers | The app is the viewer | Any number of viewer seats read the central store | +| Setup | Download and run | Provision PostgreSQL, edit `darling.json`, install the service | +| Best for | Quick triage, consultants, a handful of servers | Always-on team monitoring, larger estates, one shared store | +| Configuration | Settings UI | One JSON file (no UI) | + +Nothing is installed on the monitored SQL Servers by either edition beyond two lightweight Extended Events ring-buffer sessions and, when it is unset, a one-time `blocked process threshold` bootstrap (see [What the Service Does on Monitored Servers](#what-the-service-does-on-monitored-servers)). + +--- + +## Quick Start + +> **First time, on a box with nothing on it?** [**`docs/uat-onboarding.md`**](../docs/uat-onboarding.md) is the ordered procedure from a downloaded zip to a running service and whichever of the three surfaces you need — the WPF viewer, the web dashboard, or MCP — with a proof point at every step. This section, and the rest of this document, is the reference it links back into. + +### Prerequisites + +- **Windows** for the service host (Windows-service lifetime, DPAPI password protection) and for the viewer (WPF). Monitored servers can be SQL Server 2016–2025, Azure SQL Managed Instance, AWS RDS for SQL Server, or Azure SQL Database. +- **A PostgreSQL store — bundled or your own.** In managed mode (the shipped default, see [Managed Bundled PostgreSQL](#managed-bundled-postgresql)) the service runs its own bundled PostgreSQL 18 + TimescaleDB and no database provisioning is needed. To bring your own instead, PostgreSQL 16 or newer is recommended (developed and validated against PostgreSQL 18) with a database and a login the service can create tables in — and if that store has TimescaleDB, size its background workers before you rely on compression, because the stock PostgreSQL defaults cannot run the policies (see [Background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont)). +- **TimescaleDB is optional and auto-adopted.** If the extension is installed (or pre-created by an administrator) in the store database, the service detects it at startup and automatically converts the collector tables to hypertables with compression; without it, the service runs in plain-PostgreSQL mode, which is fully supported. No configuration flag either way. +- **Two .NET 10 runtimes on the host**, from . Both + shipped binaries are framework-dependent, and a stock Windows Server image has neither: + - **ASP.NET Core Runtime 10** for the service. Required unconditionally — the MCP package brings the + ASP.NET Core framework reference in transitively, so it is needed whether or not you ever enable MCP + or the web dashboard. + - **.NET Desktop Runtime 10** for the viewer (WPF). Not needed on a headless collector host, and not + needed for a remote seat installed from the viewer's own `Setup.exe`, which is self-contained. + + `install-darling.ps1` checks both before it installs anything: it refuses when ASP.NET Core is missing + and warns when the Desktop runtime is. The .NET 10 SDK covers both if you are building from source. + +Build from the repository root: + +``` +dotnet build Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release +``` + +``` +dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release +``` + +### Configure darling.json + +The service reads one JSON file. It resolves the path in this order: + +1. An explicit path (when a component is handed one) +2. The `DARLING_CONFIG` environment variable +3. `darling.json` next to the service binary + +Copy the shipped `darling.sample.json` (it lands next to the built binary) to `darling.json` and edit. Comments and trailing commas are allowed; property names are case-insensitive. + +Minimal working example — one server, integrated auth, bring-your-own PostgreSQL. (With the bundled store instead, replace the `postgres` block with `"postgres": { "managed": true }` and skip provisioning entirely — see [Managed Bundled PostgreSQL](#managed-bundled-postgresql).) + +```json +{ + "postgres": { + "connectionString": "Host=localhost;Port=5432;Username=darling;Database=darling" + }, + "servers": [ + { + "name": "SQL2022", + "host": "SQL2022", + "auth": "integrated", + "excludedDatabases": [] + } + ] +} +``` + +**Integrated auth (recommended).** The service connects to monitored servers as the Windows account the service runs under — there is no separate Windows credential to configure. Grant that account the [permissions below](#permissions-on-monitored-servers). The default install's virtual service account reaches *remote* servers as the collector machine's computer account (`DOMAIN\$`), so for integrated auth you will usually [run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) instead. + +**SQL auth.** Set `"auth": "sql"`, a `username`, and an `encryptedPassword` produced by the `--encrypt-password` verb: + +``` +PerformanceMonitor.Darling.Service.exe --encrypt-password +``` + +It prompts for the password on stdin (so the plaintext never lands in your shell history) and prints a base64 DPAPI blob. Paste that blob into the server's `"encryptedPassword"`. The blob is protected with **DPAPI LocalMachine scope**, so an administrator can encrypt it interactively and the service account can decrypt it later on the same machine — but it is machine-bound: run `--encrypt-password` **on the machine that will run the service**, and re-encrypt if you move `darling.json` to another machine. A plaintext `"password"` also works as a dev convenience, but the service logs a warning every time it is used. The same slot also takes an **`env:NAME` or `file:/path` reference** (#1804): the service reads the named environment variable or the file's (trimmed) contents at connect time, nothing secret lands in `darling.json`, and no warning is logged — the supported shape on non-Windows hosts, and compose-`secrets:`-friendly everywhere. A missing or empty reference target is a configuration error naming both the setting and the target, never a silent empty password. + +**excludedDatabases** (per server) removes databases from collection: per-database collectors skip them and the exclusion is spliced into the collector queries — the same filter Lite applies. There is a second, separate `alerts.excludedDatabases` list that excludes databases from blocking/deadlock/long-running-query **alert evaluation** without affecting collection. + +### Validate the Config (Pre-flight) + +Before installing the service, check that `darling.json` is well-formed and that every monitored server is reachable with the configured credentials: + +``` +PerformanceMonitor.Darling.Service.exe --test-connection +``` + +(`--validate-config` is an alias.) It validates the file, then connects to and probes each server, printing a `[PASS]`/`[FAIL]` line per server (SQL major version, engine edition, and whether the account has msdb access for failed-job alerts). It exits `0` only when the file is valid **and** every server is reachable, so it doubles as a deployment gate. + +A PostgreSQL target reports what matters there instead — version, writer or reader, Aurora or not, and **how many of the PostgreSQL collectors will actually run against it**, naming the ones that will not: + +``` + [PASS] aurora-writer: PostgreSQL 17 (server_version_num 170007), writer, Aurora — all 12 PostgreSQL collectors apply + [PASS] aurora-reader: PostgreSQL 17 (server_version_num 170007), reader (in recovery), Aurora — 9 of 12 PostgreSQL collectors apply (skipped: pg_autovacuum_stats, pg_index_usage_stats, pg_table_bloat_stats) + [PASS] selfhosted: PostgreSQL 15 (server_version_num 150012), reader (in recovery), not Aurora — 6 of 12 PostgreSQL collectors apply (skipped: pg_autovacuum_stats, pg_index_usage_stats, pg_io_stats, pg_statement_stats, pg_table_bloat_stats, pg_wait_stats) +``` + +That count comes from the same [engine and version gate](#postgresql-targets) the collector runner uses, not a separate list, so it is the real answer rather than an estimate — and it is the answer at *pre-flight*, before an empty table has to be explained weeks later. Add an explicit config path as a second argument if `darling.json` is not next to the exe and `DARLING_CONFIG` is not set. This is the same probe the Viewer's **Test Connection** button runs through the service. + +One identity caveat: the verb connects as **you**, the console user — not as the service account. For `"auth": "integrated"` servers a `[PASS]` proves the server is reachable and the config is well-formed, but the grants that matter at runtime are the *service account's*: the per-server connect lines in the service log are the real proof (see [Run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa)). + +### Run It — Console Mode + +The same executable serves interactive debugging and service installation; the Windows-service lifetime is a no-op when run from a console. + +``` +Darling\PerformanceMonitor.Darling.Service\bin\Release\net10.0\PerformanceMonitor.Darling.Service.exe +``` + +Watch the log output: you should see the config load (`Loaded configuration from ...`), the store migrate (`Postgres store ready (schema v44, ...)` — the number is whatever the current migration count is), the TimescaleDB detection result, per-server connects, and then per-collector run lines with row counts. + +### Run on Linux (Docker Compose or systemd) {#1804} + +The service is cross-platform .NET; only the **bundled zero-admin store** and DPAPI are Windows-specific. On Linux you pair the service with the official TimescaleDB image (compose, the recommended shape) or point it at PostgreSQL you already run (systemd), keeping `postgres.managed = false` either way. The Viewer stays a Windows desktop app — Linux hosts read the **web dashboard**, which the container exposes. + +**Compose (the whole stack as one deployment)** — everything lives in [`Darling/compose/`](compose/): + +```bash +cd Darling/compose +cp darling.sample.json darling.json # edit: servers, alerting, tokens +# one secret per file — see secrets/README.md for the exact list +docker compose up -d +``` + +Web dashboard on `http://:5153` behind its token, MCP (if enabled) on `:5152` behind its bearer token. The port mappings are the exposure boundary: the container-aware bind gate honors `web.network`/`mcp.network` under `managed = false` **inside a container only**, and the tokens are still mandatory. Three rules worth knowing before they bite: + +- **Nothing secret goes in darling.json.** Every secret slot — the whole `postgres.connectionString`, server `password`s, `smtp.password`, the tokens — takes an `env:NAME` or `file:/run/secrets/` reference. The compose file mounts each secret from `secrets/`. +- **Start with a fresh store volume per deployment.** The control plane is store-authoritative after the first seed, so a reused volume's enable toggles override darling.json — by design. +- **File permissions are yours on Linux.** The Windows build locks config/credentials down with ACLs; here the container boundary is the isolation, and the `secrets/` directory should be `chmod 700` with `600` files (the systemd shape should do the same for `darling.json` itself). + +**systemd + bring-your-own PostgreSQL** — download `PerformanceMonitorDarling-linux-x64-*.tar.gz` from the release, extract to `/opt/darling`, point `DARLING_CONFIG` at your config (connection string to your own PostgreSQL 15+ with TimescaleDB; the service degrades gracefully without TimescaleDB), and run `dotnet PerformanceMonitor.Darling.Service.dll` under a unit like: + +```ini +[Unit] +Description=PerformanceMonitor Darling +After=network-online.target + +[Service] +ExecStart=/usr/bin/dotnet /opt/darling/PerformanceMonitor.Darling.Service.dll +Environment=DARLING_CONFIG=/etc/darling/darling.json +User=darling +Restart=on-failure + +[Install] +WantedBy=multi-user.target +``` + +Use the same `env:`/`file:` secret references (systemd `LoadCredential=` pairs naturally with `file:`), and note `Microsoft.Data.SqlClient` needs `libgssapi-krb5-2` installed (`apt-get install libgssapi-krb5-2`) — the container image carries it already. + +### Install as a Windows Service + +**Scripted (recommended):** the packaged zips ship `install-darling.ps1` beside the service exe. Extract the zip to its final location (e.g. `C:\PerformanceMonitorDarling`), then from an elevated PowerShell in that folder run `.\install-darling.ps1`. It checks the install location and refuses anywhere the service could not read itself (see [below](#the-install-location-has-to-be-machine-scoped)), checks for `darling.json` (copying the sample and stopping for you to edit it on first run), runs the `--test-connection` pre-flight, registers the Event Log source, creates the service under the virtual account (or upgrades an existing install's binPath in place, preserving config/store/credentials), starts it, and creates Desktop + Start Menu **Darling Viewer** shortcuts (pin to taskbar from the Start Menu entry — Windows does not allow programmatic pinning). `uninstall-darling.ps1` reverses it, deliberately leaving the store/config in place unless you pass `-PurgeData`. + +#### Upgrading an existing install + +`install-darling.ps1` registers a service; it does not lay a new build over an old one. That step — stop, back up, copy, start, verify — is `upgrade-darling.ps1`, shipped in the same zip. + +> **`upgrade-darling.ps1` does not exist in 3.5.0 or earlier.** It was added after 3.5.0 was tagged, so a build up to and including 3.5.0 does not contain it and neither does its zip. If you are upgrading FROM one of those, see [upgrading from a build that predates the script](#upgrading-from-a-build-that-predates-the-script) below and use the manual procedure — the steps in this section describe a script you will not have. + +Extract the new zip to a **staging** folder and run *its* copy: + +```powershell +Expand-Archive PerformanceMonitorDarling-3.5.1.zip -DestinationPath C:\staging\3.5.1 +C:\staging\3.5.1\upgrade-darling.ps1 -Source C:\staging\3.5.1 +``` + +It resolves the install directory from the registered service, verifies the zip's SHA256 when you point it at one (`-Source ...\PerformanceMonitorDarling-3.5.1.zip`, checked against `-Sha256` or a `SHA256SUMS.txt` beside it), backs the install root's files up to `_rollback_manual_`, prunes the backups past the newest `-KeepRollbacks` (3), lays the new build down, confirms `darling.json` is byte-identical, and starts the service. Re-running after a failure is safe and is the intended recovery: a backup taken in the last `-BackupWindowMinutes` (60) is reused rather than replaced, so a re-run cannot overwrite the good pre-upgrade copy with a copy of a half-upgraded tree. + +**It never kills a process**, and it checks for them twice. Before stopping anything it names processes running out of the install tree that a service stop will *not* close — your own `psql.exe`, a shell sitting in the folder, a Darling Viewer you left open — and refuses, costing nothing but a re-run. After the service is down it checks again with no exclusions; anything still there is usually a postmaster that outlived the stop, and that is exactly what must not be killed (the bundled PostgreSQL lives under `pg-runtime` and killing it takes the store down). Give it a few seconds and re-run. + +**A rollback backup holds the install root's *files*, not `viewer\`, `wwwroot\`, `runtimes\` or `pg-runtime\`** — that is what keeps one to ~120 MB instead of ~1 GB. It matters in one case: if a copy dies partway, those subdirectories can be left mixed old-and-new, and a full revert is the *previous version's zip re-extracted* followed by the backup's files over the top. The script says so at the point of failure. + +It **refuses to run from the install directory itself** — the copy would overwrite the script PowerShell is reading — which is why the staging folder above is not optional. + +**Files the new build no longer ships.** The copy is an *overlay*: it writes what the new build ships and deletes nothing else, so a dependency that went away, an assembly that changed name, or a whole `runtimes\\lib\\` subtree stranded by a target-framework move stays in the tree forever — and those are directories .NET probes for assemblies ([#2529](https://github.com/erikdarlingdata/PerformanceMonitor/issues/2529)). It has happened: the Lite package dropped 44 shipped files across twelve consecutive releases, 43 of them in one step. After every successful copy the script writes `darling-install-manifest.txt` into the install root recording the files that copy laid down, and the next upgrade diffs its own payload against it and **names** whatever an earlier build shipped and this one does not. Add `-RemoveStaleFiles` to delete them rather than only list them — off by default for now, so you can see the answer on your own boxes before anything acts on it. + +Everything it can name provably came out of one of our own zips, because the manifest is DERIVED from the payload rather than maintained by hand: `darling.json`, its `.bak-*` copies, the DPAPI credential blobs, the `_rollback_manual_*` backups and `pg-runtime\` were never in a payload, so they are never in the manifest and can never be nominated. When it cannot tell — no manifest yet, one it cannot parse, a source it cannot read — it removes nothing and says so on its own line. Deleting `darling-install-manifest.txt` is safe: the next upgrade reports that it cannot tell, removes nothing, and writes a fresh one. + +**Backups pile up, and nothing used to remove them.** A dogfood box was found carrying 46 of them, 5.48 GB, the oldest three weeks old, with the service naming every one on every start ([#2525](https://github.com/erikdarlingdata/PerformanceMonitor/issues/2525)). Retention above fixes new deploys; boxes that already have a backlog clear it with the installed copy, which needs no staging folder and does not stop the service: + +```powershell +C:\PerformanceMonitorDarling\upgrade-darling.ps1 -ListRollbacks # show what would go +C:\PerformanceMonitorDarling\upgrade-darling.ps1 -PruneOnly # remove all but the newest 3 +``` + +The service reports the set once per start with a count, a total and that command — informational while you are within retention, a warning past it. It never deletes one itself: it did not create them. + +**The install is not verified when the service reaches Running.** That means the process started, not that it collects. The script prints the post-start checklist; work it about 10–15 minutes later against the store, and hold the upgrade unverified until every line passes. + +#### Upgrading from a build that predates the script + +`upgrade-darling.ps1` was added after 3.5.0, so upgrading **from 3.5.0 or earlier** is a manual +procedure. It is the same sequence the script automates ([#2593](https://github.com/erikdarlingdata/PerformanceMonitor/issues/2593)). + +**Nothing you need to preserve lives in the install directory.** The managed store, the DPAPI +credential blobs and the logs are all under `C:\ProgramData\PerformanceMonitorDarling`, and anything +encrypted with `--encrypt-password` survives because those blobs are DPAPI **machine** scope rather than +account scope. The only file that has to travel is `darling.json`. That also means renaming the install +folder does **not** back up your data — if you want a data rollback point, snapshot the volume or stop +the service and copy that ProgramData folder. + +1. **Verify the download first.** `Get-FileHash -Algorithm SHA256` against `SHA256SUMS.txt` from the + same release page. This is the one step with no recovery if it is skipped. +2. `Stop-Service 'PerformanceMonitor Darling' -Force`, then **poll until `Status` is actually `Stopped`**. + Requesting a stop is not the same as it having stopped, and laying a build over a live tree is where + this goes wrong. +3. Rename the current install folder aside (e.g. `...\PerformanceMonitorDarling_3.3`). This is your + rollback. +4. Extract the new zip to the **original** folder name. +5. Copy `darling.json` from the renamed folder into the new one, then **hash it on both sides and confirm + they match**. A config silently altered mid-upgrade is very hard to notice afterwards. +6. Run `.\install-darling.ps1` from the new folder. For an existing install it re-points the service's + binPath in place and preserves config, store and credentials. +7. Start the service, then **verify** — see the post-start checklist below. Reaching `Running` means the + process started, not that it collects, and an upgrade that crosses several schema migrations has more + than usual to get wrong on first start. + +A clean extract like this has one advantage over the scripted overlay: files a newer build stopped +shipping cannot accumulate, which is the problem `-RemoveStaleFiles` exists to clean up on an +overlay-upgraded tree. + +#### The install location has to be machine-scoped + +Extract to a local, machine-scoped path — `C:\PerformanceMonitorDarling` is the documented one. **Not** anywhere under a user profile (`C:\Users\...`, including your Desktop or Downloads), and not a UNC path or a mapped drive. + +The service runs as the unprivileged virtual account `NT SERVICE\PerformanceMonitor Darling`, never LocalSystem, because the bundled PostgreSQL refuses to run with administrative privileges. That account is not you, not SYSTEM, and not Administrators — and a user profile grants access to about those three and nobody else, so the service cannot read its own program files there. It installs cleanly and then fails: `initdb.exe` dies at `0xC0000135` (STATUS_DLL_NOT_FOUND) before it can report anything (#2185). A folder created under `C:\` inherits read + execute for `BUILTIN\Users` instead, which the virtual account is a member of, which is why the documented location works. Network paths fail for a related reason: a virtual account [reaches the network as the computer account](https://learn.microsoft.com/en-us/sql/database-engine/configure-windows/configure-windows-service-accounts-and-permissions#virtual-accounts) rather than as you, and a mapped drive letter belongs to your logon session, which a service does not share. + +`install-darling.ps1` refuses a fresh install in any of these locations rather than leaving you a service that cannot start. A service registered by hand instead — the manual `sc create` path below, which the installer never sees — gets the same diagnosis from the service itself: on start, ahead of reading `darling.json` and long before the store bootstrap, it logs one critical line naming the path, why its own account cannot read it, and where to move it, so the cause is above the failure rather than three messages downstream of it. To move an existing install, stop the service, move the folder, and re-run `install-darling.ps1` from the new location — it updates the service's binPath in place and leaves your `darling.json`, store data, and credentials alone. + +**Manual:** publish (or copy the build output) to a stable path, put `darling.json` next to the exe (or set `DARLING_CONFIG` as a machine environment variable), then register it: + +``` +dotnet publish Darling/PerformanceMonitor.Darling.Service/PerformanceMonitor.Darling.Service.csproj -c Release -o C:\PerformanceMonitorDarling +``` + +``` +sc create "PerformanceMonitor Darling" binPath= "C:\PerformanceMonitorDarling\PerformanceMonitor.Darling.Service.exe" start= auto obj= "NT SERVICE\PerformanceMonitor Darling" +``` + +``` +sc start "PerformanceMonitor Darling" +``` + +Also register the service's Windows event source once, from the same elevated shell — event-source registration requires elevation, and the virtual service account cannot do it itself (without this, Event Log diagnostics are silently dropped; the file log under `%ProgramData%\PerformanceMonitorDarling\logs` works regardless): + +``` +powershell -NoProfile -Command "New-EventLog -LogName Application -Source 'PerformanceMonitor Darling' -ErrorAction SilentlyContinue" +``` + +The `obj=` clause runs the service under a **virtual service account** (`NT SERVICE\` — password-less, per-service SID, unprivileged; the same convention SQL Server itself uses). That is the right account for SQL-auth monitoring, and with `postgres.managed = true` it is more than a preference: PostgreSQL refuses to execute with administrative privileges, so don't run the service as LocalSystem — a least-privilege account keeps the bundled store's initdb/start path on ground PostgreSQL supports. For integrated auth to monitored servers, [run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) instead. Note the space after `binPath=`, `start=`, and `obj=` — `sc` requires it. + +One managed-mode handoff gotcha: if you test-drove the service from a console first, the bundled store's data directory belongs to *your* account, and the service account may not be able to write it. Point the service at a fresh `postgres.dataDirectory` (or delete the test directory) rather than fighting ACLs. + +#### Run the service as a domain account or gMSA + +With `"auth": "integrated"`, the monitoring identity **is** the service's Log On account — nothing in `darling.json` names a Windows account, and there is no separate credential to set. The default virtual account carries only the *machine* identity onto the network (remote servers see `DOMAIN\$`), so for integrated auth against remote servers you almost always want a real AD service account or, better, a gMSA. Switching is a Windows-side change plus a SQL-side grant, with one file-permission step in the middle that bites everyone who skips it: + +1. **Change the Log On account.** Stop the service, then **Services.msc → PerformanceMonitor Darling → Log On → This account** — that route also grants the account the *Log on as a service* right automatically. Or from an elevated prompt (with `sc config` you grant *Log on as a service* yourself, via secpol.msc or GPO): + + ``` + sc config "PerformanceMonitor Darling" obj= "DOMAIN\svc-account" password= "ThePassword" + ``` + + A gMSA works the same way with an empty password: `obj= "DOMAIN\gmsa-name$" password= ""`. Keep the account **out of the local Administrators group**: with `postgres.managed = true` the bundled PostgreSQL refuses to run with administrative privileges, exactly as it refuses LocalSystem. + +2. **Grant the account on every monitored server** — a Windows login holding the same [permissions below](#permissions-on-monitored-servers) (the `GRANT`s there apply to a Windows login unchanged): + + ```sql + USE [master]; + CREATE LOGIN [DOMAIN\svc-account] FROM WINDOWS; + ``` + +3. **Re-grant the service's own files — the step people miss.** The service deliberately locks its files down to SYSTEM, Administrators, and the account it was *running as*; the new account is on none of those ACLs, and the service will fail to read its config or write its store. One-time, from an elevated prompt, before starting the service: + + ``` + icacls "C:\ProgramData\PerformanceMonitorDarling" /grant "DOMAIN\svc-account:(OI)(CI)F" + icacls "C:\PerformanceMonitorDarling\darling.json" /grant "DOMAIN\svc-account:F" + ``` + + Adjust the second path to wherever `darling.json` sits beside the service exe; the first covers the logs and, in managed mode, the store's data directory. On its next start the service re-asserts the tight ACL itself — now including the new account — so this does not need repeating. + + In managed-store mode there is one more, and it needs **ownership**, not a grant: the store's superuser credential `pg-credential.dpapi` (beside the data directory, under `C:\ProgramData\PerformanceMonitorDarling` by default) is trusted only when *owned* by SYSTEM, Administrators, or the service account — an anti-pre-plant check — and `icacls /grant` changes permissions, never ownership, so after the switch the file is still owned by the *previous* service account and the service refuses it. Hand ownership to Administrators (trusted across any future account change, which is why not the new account itself) and grant the new account on the file directly — its ACL is protected and does **not** inherit the folder grant above: + + ``` + takeown /f "C:\ProgramData\PerformanceMonitorDarling\pg-credential.dpapi" /a + icacls "C:\ProgramData\PerformanceMonitorDarling\pg-credential.dpapi" /grant "DOMAIN\svc-account:F" + ``` + + The sibling role credentials (the admin/viewer/mcp `.dpapi` files) hit the same ownership check but self-heal — a role password can be re-asserted, a superuser's cannot — so expect one-time `discarding and regenerating` warnings on the first start, not faults. + +4. **Start the service and verify from its log** (`%ProgramData%\PerformanceMonitorDarling\logs`): the per-server connect lines are the proof that the *service account's* grants work. `--test-connection` from your console runs as you, not the service account — see the [pre-flight note above](#validate-the-config-pre-flight). + +Nothing else moves: anything encrypted with `--encrypt-password` (SQL-auth server passwords, SMTP) survives the account change, because those blobs are DPAPI **machine**-scope, not account-scope — and collected data is untouched. Later `install-darling.ps1` upgrades preserve a custom Log On account and harden `darling.json` for the account the service actually runs as. + +### What the Service Does on Monitored Servers + +On each successful connect, the service: + +1. **Probes the server** — one query against `sys.dm_os_sys_info` / `SERVERPROPERTY()` for version, engine edition (box / Managed Instance / Azure SQL DB), AWS RDS detection, and msdb access. It is the same detection query Lite runs, so both editions classify a server identically. +2. **Ensures two Extended Events ring-buffer sessions** (created if missing, started if stopped; ~4 MB ring buffer each, no files written on the server): + - `PerformanceMonitor_Deadlock` — `xml_deadlock_report`, server-scoped on on-prem/Managed Instance/RDS; `database_xml_deadlock_report`, database-scoped on Azure SQL Database. + - `PerformanceMonitor_BlockedProcess` — `blocked_process_report`, server-scoped (database-scoped on Azure SQL Database). +3. **Bootstraps the blocked-process threshold** — if `blocked process threshold (s)` is `0`, the service sets it to `5` via `sp_configure`. On AWS RDS `sp_configure` is unavailable; the attempt is tolerated and logged, and you set the threshold through an RDS Parameter Group instead (Azure SQL Database has a fixed 20-second threshold). +4. **Runs the on-connect config snapshots once** (`server_config`, `database_config`, `database_scoped_config`, `trace_flags`, `server_properties`), then runs all scheduled collectors on the shared default cadences. + +Every failure in steps 2–3 is tolerated and logged: the deadlock/blocked-process collectors simply read zero rows until the sessions exist (and blocked-process reports only start arriving once the threshold is set). Monitoring queries connect with a 15-second connect budget and an application name of `PerformanceMonitorDarling`; connection encryption fails closed to `Mandatory` when the configured mode is unrecognized. + +### Permissions on Monitored Servers + +Darling needs the **same target-server grants as Lite**, so the copy-paste block lives in one place for both: **[Permissions in the root README](../README.md#lite--darling-on-premises)** — `VIEW SERVER STATE`, `CONNECT ANY DATABASE`, `VIEW ANY DEFINITION`, `ALTER ANY EVENT SESSION`, and the optional `ALTER TRACE`, `ALTER SETTINGS`, and msdb job-table grants, verified live against SQL Server 2025 with a scratch login carrying exactly them ([#1823](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1823)). That block is authoritative; this section is the Darling-specific reading of it. Keeping one list instead of two is deliberate — a second copy is how the old one went stale. + +**The one Darling-specific line:** for `"auth": "integrated"` the grants go to the Windows account **the service runs as**, so use `CREATE LOGIN [DOMAIN\svc-account] FROM WINDOWS;` in place of the block's `CREATE LOGIN ... WITH PASSWORD`. Everything after it is unchanged. See [Run the service as a domain account or gMSA](#run-the-service-as-a-domain-account-or-gmsa) for which account that actually is — it is not the one you ran `--test-connection` as. + +What each grant buys you, and what breaks without it: + +| Grant | Why | If missing | +|---|---|---| +| `VIEW SERVER STATE` | All DMV collectors (wait stats, query stats, memory, CPU, file I/O, sessions, etc.) and the connect probe | Collection fails — this one is required | +| `ALTER ANY EVENT SESSION` | Create/start the two XE sessions | Logged; deadlock and blocked-process collectors read zero rows (an admin can pre-create the sessions instead) | +| `CONNECT ANY DATABASE` | The per-database collectors (`database_scoped_config`, `query_store_health`, `index_object_stats`, `database_size_stats`, `query_store_stats`) enter each database via `EXECUTE [db].sys.sp_executesql` | Databases the login cannot enter are skipped; without the grant that is every user database | +| `VIEW ANY DEFINITION` | Catalog-view row visibility everywhere: `sys.tables` / `sys.indexes` / `sys.objects` for the index and object collectors, `sys.dm_db_partition_stats`, and the AG catalog views (`sys.availability_groups`, `sys.availability_replicas`) | **Silently zero rows** — catalog views hide rows rather than erroring, so missing objects look exactly like empty databases, and a real AG cluster looks identical to a server with no AGs | +| `ALTER SETTINGS` | The `sp_configure` blocked-process-threshold bootstrap | Logged; set the threshold yourself (or via RDS Parameter Group) | +| `ALTER TRACE` | The `default_trace_events` collector — `sys.traces` / `fn_trace_gettable` accept nothing less | `PERMISSIONS` skip in collection health; the default-trace tab stays empty | +| msdb job-table `SELECT`s + `agent_datetime` `EXECUTE` | `running_jobs` / `job_history` / `agent_status` collectors and the failed/long-running-job alerts — all direct table reads; `SQLAgentReaderRole` alone leaves every one failing with error 229 | Skipped gracefully — logged as a permissions skip, alerts return no jobs | +| `DBCC TRACESTATUS` permission | `trace_flags` snapshot | Degrades to zero rows with a warning | + +The msdb grants live inside a system database SQL Server setup can rewrite — re-check them after a CU or version upgrade. + +**Azure SQL Database:** connect to the one database you monitor (set the server entry's `"database"`), using a contained user with `VIEW DATABASE STATE` and `VIEW DEFINITION`, matching the product's existing Azure guidance. The XE sessions are created database-scoped there (`ALTER ANY DATABASE EVENT SESSION`); SQL Agent collectors are skipped automatically. + +Collectors that hit a permission error (SQL errors 229/297/300, plus 8189 from `sys.traces`) log a `PERMISSIONS` row in `collection_log` and retry on their next scheduled run — one denied collector never stops the rest. + +#### Which collectors run on which platform + +Every collector declares its own applicability in code (`AppliesTo(CollectorTargetInfo)`), so this is not a hand-maintained list of 36 rows — the collectors fall into five groups, and a collector outside its supported platform is **skipped before it runs**, not failed and logged every cycle. + +| Runs on | Collectors | Gate | +|---|---|---| +| Everything | wait stats, CPU utilization, memory (stats/clerks/grants), file I/O, tempdb, latches, spinlocks, plan cache, session summary, plus blocking, deadlocks, blocked-process reports, DMV blocking snapshots, perfmon, query snapshots, procedure stats, index/object stats, long-query completions, database config/scoped-config/size, server properties, session stats, waiting tasks | no gate | +| On-prem, Managed Instance, RDS — **not** Azure SQL DB | CPU scheduler stats, default trace events, memory pressure events, server config, system health events, trace flags | `!IsAzureSqlDb` | +| On-prem and Managed Instance, needs msdb | job history | `!IsAzureSqlDb && HasMsdbAccess` | +| On-prem and Managed Instance, needs msdb — **not** RDS | agent status, running jobs | `!IsAzureSqlDb && !IsAwsRds && HasMsdbAccess` | +| SQL Server 2016+ (or any Azure flavour) | query stats, Query Store stats | `SqlMajorVersion >= 13 \|\| IsAzureSqlDb \|\| IsAzureManagedInstance` | + +Notes: + +- **Azure SQL DB** is the most restricted target: the six `!IsAzureSqlDb` collectors read server-scoped DMVs or on-disk artifacts that do not exist there, and the SQL Agent collectors have no Agent to read. Nothing about that is a permission problem, so it is not reported as one. +- **AWS RDS** blocks direct `msdb` job reads specifically; the rest of the SQL Agent surface is unaffected. +- **`HasMsdbAccess`** is probed per server at connect and is exactly `HAS_DBACCESS('msdb')` — *any* access to msdb, not a specific role or table grant. Losing msdb access later moves those collectors from running to skipped without an error storm. A login that can enter msdb but lacks `SELECT` on the job tables passes this probe and is caught one layer down as a `PERMISSIONS` skip instead. +- An unknown version (`SqlMajorVersion == 0`, i.e. detection has not completed yet) is treated as capable rather than skipped, so a collector is never silently dropped because a probe was slow. + +If a tab or column is empty and you expect data, check **Collection Health**: a collector skipped for platform reasons shows no runs at all, whereas one denied by permissions logs `PERMISSIONS` and is classified `NO_PERMISSIONS`. Those are different problems with different fixes — the first is expected on that platform, the second is a grant to add from the table above. + +--- + +## Configuration Reference + +All sections except `postgres` and `servers` are optional — omit a section (or any key) to get the defaults listed here. Defaults deliberately mirror a fresh Lite install. + +### postgres + +Two mutually exclusive modes — setting both `managed: true` and `connectionString` is a validation error: + +| Key | Default | Notes | +|---|---|---| +| `managed` | `false` | `true` runs the bundled PostgreSQL + TimescaleDB (Windows only; see [Managed Bundled PostgreSQL](#managed-bundled-postgresql)). The connection string is derived, never configured. | +| `port` | `5641` | Managed mode only: the loopback port the bundled server listens on. Deliberately uncommon so it coexists with any PostgreSQL (5432) already on the machine. | +| `dataDirectory` | *(null)* | Managed mode only: the cluster's data directory. `null` means `%ProgramData%\PerformanceMonitorDarling\pg`. | +| `connectAs` | `"admin"` | Managed mode only: which least-privilege role the Viewer connects as — `"admin"` (reads everything + manages mute rules and dismisses alerts) or `"viewer"` (read-only; those write actions are hidden/disabled). See [Security & Least-Privilege Roles](#security--least-privilege-roles). Ignored in bring-your-own mode (the connection string picks the role). | +| `connectionString` | *(required unless managed)* | Npgsql connection string for a store you provision yourself, e.g. `Host=localhost;Port=5432;Username=darling;Password=...;Database=darling`. You own that cluster's settings: if it has TimescaleDB, size its [background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont) — managed mode does this for you, this mode does not. | + +### servers (array, at least one entry) + +| Key | Default | Notes | +|---|---|---| +| `name` | `""` | Display name; falls back to `host` | +| `host` | *(required)* | Server/instance to monitor | +| `engine` | `"sqlserver"` | `"sqlserver"` or `"postgres"` (`postgresql` / `pg` / `aurora-postgresql` also accepted). Configuration rather than something probed, because it decides which driver builds the connection string before there is a connection to ask. An omitted or unrecognized value means SQL Server, so every existing `darling.json` keeps its exact present behaviour — see [PostgreSQL targets](#postgresql-targets) | +| `port` | *(driver default)* | PostgreSQL targets on a non-default port. SQL Server carries its port in the host as `host,1433` instead, and that convention is left alone | +| `database` | *(none)* | Azure SQL Database only: the one database this entry monitors (also part of the server's storage identity). PostgreSQL targets connect to the maintenance database and read cluster-wide catalogs | +| `auth` | `"integrated"` | `"integrated"` or `"sql"` | +| `username` | *(none)* | Required for `"sql"` | +| `encryptedPassword` | *(none)* | DPAPI blob from `--encrypt-password` (preferred) | +| `password` | *(none)* | A literal (dev only, warned on every use) or an `env:NAME` / `file:/path` reference (#1804) — references are the supported non-Windows shape and are not warned | +| `readOnlyIntent` | `false` | Route to a readable AG secondary (`ApplicationIntent=ReadOnly`) | +| `trustServerCertificate` | `false` | | +| `encryptMode` | `"Mandatory"` | `Mandatory` / `Strict` / `Optional`; unknown values fail closed to `Mandatory` | +| `multiSubnetFailover` | `false` | | +| `excludedDatabases` | `[]` | Databases excluded from collection | + +### capturePlans (boolean, optional) + +| Key | Default | Notes | +|---|---|---| +| `capturePlans` | `true` | Capture execution plans into `query_stats.query_plan_xml` and `query_store_stats.query_plan_text`. PostgreSQL TOAST compresses the plan text transparently (LZ4 on the managed store) and TimescaleDB chunk compression squeezes it further, so plans are cheap to keep — unlike Lite, which stores to DuckDB/Parquet and deliberately never captures them. Set `false` to skip plan capture (e.g. to shave storage across a very large fleet). | + +### collectSchemaChangeEvents (boolean, optional) + +| Key | Default | Notes | +|---|---|---| +| `collectSchemaChangeEvents` | `true` | Record `Object:Created` / `Object:Altered` / `Object:Deleted` schema-change (DDL) events in the built-in default-trace collector. Set `false` on a noisy or benchmark box where a create/drop-happy workload floods the viewer's **System Events > Default Trace** tab — e.g. HammerDB's TPC-H Query 15 creates and drops a `revenue` view thousands of times, and the collector faithfully records every create/delete. Only the Object DDL slice is suppressed; file auto-grow/shrink, ErrorLog, and security-audit events are still collected. The shared collector's equivalent of the full Dashboard's `@include_object_events`. A file-only knob (not stored in the control plane): edit and restart. | + +### alerts + +The shared alert engine's switches and thresholds. Every default mirrors Lite's alert defaults exactly, so an empty section alerts like a fresh Lite install. `enabled: false` turns off all alert evaluation **and** scheduled-analysis finding notifications (the analysis itself still runs and persists findings). + +| Key | Default | Meaning | +|---|---|---| +| `enabled` | `true` | Master switch for alert evaluation + finding notifications | +| `cpuEnabled` | `true` | | +| `cpuThresholdPercent` | `80` | | +| `cpuMode` | `"total"` | `"total"` = SQL + other processes; `"sql"` = SQL process only | +| `blockingEnabled` | `true` | | +| `blockingCountThreshold` | `1` | Blocked-process count (rolling window) that trips the alert | +| `blockingWaitSecondsThreshold` | `0` | Total blocked wait, in seconds, summed across the latest blocking snapshot; `0` = off. A second gate beside the count one, because a count cannot tell one session blocked for an hour from one blocked for a second. Reports as its own "Blocking Wait Time" alert, and unlike the count gate it is level-triggered: it re-fires every cooldown while the wait stays above the threshold and clears when it drops below | +| `deadlockEnabled` | `true` | | +| `deadlockCountThreshold` | `1` | Deadlock count (rolling window) that trips the alert | +| `poisonWaitEnabled` | `true` | THREADPOOL / RESOURCE_SEMAPHORE / RESOURCE_SEMAPHORE_QUERY_COMPILE | +| `poisonWaitThresholdMs` | `500` | Average ms per wait | +| `longRunningQueryEnabled` | `true` | | +| `longRunningQueryThresholdMinutes` | `30` | | +| `tempDbSpaceEnabled` | `true` | | +| `tempDbSpaceThresholdPercent` | `80` | | +| `lowDiskEnabled` | `true` | Volume free space; graded CRITICAL when critically low | +| `lowDiskThresholdPercent` | `10` | Fire below X% free; `0` disables this dimension (clamped 0–100) | +| `lowDiskThresholdGb` | `5` | Fire below X GB free; `0` disables this dimension | +| `longRunningJobEnabled` | `true` | SQL Agent job running long vs. its history | +| `longRunningJobMultiplier` | `3` | Fires at 3x the job's historical average | +| `failedJobEnabled` | `true` | Live msdb check for recently failed jobs | +| `failedJobLookbackMinutes` | `60` | Clamped 1–1440 | +| `cooldownMinutes` | `5` | Minimum minutes between repeats of the same alert condition (clamped 1–120) | +| `excludedDatabases` | `[]` | Excluded from blocking/deadlock/long-running-query **alert evaluation** (collection unaffected) | + +Not configurable (hardcoded to Lite's defaults until someone needs a knob): the long-running-query read shape (top 5 results; the five noise filters — sp_server_diagnostics, WAITFOR, backups, misc waits, CDC — all on) and the analysis-finding notification policy (notify at severity >= 1.5, 6-hour per-finding cooldown). + +### smtp + +Email delivery is enabled when `host`, `from`, and `to` are all set — there is no separate enable flag. + +| Key | Default | Notes | +|---|---|---| +| `host` | `""` | | +| `port` | `587` | | +| `useSsl` | `true` | | +| `username` | *(none)* | For authenticated relays | +| `encryptedPassword` | *(none)* | Same `--encrypt-password` DPAPI pattern as SQL auth | +| `password` | *(none)* | A literal or an `env:NAME` / `file:/path` reference (#1804) — the non-Windows email path | +| `from` | `""` | | +| `to` | `""` | Comma-separated recipients | +| `emailCooldownMinutes` | `15` | Email/webhook channel cooldown (clamped 1–120) | + +### webhooks + +A channel is enabled by a non-empty URL. + +| Key | Default | Notes | +|---|---|---| +| `teamsUrl` | `""` | Teams incoming webhook | +| `teamsProxy` | `""` | Optional proxy address | +| `slackUrl` | `""` | Slack incoming webhook | +| `slackProxy` | `""` | Optional proxy address | + +### mcp + +The embedded MCP server, over Streamable HTTP bound to `localhost` by default (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan) to reach it — and the store — from the LAN). It exposes the same tool names Lite and the Dashboard expose, plus small Darling-only WRITE surfaces — Custom Views management, alert tuning, and server onboarding (see the last three bullets): + +- **Six diagnostic-analysis tools** — `analyze_server`, `get_analysis_facts`, `compare_analysis`, `audit_config`, `get_analysis_findings`, `mute_analysis_finding`. +- **Five plan-analysis tools** — `analyze_query_plan` (by `query_hash`), `analyze_procedure_plan` (by `sql_handle`), `analyze_query_store_plan` (by `database_name` + `query_id`), `analyze_plan_xml` (raw showplan XML, no fetch), and `get_plan_xml` (raw stored plan XML by `query_hash`). These run the shared execution-plan analyzer over the plan XML the collectors already captured into the store — a stored-plan read, never a live query against the monitored server. `analyze_query_plan`/`get_plan_xml` accept an optional `database_name`, and `analyze_query_store_plan` an optional `plan_id`, to pin the exact stored plan when the caller knows it. +- **Fifteen core data-read tools** — the diagnostic reads an assistant needs to investigate a server, each a stored read of the collected data (never a live query against the monitored server): + - *Resource metrics* — `get_cpu_utilization`, `get_wait_stats`, `get_wait_trend`, `get_wait_types` (the distinct observed wait types, to pick one for `get_wait_trend`), `get_memory_stats`, `get_memory_clerks`, `get_file_io_stats`, `get_tempdb_trend`, `get_perfmon_stats`. + - *Query performance* — `get_top_queries_by_cpu`, `get_top_procedures_by_cpu`, `get_query_store_top` (these hand back the `query_hash` / `sql_handle` / `query_id` + `plan_id` keys the plan-analysis tools consume). + - *Discovery / health* — `list_servers` (with collection-freshness status, and the [declared peer stores](#peers) when the fleet is split across several Darling boxes), `get_collection_health`, `get_server_properties`. + + These are the tools the analysis findings' `next_tools` recommendations point at, so a client following a finding's advice resolves them on this same server. Result shapes match Lite's (the store is Lite's collector schema); where Lite and the Dashboard's shapes diverge, Darling follows Lite — the shape its collector-mirror store can serve faithfully. +- **Twenty diagnostic-depth data-read tools** — deeper reads for a blocking / deadlock / session / configuration / storage investigation, each a stored read: + - *Blocking / deadlocks* — `get_blocking` (blocked/blocking pairs from the blocked-process-report XE + the always-on DMV fallback), `get_deadlocks`, `get_deadlock_detail` (raw graph XML), `get_blocked_process_xml` (raw report XML), and the per-minute count series `get_blocking_trend` / `get_deadlock_trend`. + - *Sessions* — `get_session_stats` (latest per-application connection counts), `get_active_queries` (captured running-query snapshots), `get_waiting_tasks`. + - *Config* — the change history `get_server_config_changes`, `get_database_config_changes`, `get_trace_flag_changes`, plus the latest-snapshot pair `get_database_scoped_config` / `get_query_store_health` (per-database Query Store health: actual vs desired state, readonly_reason decoded, storage vs cap) and the current-config snapshots `get_server_config` / `get_database_config` / `get_trace_flags` (what sp_configure / sys.databases / the active trace flags are set to **right now** — the companion to the `*_changes` diffs, which are empty on a stable server). + - *Index / object* — `get_table_index_sizes` (size + growth), `get_index_usage` (Unused / Write-only / Active), `get_object_locking` (lock/latch contention), `get_database_sizes`. + + The three config-change tools diff the store's config snapshots. This edition captures configuration **when the service connects** to a server (not on a fixed schedule), so a change is detected between two connect snapshots and at least two are needed — a stable, always-connected deployment may show no changes until the next connect. They emit only the values the collectors capture; the Dashboard's `requires_restart` / setting `description` / `setting_type` / generated change-narrative enrichment is not collected here and is omitted. The Dashboard's `get_blocking_deadlock_stats` aggregate is **not** hosted (Darling has no blocking/deadlock rollup table — use `get_blocking` / `get_deadlocks` for the raw events). + +- **Eight resource-contention + jobs data-read tools** — deeper reads for an internal-contention / worker-thread / plan-cache / SQL Agent investigation, each a stored read of the latest collected snapshot: + - *Latch / spinlock* — `get_latch_stats` (top latch classes by wait time, per-second rates), `get_spinlock_stats` (top spinlocks by collisions). + - *Memory grants* — `get_resource_semaphore` (workspace-memory target / max-target ceiling vs granted / used), `get_memory_grants` (per-pool grant detail), `get_memory_pressure_events` (RING_BUFFER_RESOURCE_MONITOR notifications — the process/system pressure indicators, not on Azure SQL DB). + - *Plan cache / scheduler* — `get_plan_cache_bloat` (single-use vs multi-use + bloat level), `get_cpu_scheduler_pressure` (runnable queue, worker utilization, pressure level). + - *Jobs* — `get_running_jobs` (running SQL Agent jobs vs historical average / p95). + + The Dashboard's per-class latch `severity` / `description` / `recommendation`, spinlock `description`, plan-cache `bloat_level`, and CPU-scheduler `pressure_level` / `recommendation` are the Dashboard / reporting-view CASE derivations (not collected columns), reproduced service-side so the full result shape is served. Darling's delta collectors store no `sample_interval_seconds`, so per-second latch/spinlock rates are derived from the collection interval, and the Dashboard's `get_resource_semaphore` `sample_interval_seconds` is not emitted for the same reason (`max_target_memory_mb`, the workspace-memory ceiling, is added since the store carries it). + +- **Twelve PostgreSQL data-read tools** — the read surface for a PostgreSQL target's collectors, each a stored read (see [PostgreSQL targets](#postgresql-targets)): + - *Waits and queries* — `get_pg_wait_stats` (top wait events in the window, decoded to type + event name), `get_pg_top_queries` (query shapes by total execution time, carrying Aurora's storage-vs-cache I/O split and per-statement peak memory). + - *Outage predictors* — `get_pg_wraparound_risk` (XID and MultiXact freeze headroom per database), `get_pg_xmin_horizon` (why vacuum is reclaiming nothing, attributed to the specific holder), `get_pg_replication_slots` (slot health, including whether retained WAL is still growing). + - *Maintenance* — `get_pg_autovacuum_health` (tables behind on vacuum or analyze, ranked by how far past each table's OWN trigger threshold it is — the ratio, not the dead-tuple count, because the same count is routine on a large table and urgent on a small one). + - *I/O attribution* — `get_pg_io_stats` (reads, hits, extends and evictions by backend type, object and context). The context dimension is the one with no SQL Server counterpart and the one that changes the remedy: it separates ordinary buffer-pool misses, where more `shared_buffers` or a better index helps, from sequential scans that deliberately bypass the pool through a small ring buffer, where neither will. + - *Per-database counters* — `get_pg_database_stats` (temp-file spills, cache hit ratio, deadlocks, and the commit/rollback split, from `pg_stat_database`). **Temp files are the reason it exists**: they are work `work_mem` could not hold, which is the most common cause of a PostgreSQL query being slow for a reason its plan shape does not show, and on stock PostgreSQL it is the only temp-file evidence available anywhere. A statistics RESET is reported as a reset — an explicit count and a lower-bound caveat — rather than surfacing as a negative rate or a spike. + - *Storage* — `get_pg_table_bloat` (the per-table bloat ESTIMATE, suppressed rather than captioned when its statistics cannot be trusted) and `get_pg_index_usage` (per-index scan counts with the constraint, replica-identity and validity facts that decide whether an unscanned index can be dropped at all). + - *Sessions and contention* — `get_pg_blocking` (blocking chains that were SAMPLED, with the root blocker attributed and both sides' state captured — PostgreSQL has no engine-side blocked-process recorder, so blocking shorter than the interval leaves no trace anywhere) and `get_pg_session_states` (who is holding a transaction open, and whether they actually pin the xmin horizon). **That second half is a different question from the first**, and the tool exists to keep them apart: an `idle in transaction` session under READ COMMITTED that has only read, or whose write matched no rows, holds neither a snapshot nor a transaction id and starves vacuum of nothing — both measured on a live instance. So the read gates its causal claim on `peak_horizon_age`, where `-1` means the session pinned NOTHING, and tells you so rather than letting a long duration imply otherwise. Pairs with `get_pg_xmin_horizon`, which names the CLASS of holder where this names the session. + + These are separate tools rather than widened SQL Server ones. PostgreSQL's waits are a two-level type/event taxonomy with no signal-wait concept reported in microseconds, and the wraparound / horizon / slot signals have no SQL Server counterpart at all — sharing a result shape would mean lying about a unit or emitting mostly-null columns. The three outage predictors are the ones worth wiring to a pager: each names a condition that stops the server outright, and each is silent until it is nearly too late. + +- **Five trend data-read tools** — windowed time-series siblings of the core reads, each a stored read of the collected series over the window (BOTH-sides, naive-UTC): + - `get_memory_trend` (total / target server memory, buffer pool, plan cache over time), `get_perfmon_trend` (a single counter's value + delta, `counter_name` required), `get_file_io_trend` (per-database read/write latency, top-10 busiest files), `get_query_trend` (one query's per-collection history by `query_hash` + `database_name`), `get_query_duration_trend` (overall elapsed-ms/sec + executions/sec). + + Each mirrors the viewer's proven chart read (byte-identical Postgres SQL); the shape follows Lite where the SKUs diverge. `get_perfmon_trend` reproduces Lite's miss vocabulary (Page Life Expectancy is intentionally not collected; an unknown counter hands back the collected names). `get_memory_trend` carries a `total_granted_mb` field for field-for-field parity with Lite, where its memory_stats-only read leaves it 0 (the grant overlay is a separate chart series). + +- **Eight system-health parse-on-read tools** — the Dashboard's `get_health_parser_*` family, over Darling's raw `system_health_events`: + - `get_health_parser_system_health` (corruption + contention counters), `get_health_parser_severe_errors` (severity ≥ 19, with `database_id` resolved to a name), `get_health_parser_scheduler_issues`, `get_health_parser_memory_conditions`, `get_health_parser_memory_broker`, `get_health_parser_memory_node_oom`, `get_health_parser_cpu_tasks`, `get_health_parser_io_issues`. + + Where the Dashboard reads its server-side-parsed `collect.HealthParser_*` tables, these shred the raw extended-event XML **on read** with the shared `SystemHealthParser` (the same parser the viewer's System Events tab uses) and gate with the service-side twin of the viewer's `SystemEventSignificance` — returning the same SIGNIFICANT warning set the Dashboard surfaces (sp_HealthParser at `@warnings_only = 1`). `get_health_parser_system_health` is the one UNGATED category (its counter series plots every snapshot). Each row carries the full sp_HealthParser column set keyed on the event's `event_time`; the tools window on `event_time` (the event's real time), so "last 24 hours" means events that happened in the last 24 hours. + +- **Five alert + health-overview tools** — the fleet-triage reads the fleet edition previously lacked, each a stored read over the monitoring store (no live hit): + - *Alerts* — `get_alert_history` (what fired, value vs threshold, delivery success/failure, muted — fleet-wide by default, or scoped to a server), `get_alert_settings` (the current alert config the service is using — per-alert enable/thresholds, cooldown, excluded databases, delivery mode, analysis cadence), `get_mute_rules` (the alert mute rules in force, so a suppressed server is distinguishable from a healthy-quiet one). + - *Health overview* — `get_server_summary` (one-shot per-server CPU / memory / recent blocking / recent deadlocks), `get_daily_summary` (a day's composite health band — Healthy / Warning / Critical — folded through the shared `DailyHealthBandCalculator`, plus the signals behind it). + +- **Eight Custom Views tools (Darling-only)** — discover, create, and manage the saved dashboards/notebooks a user composes from the curated measure catalog (the same views the web viewer's editor builds), stored in `config.custom_views`. None touches a monitored SQL Server or the collected performance data — the write tools write only view definitions to the monitoring store. + - *Discover* — `describe_custom_view_catalog` (the compose vocabulary — measures with their source/kind/valid-aggregates/allowed-dimensions/units/per-server-type availability, dimensions, unit families, aggregates, time buckets, filter ops, and viz types). An MCP client calls this FIRST so a composed panel uses only legal identifiers instead of guessing at names; it returns the SAME `/api/catalog` vocabulary the web composer's picker binds to. Read-only static reference — no store, no server. + - *Read* — `list_custom_views` (summaries: id, name, description, kind, version), `get_custom_view` (one view's full definition + version). + - *Author* — `validate_custom_view` (dry-run a definition against the catalog + composer rules, no save), `create_custom_view` (validate then save), `update_custom_view` (validate then replace in place, optimistic-concurrency on `version`), `delete_custom_view`. + - *Self-test* — `run_custom_view_panel` (compile + run a single composed panel and return `{sql, rows, annotations}` — the composer's live preview, for checking a generated panel's data before saving). + + The create/update/delete tools are the one view-authoring **write** surface; create/update run the SAME `ValidateDefinition` authority as `validate_custom_view`, so an invalid definition is rejected before it stores; every tool routes through the SAME store + validator + compile-and-run + catalog the web viewer's editor uses (no divergent second implementation). This write surface is part of what the MCP token gates — see [What a token can reach](#opt-in-network-endpoints-lan) below. + +- **Three alert-tuning write tools (Darling-only)** — `update_alert_settings`, `create_mute_rule`, and `delete_mute_rule` let an MCP client TUNE the alert engine the fleet shares — the SAME config `get_alert_settings` / `get_mute_rules` read and the Viewer's Settings window writes. `update_alert_settings` is a PARTIAL update of the single global settings row: read via `get_alert_settings`, change fields, and send only those back in the same nested shape; every field is validated against the SAME ranges/enums the Settings window enforces BEFORE any write, an out-of-range or unknown field returns `{status:"invalid"}` and writes nothing, and the write self-bumps `config_version` so the running service hot-reloads within one collection sweep. `create_mute_rule` / `delete_mute_rule` reuse the SAME `PgMuteRuleStore` `get_mute_rules` reads through (and the same GUID id-generation the Viewer's mute-create path uses). None touches a monitored SQL Server or the collected data — only the shared alert configuration; SMTP/webhook delivery credentials are out of scope (the `mcp` role cannot read or write the secret columns). It is part of what the MCP token gates — see [What a token can reach](#opt-in-network-endpoints-lan) below. + +- **Two server-onboarding write tools (Darling-only)** — `add_servers` (BULK) and `remove_server` let an MCP client stand up or tear down FLEET monitoring conversationally ("monitor these twenty servers with this login"), the service-side twin of the Viewer's Add / Manage Servers dialogs. `add_servers` takes a JSON **array** of server objects (`host` required; optional `display_name` / `database` / `read_only_intent` / `multi_subnet_failover`; `auth` `Windows`/`SQL` with `username`+`password` for SQL; and the exposed TLS options `encrypt_mode` `Optional`/`Mandatory`/`Strict` + `trust_server_certificate`) and processes them **in order**: it validates each entry, PROBES the connection in-process (reusing the same `DarlingServerConnector.ProbeAsync` the `--test-connection` verb runs — the service holds the network path + credentials, so no `test_connect` command plane is needed), skips a case-folded duplicate (`duplicate`) of an already-monitored server or an earlier entry, DPAPI-encrypts the SQL password (the service identity, so it round-trips at collection time), and INSERTs the row mirroring the service's own seed shape. A server that fails to connect is `connection_failed` and the batch continues; Entra/MFA/Service-Principal/Managed-Identity auth is `invalid` (the service connects with Windows or SQL only). `remove_server` DELETEs a monitored server by name (resolved the same way every `server_name` is) — already-collected history is kept. Both write only the monitoring store's `config.config_monitored_servers` registry; neither runs anything on a monitored server beyond the one-time probe. **The SQL password travels to the endpoint inside `add_servers`' request** and is DPAPI-encrypted at rest (never returned) — it is part of what the MCP token gates, and it puts a credential on the wire; see [What a token can reach](#opt-in-network-endpoints-lan) below. + +| Key | Default | Notes | +|---|---|---| +| `enabled` | `false` | **Off by default** — a headless service does not open a local port unless you ask | +| `port` | `5152` | Chosen so all three editions coexist on one machine (Dashboard 5150, Lite 5151) | + +Register with Claude Code: + +``` +claude mcp add --transport http --scope user sql-monitor-darling http://localhost:5152/ +``` + +If the port is already in use at startup, the MCP server logs an error and does not start; collection is unaffected. + +### web + +The embedded read-only **web dashboard** — a browser view of the monitoring store, served over HTTP on its OWN port (default **5153**), separate from the MCP server. It is a distinct surface from [`### mcp`](#mcp): its own enable flag, port, token, and exposure block, because the two gate different blast radii (the MCP token guards `analyze_server`'s **live outbound** connections to your monitored SQL Servers; the web dashboard is **read-only over the collected store**). It connects to the store as the least-privilege `viewer` role. Loopback-only by default; see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan) to reach it from the LAN. + +| Key | Default | Notes | +|---|---|---| +| `enabled` | `false` | **Off by default** — a headless service does not open a local port unless you ask | +| `port` | `5153` | Chosen so all four local surfaces coexist on one machine (Dashboard 5150, Lite 5151, Darling MCP 5152) | + +Once enabled, open `http://localhost:5153/` in a browser on the service host. Like the MCP server, `enabled`/`port` here are the file SEED; after first start they live in the control plane and the Viewer's Settings toggles them LIVE (the service starts/stops/rebinds the dashboard within seconds — no restart). If the port is already in use at startup, the web host logs an error and retries on a calm cadence; collection is unaffected. + +**What you see.** The dashboard opens on a **Fleet Overview**: a card per enabled server with a status dot, six per-metric health bands (CPU, threads, memory, blocking, deadlocks, collectors), and its last collection time — all banded server-side, so the browser only renders (a server that has never reported shows an amber "Awaiting first collection", never a red offline). Above the cards a worst-first "Needs attention" list surfaces the servers to look at, or an all-healthy line when there is nothing to chase. Click a card to **drill into one server**: an overview, wait stats with a trend for the heaviest wait, active queries, a CPU chart, memory and file-I/O trends, and collection health — the same collected data the viewer shows, over inline charts. A fleet-wide **Alert History** page (with a server filter box) rounds out phase 1, and a **Custom Views** page hosts the composer for the saved dashboards/notebooks in `config.custom_views` — the dashboard's one write path. Otherwise it is a read-only view — no settings, no live-server queries — and it refreshes every 60 seconds (pausing while the tab is hidden). The frontend ships fully self-contained (no CDN, no fonts, no remote anything), so it works on an air-gapped host with no internet access. + +### peers + +**Declared peer stores** — optional, and only relevant when the fleet is split across **several Darling boxes**, one store each (SQL Server primaries on one box, their readable replicas on another, PostgreSQL on a third). Each box's MCP server answers over **its own** store only, so a server monitored by a sibling resolves as not-found — which an agent cannot tell apart from *"nobody monitors this server."* Declaring the siblings fixes that at the three places an agent forms its picture of the fleet. + +**Disclosure only.** There is no address and no credential in this block, and nothing behind it: the service never contacts a peer, cannot read a peer's data, and cannot tell whether a peer is even running. A peer is a **name** plus a **sentence**, so an agent (or its human) can pick the right endpoint. Everything here is sent verbatim to every connected MCP client, so the service **refuses to start** if any peer text looks like a connection string or credential. + +| Key | Default | Notes | +|---|---|---| +| `thisStoreCovers` | `""` | One sentence naming what THIS store monitors — the anchor the peer list is relative to | +| `stores[].name` | — | **Required.** Whatever an operator would recognize (the box name, "the use1 store") | +| `stores[].covers` | `""` | A short sentence naming what that store monitors. Human prose — never parsed, only shown | +| `stores[].matches` | `[]` | Optional server-name **substrings** that store monitors, case-insensitive. The only machine-checked field | + +```jsonc +"peers": { + "thisStoreCovers": "the 42 us-east-1 SQL Server primaries", + "stores": [ + { + "name": "prod-sql-use2-monitor-01", + "covers": "the readable replicas of those same 42 primaries, in-region from us-east-2", + "matches": ["use2"] + }, + { "name": "prod-sql-pg-monitor-01", "covers": "the Aurora PostgreSQL clusters", "matches": ["-aurora-"] } + ] +} +``` + +What it changes, with peers declared: + +- **The MCP instructions** gain a Fleet Coverage section, high enough that an agent reads which store it is talking to before it reads the tool census. +- **`list_servers`** gains `this_store_covers`, a `peer_fleets` array, and a `peer_note`. Both are always present: an *empty* `peer_fleets` has two very different meanings (this really is the only store, or nobody declared the siblings) and the service cannot tell them apart, so `peer_note` says exactly that rather than letting an empty array read as "this is the whole fleet." An **empty registry** answers in prose rather than JSON, and carries the peer list too — a store with nothing registered is a fresh or just-restarted box, which is the worst place to drop the disclosure. +- **The server-resolution miss** appends the disclosure to the existing "Could not resolve server. Available servers:" listing, naming the peer whose declared coverage matches — so *not monitored here* stops looking like *not monitored anywhere*. + +`matches` is deliberately plain substrings, no globbing and no regex: it exists to answer "which region/role prefix is this name?", and a pattern language would be a config surface with its own failure modes. Blank entries are dropped — an empty substring matches every name, which would make one peer claim the whole fleet. A peer with no `matches` is still disclosed everywhere; it just cannot be singled out on a miss, and the miss message says so instead of implying the server is unmonitored. + +A **file-only** block (not seeded into the control plane): it describes the deployment topology of *this* box, which must not be editable from a peer's Viewer. An edit takes effect on the next service restart. There is deliberately **no cross-store connectivity** here — actual federated reads (auth between stores, latency, partial failures) are a much larger surface, and may never be worth building if disclosure alone makes the split legible. + +**Declaring nothing changes nothing, with one exception worth knowing about on upgrade.** The instructions, the resolution-miss message, and `list_servers`' empty-registry sentence are byte-for-byte what they were. But `list_servers`' JSON envelope carries `this_store_covers`, `peer_fleets` and `peer_note` on *every* response, declared or not — so a script comparing that tool's exact shape sees three new keys even if you never write a `peers` block. That is deliberate: an empty `peer_fleets` means *either* "this is the only store" *or* "nobody declared the siblings", and a note that only appeared when peers were declared would say nothing in precisely the case that produces the wrong conclusion. + +**A `peers` block that fails validation is refused whole, and nothing is disclosed** — not the valid subset. An unfinished block that asserts coverage which may be wrong is worse than no block, and the service logs each problem at Critical. The check runs inside the publish rather than only in config validation, because the MCP host loads its own config and deliberately never validates it (its fail-closed checks are host-local), so validation alone would leave the one path that actually broadcasts uncovered. + +### No Schedule Knobs, by Design + +There are deliberately **no collection-schedule or retention settings** in `darling.json`. The service consumes the shared per-collector defaults (`CollectorScheduleDefaults`) — the same cadences and retention horizons a fresh Lite install uses, identity-pinned by tests so the two editions cannot drift. If a schedule knob is ever genuinely needed, it will be added then, not speculatively. + +--- + +## PostgreSQL Targets + +Darling monitors PostgreSQL alongside SQL Server. For the ordered procedure with a proof point at each step, see [**the first-target runbook**](../docs/postgres-first-target-runbook.md); this section is the reference for what each piece does. + +Add `"engine": "postgres"` to a `servers` entry and that target is collected by the PostgreSQL collectors instead of the T-SQL ones: + +```json +{ + "name": "orders-prod", + "engine": "postgres", + "host": "orders-prod.cluster-abc123.us-east-1.rds.amazonaws.com", + "auth": "sql", + "username": "darling_monitor", + "encryptedPassword": "" +} +``` + +**Which path registers a target depends on the store, not the file.** `darling.json` seeds `config.config_monitored_servers` once, when it is empty; after that the registry is authoritative and a darling.json edit adds nothing. So a fresh install declares its PostgreSQL targets in the file, and an existing one adds them with the [`add_servers`](#mcp) tool (or the Viewer's Add Server dialog), which takes effect within one collection sweep without a restart. The registry carries `engine` and `port` per row, so a target keeps its engine across restarts and reloads. + +`auth` must be `"sql"` — PostgreSQL has no integrated-authentication path here, and an entry asking for it fails [`--test-connection`](#validate-the-config-pre-flight) rather than waiting to fail at first connect. `add_servers` enforces the same rule, and unlike the file parser it REFUSES an unrecognized `engine` rather than resolving it to SQL Server: the file's leniency keeps one bad line from stopping the whole fleet at startup, while onboarding is a single deliberate act where a silent fallback would surface as a connection failure against the wrong port. Password handling is identical to a SQL Server entry: `--encrypt-password` produces the DPAPI blob, and the `env:NAME` / `file:/path` references work the same way. TLS defaults to full certificate verification (`SslMode=VerifyFull`); `trustServerCertificate` relaxes it to `Require`, which is the setting Aurora usually needs since it presents an RDS CA a stock trust store does not know, and `"encryptMode": "Optional"` relaxes it further to `Prefer`. + +**One store, both engines.** The PostgreSQL collectors write to the same store as the SQL Server ones, into their own tables, on the same naive-UTC contract and the same `server_id` identity. Nothing is partitioned by engine — a mixed fleet is one store, one viewer, one MCP endpoint. + +**A collector table must never be named after a `pg_catalog` object.** `pg_catalog` is searched implicitly and *first*, ahead of every entry in `search_path`, so an unqualified reference to such a name resolves to the system object no matter what the store holds. It fails loudly in one place — `CREATE INDEX` on a view is 42809, which aborts the migration and leaves the store unusable — and silently everywhere else: a reader's `FROM ` would return the monitoring store's own system view instead of collected history, so the tool reports nothing and any alert behind it never fires. That is why the slot collector stores into `pg_replication_slot_stats` while still being *named* `pg_replication_slots` after the view it reads (the same split as `query_store` → `query_store_stats`). A live-store test asserts no collector table shadows a catalog object, against the real catalog rather than a hardcoded reserved list. + +**A collector never runs against the wrong engine.** Every definition declares its `TargetEngine`, and both SKUs check it before dispatch, so a PostgreSQL target is never sent T-SQL and a SQL Server target never sees `pg_stat_statements`. A store monitoring only SQL Server still carries the five PostgreSQL tables, empty; nothing else about it changes. + +### Permissions on a PostgreSQL target + +One role covers every collector: + +```sql +CREATE ROLE darling_monitor WITH LOGIN PASSWORD ''; +GRANT pg_monitor TO darling_monitor; +``` + +`pg_monitor` is the standard PostgreSQL monitoring role — it bundles `pg_read_all_stats`, `pg_read_all_settings`, and `pg_stat_scan_tables`. Without it the statistics views still return rows, but only for the connecting user's own backends, which silently turns fleet monitoring into self-monitoring. On Amazon Aurora and RDS the same grant works: `GRANT pg_monitor TO darling_monitor;` as an `rds_superuser`. No superuser is needed, and nothing is created on the monitored server — unlike a SQL Server target, there are no Extended Events sessions to provision and no server setting to bootstrap. + +`pg_stat_statements` must be present for `pg_statement_stats`, which means the extension in `shared_preload_libraries` (a restart, or a parameter-group change plus reboot on Aurora/RDS) and `CREATE EXTENSION pg_stat_statements;` in the database Darling connects to. The extension tracks **all** databases in the cluster keyed by `dbid`, so one installation in the connect database covers the whole instance. The other six collectors need nothing installed — they read core catalogs and Aurora's built-in functions. + +### What gets collected + +| Collector | Source | Cadence / retention | Why it exists | +|---|---|---|---| +| `pg_wait_stats` | `aurora_stat_system_waits()` | 1 min / 30 d | **Aurora only.** Core PostgreSQL has no cumulative wait counters at all — `pg_stat_activity.wait_event` is an instantaneous sample — so there is no equivalent to `sys.dm_os_wait_stats` to read on a non-Aurora target | +| `pg_statement_stats` | `aurora_stat_statements()` | 1 min / 30 d | **Aurora only.** Per-query-shape totals, matching `query_stats`' cadence. Aurora's function adds the storage-vs-cache I/O split (`storage_blks_read` / `orcache_blks_hit`) and per-statement peak memory, neither of which core PostgreSQL exposes | +| `pg_wraparound_stats` | `pg_database`, `pg_class` | 5 min / 90 d | XID and MultiXact freeze headroom per database. The highest-consequence signal PostgreSQL has and one with no SQL Server counterpart: run out of transaction IDs and the server stops accepting writes. Freeze headroom moves in autovacuum-sized steps rather than continuously, so 5 minutes is ample and 90 days shows the age trend against the actual freeze threshold | +| `pg_xmin_horizon` | `pg_stat_activity`, `pg_replication_slots`, `pg_stat_replication`, `pg_prepared_xacts` | 1 min / 30 d | Why vacuum is reclaiming nothing. Four unrelated causes produce an identical symptom and need completely different fixes, so this attributes the specific holder instead of reporting the number. Per-minute because a holder is the fast-moving leading indicator — the useful answer is which session or slot appeared minutes ago | +| `pg_replication_slots` → `collect.pg_replication_slot_stats` | `pg_replication_slots` | 1 min / 90 d | Slot health and retained WAL. An abandoned slot retains WAL without bound by default, filling the volume and stopping the server, and it grows at whatever rate the server writes WAL — hours on a busy writer, not days | +| `pg_autovacuum_stats` | `pg_stat_user_tables`, `pg_class` | 60 min / 90 d | **Writers only**, and per database. Per-table autovacuum state. Stores each table's own computed trigger threshold beside its dead-tuple count, honouring per-table `reloptions` overrides rather than only the GUCs — without the threshold a dead-tuple count is not actionable, since the same count is routine on a large table and urgent on a small one | +| `pg_io_stats` | `pg_stat_io` | 1 min / 30 d | I/O attributed to a `(backend_type, object, context)` triple rather than to a file — who did it, to what, and why. PostgreSQL 16+; valid on a standby. Every counter column is nullable on purpose: PostgreSQL uses NULL for "does not apply to this combination", and on Aurora the whole write side is NULL because backends there do not write data files. The read reports whether write counters are TRACKED, so absent writes cannot be misread as zero writes | +| `pg_blocking` → `collect.pg_blocking_edges` | `pg_stat_activity`, `pg_blocking_pids()` | 1 min / 30 d | Who is blocked, by whom, and what state each side was in — stored as an edge list, one row per (blocked, blocking) pair, so the read layer can assemble chains and name the root. **This is a SAMPLE, not an event log**, and that is the one thing to carry away: SQL Server's blocked-process report is written by the engine when blocking crosses a threshold, whereas PostgreSQL records nothing unless something asks, so blocking shorter than the interval is never seen. Valid on a standby, where recovery conflicts are blocking that happens nowhere else. One minute is the floor worth paying for: `pg_blocking_pids()` takes ShareLock on the lock manager partitions per call, so it is evaluated only for backends already waiting on a lock | +| `pg_database_stats` | `pg_stat_database` | 1 min / 30 d | Four questions off one cluster-wide view: temp-file spills (`temp_files` / `temp_bytes`), cache hit ratio (`blks_hit` / `blks_read`), a server-recorded `deadlocks` count, and the commit/rollback split. **Every major and every target shape** — the newest column selected arrived in PostgreSQL 9.2, it is core rather than Aurora, and a standby is included deliberately, because a sort on a read replica spills the way it does on a writer. Stored PER DATABASE rather than aggregated, because `stats_reset` is per database: an aggregate would have to pick one reset timestamp for rows that legitimately disagree, and one database's reset would corrupt the cluster's delta with nothing left in the data to say so. `stats_reset` is captured so the read can report a reset AS a reset | +| `pg_index_usage_stats` | `pg_stat_user_indexes`, `pg_statio_user_indexes`, `pg_index` | 24 h / 90 d | Per-index scan counts, sizes, and the constraint/replica-identity/validity facts that decide droppability. **Writers only, and per database** — a replica reports its OWN scan counts rather than the writer's, so an index that is load-bearing on the primary reads as unused there, which is the single worst answer this surface can give. `last_idx_scan` is PostgreSQL 16+ and is substituted with a typed NULL below it so the row shape does not change across a mixed-version fleet. Daily rather than hourly, matching `index_object_stats`: "has anything scanned this index" is a structural question, and an hourly sample would record the same catalog facts 24 times a day at 24x the fan-out connections. 90 days of retention because the RETENTION WINDOW IS THE EVIDENCE — an index can only be called unused for as long as we have been watching it, and 30 days cannot clear a monthly report. Indexes below 64 KB are not collected (nothing to model), except invalid ones, which are a finding at any size | +| `pg_table_bloat_stats` | `pg_class`, `pg_stats`, `pg_stat_user_tables` | 60 min / 90 d | The statistics-based per-table bloat ESTIMATE, plus measured heap/TOAST/index sizes and the dead-tuple counts. **Never reads the relation**, which is what makes it affordable on a cadence: measured at 44 ms for a 2,001-table database against 860 ms for `pgstattuple` over the same tables, and that 19x is a floor on the ratio because those tables were tiny and fully cached. Hourly to match `pg_autovacuum_stats` deliberately — this measures the DAMAGE whose CAUSE that one measures, and correlating them needs a common grain. **Writers only, and per database.** Tables below 1 MB are not collected. The estimate is suppressed rather than captioned when its inputs cannot be trusted; see V85 | +| `pg_session_states` | `pg_stat_activity` | 1 min / 30 d | Which sessions are holding a transaction open, for how long, and — the part that decides whether it matters — whether that transaction is what pins the xmin horizon. **`idle in transaction` is NOT automatically a horizon holder**, which is the whole reason this stores an age rather than a state: measured on live PostgreSQL 16.15, a READ COMMITTED transaction that only read and one whose `UPDATE` matched zero rows both sit idle in transaction indefinitely with `backend_xmin` AND `backend_xid` NULL, pinning nothing at all. `horizon_age` is the GREATER of the two ages and is `-1` when the session holds neither — not `0`, which would read as "holds the newest possible xid". **No raw statement text is stored**: `pg_stat_activity.query` carries literal parameter values, so the normalised `query_id` and a whitelisted command keyword are stored instead. **A SAMPLE, like `pg_blocking`** — a transaction that opened and closed between two cycles is invisible, and the read carries its own capture counts so that is visible rather than assumed. Valid on a standby, deliberately: a standby holds its own transactions, and with `hot_standby_feedback` on its xmin propagates to the primary | + +The two Aurora-only collectors are gated on Aurora detection, not on configuration: the connect probe looks for `aurora_version()` and the gate follows what it finds. Point Darling at self-managed PostgreSQL and the four core-catalog collectors run while those two sit out. + +`pg_autovacuum_stats` additionally gates OFF on a standby (`pg_is_in_recovery()`), and the reason is worth knowing because it is not a permissions or availability problem. `pg_stat_user_tables` reads fine on a replica and reports **all zeros**: measured on Aurora 17.7, the same cluster, database and 15 tables, the writer reported 13,654,458 dead tuples and 150,790,506 live tuples while the reader reported 0 for every tuple counter. Those are the writer's stats-collector numbers and they are not replicated. Ungated, a replica target would return no rows, the activity filter would read that as "nothing has pending work", and you would get a confident report of perfect autovacuum health for a cluster 13 million dead tuples behind. For the same reason, treat an empty `get_pg_replication_slots` result from a replica as per-instance rather than cluster-wide — slots live on the writer. + +Cadences and retention are the shared defaults, with no knobs, exactly as for SQL Server. + +Three of the twelve are outage predictors rather than performance metrics, which is deliberate: PostgreSQL's most damaging failures are quiet, slow, and fully predictable days ahead, and nothing in the engine raises its hand about them. Every one of the twelve has an MCP tool (see [the tool list](#mcp)), a web tab and — since #2530 — a WPF viewer tab: a PostgreSQL target gets **seven** inner tabs (Overview, Activity, Vacuum, Waits, I/O, Replication, Storage) in place of the nineteen SQL Server ones, chosen from `collect.servers.engine_kind`. The placement is derived from the collector catalog in both directions, so a thirteenth PostgreSQL collector cannot ship without a screen. + +### What it does not do yet + +Plan capture has no PostgreSQL equivalent in the store yet, and that is the largest remaining gap — plans are why people open a database monitoring tool. Alerting and scheduled analysis are still SQL-Server-shaped, so a PostgreSQL target collects and is readable through MCP, the web dashboard and the viewer, but does not yet raise alerts or produce analysis findings. The PostgreSQL tabs on both UIs are tables over the window: no charts, no correlated timeline, and no drill-down from a blocking root to the sessions behind it. **The three Tier 0 outage predictors DO alert.** Wraparound risk, a blocked vacuum horizon and replication-slot retention are evaluated on the alert cadence and delivered through the same deliverer, history and mute rules as every SQL Server alert, so they land in the same places and obey the same suppression. They ride alongside the shared engine rather than inside it, via a separate `IPostgresAlertReadAdapter` consulted only for PostgreSQL targets — Lite has no PostgreSQL target, and extending the shared adapter would have left it implementing three methods that can only return empty. Thresholds are derived from the server's own settings (wraparound grades against that cluster's `autovacuum_freeze_max_age`, not a constant) and are not yet configurable; see [`docs/postgres-alerting-design-note.md`](../docs/postgres-alerting-design-note.md). + +Scheduled ANALYSIS is still SQL-Server-shaped, so a PostgreSQL target does not yet produce analysis findings. + +Collector FAILURES are classified, though. A PostgreSQL fault is routed through the same `ITargetProvider.Classify` the engine seam exposes, so a persistent, operator-actionable condition records as a non-fatal skip with an explanation instead of logging `ERROR` every cycle — a `pg_stat_statements` view that was never created (SQLSTATE 42P01), a source Aurora does not implement (0A000), a feature switched off in the parameter group (55006). The message says which kind it is, because the store's non-fatal bucket is named PERMISSIONS and none of those is a missing grant. Connection-level failures (the 08 class, 57P0x) still force a reconnect and reprobe; a `statement_timeout` (57014) deliberately does not, since dropping the connection over a slow query would turn a tuning problem into a reconnect storm. + +The per-database fan-out itself is done — `pg_autovacuum_stats` is the collector that exercises it. Worth knowing what it costs: a SQL Server collector can reach another database without reconnecting (`EXECUTE [db].sys.sp_executesql`), while a PostgreSQL connection is bound to one database for its lifetime, so a per-database PostgreSQL collector is necessarily one connection per database per cycle. That is why its cadence is hourly and why new per-database collectors should be added deliberately rather than by default. + +--- + +## Operations + +### The Store + +The service migrates the store itself at startup — plain versioned SQL scripts, each applied once inside its own transaction, tracked in `darling_schema_version`, safe under concurrent starters (advisory-locked). Current schema is **v73** — `StorageVersion.SchemaVersion` is the source of truth and a test pins it to the highest rung in the ladder. + +The notable rungs are below. For the **complete** current schema, read `Darling/Darling.Tests/Fixtures/migration-ladder-*.sql` — the whole ladder as resolved SQL, regenerated per release; it is generated, so don't hand-edit it. + +| Version | Contents | +|---|---| +| **V1** — collector tables | One table per collector, all 54, generated from the shared collector definitions (column-for-column identical to Lite's DuckDB schema): `wait_stats`, `latch_stats`, `spinlock_stats`, `query_stats`, `procedure_stats`, `query_store_stats`, `query_snapshots`, `plan_cache_stats`, `cpu_utilization_stats`, `cpu_scheduler_stats`, `file_io_stats`, `memory_stats`, `memory_clerks`, `memory_pressure_events`, `tempdb_stats`, `perfmon_stats`, `deadlocks`, `blocked_process_reports`, `dmv_blocking_snapshots`, `memory_grant_stats`, `waiting_tasks`, `session_stats`, `session_summary_stats`, `running_jobs`, `database_size_stats`, `index_object_stats`, `server_properties`, `system_health_events`, the four config snapshots (`server_config`, `database_config`, `database_scoped_config`, `trace_flags`), and the twenty PostgreSQL tables listed under V63–V69, V71, V83–V94 below | +| **V2** — observability | `servers` (registry, upserted on every successful connect: identity, display name, engine edition, major version) and `collection_log` (one row per collector run: SUCCESS / PERMISSIONS / ERROR, row count, SQL-phase and storage-phase timings) | +| **V3** — alerting | `config_alert_log` (one history row per fired alert), `config_edge_trigger_watermarks` (restart-surviving edge-trigger and failed-job watermarks), `config_mute_rules` (alert mute rules; starts empty) | +| **V4** — analysis | `analysis_findings` (persisted findings incl. the stored remediation action), `analysis_muted` (muted finding patterns), and 17 `v_
` passthrough views so the shared analysis SQL runs verbatim against this store | +| **V5** — viewer passthrough views | The five remaining `v_*` passthrough views (`v_running_jobs`, `v_server_config`, `v_database_scoped_config`, `v_trace_flags`, `v_collection_log`) that complete the viewer's read layer | +| **V6** — memory passthrough views | `v_memory_clerks` and `v_memory_pressure_events`, the two views the Memory tab reads | +| **V7** — plan-capture columns | Nullable plan-XML columns for the viewer's View Plan surfaces: `procedure_stats.query_plan_xml`, `blocked_process_reports.blocked_query_plan_xml` / `blocking_query_plan_xml`, `deadlocks.victim_query_plan_xml` | +| **V8** — schema split (collect/config) | Moves the tables into the `collect` and `config` schemas (least-privilege security split); the shared SQL keeps using bare names, resolved via `search_path = collect, config, public` | +| **V9** — inventory + cost fields | `server_properties` inventory columns (`sqlserver_start_time`, `host_os_version`, `ag_replica_role`) and `servers.monthly_cost_usd` (the FinOps per-server budget) | +| **V10** — latch + spinlock collectors | `latch_stats` and `spinlock_stats` tables plus their `v_*` views | +| **V11** — CPU scheduler + plan cache collectors | `cpu_scheduler_stats` and `plan_cache_stats` tables plus their `v_*` views | +| **V12** — session summary collector | `session_summary_stats` (server-wide connection-leak / idle signal) table plus its `v_*` view | +| **V13** — system health events collector | `system_health_events` (raw `system_health` Extended Events capture) table plus its `v_*` view | +| **V14** — refresh passthrough views | `CREATE OR REPLACE` on every `v_*` view so a store upgraded across a column-adding migration picks up the new columns (Postgres freezes a view's `SELECT *` expansion at create time) | +| **V15** — index metadata columns | Per-index definition columns on `index_object_stats` (ordered key/included column lists, filter, uniqueness/constraint/FK flags, `is_disabled`, and the reconstruct-a-CREATE options — compression, fill factor, page/row locks, etc.) for monitor-side UNUSED/DUPLICATE index analysis, and refreshes `v_index_object_stats` | +| **V16** — server UTC offset | Nullable UTC-offset column on `server_properties` so the viewer can render timestamps in the monitored server's own local time (the Server-time display mode ported from Lite; Server-time = stored naive-UTC + this offset) | +| **V17** — config control plane | The viewer-writable DESIRED-state tables (`config_service`, `config_monitored_servers`, `config_alert_settings`, `config_collector_schedules`) plus a `config_version` reload beacon — statement-level bump triggers increment it on any write, and the service polls that one integer each sweep and reloads only when it changes. Server secrets are DPAPI blobs, never plaintext | +| **V18** — alert delivery mode | Global `delivery_mode` (Summary / PerEvent) + `per_event_max` on `config_alert_settings`, plus a nullable per-server `alert_delivery_mode_override` on `config_monitored_servers` (null = inherit the global), resolved through the shared `AlertDeliveryModeResolver` (#1236 / #1141) | +| **V19** — analysis state marker | `collect.analysis_state` — the service-produced per-server "insufficient data" marker (with message + time) the viewer reads, so a not-enough-history analysis pass surfaces a reason instead of a blank | +| **V20** — alert tuning knobs | The previously-hardcoded alert tuning the viewer now customizes on `config_alert_settings`: the long-running-query read shape (`long_running_query_max_results` + five noise-filter opt-outs the shared `AlertEngine` forwards) and `notify_connection_changes` (the Server-Unreachable / Restored connect-edge gate) | +| **V21** — default trace events collector | `default_trace_events` table + its `v_*` view — the significant Default Trace events (file growth, ErrorLog, security audit, optional Object DDL) the viewer's System Events tab reads | +| **V22** — index-object latest index | The engine-agnostic `idx_index_object_stats_latest` partial index backing the latest-capture-per-index reads | +| **V23** — collection-log hypertable | Converts `collection_log` to a TimescaleDB hypertable (an object-invisible no-op on plain PostgreSQL) | +| **V24** — job history collector | `job_history` table + its `v_*` view — the SQL Agent Job History surface (#1433) | +| **V25** — agent status collector | `agent_status` table + its `v_*` view — SQL Agent up/down status (#1433) | +| **V26** — generic webhook channel | The generic-webhook columns on `config_notification` (`generic_url`, `generic_headers`, `generic_body_template`, `generic_proxy`) for POSTing alerts to any endpoint (#1506) | +| **V27** — deadlocks database name | `deadlocks.database_name` (the Azure SQL DB per-database deadlock-capture watermark key, #1535) and a refreshed `v_deadlocks` | +| **V28** — Query Store replica role | `query_store_stats.replica_role` (SQL Server 2022+ AG secondary-replica attribution, #1546) and a refreshed `v_query_store_stats` | +| **V29** — long-query completions collector | `collect.long_query_completions` + its index — the opt-in long-running-query completion trace's store table (#1496) | +| **V30** — web dashboard config | `config_service.web_enabled` + `web_port` — the read-only web dashboard's live enable/port toggle, the twin of `mcp_enabled`/`mcp_port` (#1562) | +| **V46** — automatic plan correction | `collect.plan_correction` + its index — the #1952 collector's store table (FORCE_LAST_GOOD_PLAN enablement plus the engine's live recommendation set). Additive and view-less, so a fresh store gets it from V1's generated schema and V46 is what an already-existing store gets | +| **V47** — ADR persistent version store | `collect.pvs_stats` + its index + the `v_pvs_stats` passthrough view — the #1951 ADR version-store collector's store table. A fresh store gets the table from V1's generated schema; V47 is what an already-existing store gets, and the view is what keeps the Darling viewer's FinOps read byte-identical to Lite's | +| **V61** — per-fingerprint occurrence counters | `config.config_incident_occurrences` — the accumulator’s memory for the monotonic count behind an alert incident (#2216). The count that rides on an incident is a GAUGE (it falls as events age out of the read window), so a consumer seeing only throttled deliveries cannot recover how many events happened between two of them. A NEW table, not columns on `config_edge_trigger_watermarks`: the key is wrong (per (server, metric) vs per (server, metric, fingerprint)) and Lite writes that row with a PARTIAL `INSERT OR REPLACE` column list, so an added column would zero itself every time an alert fired | +| **V62** — plan-XML codec knob | `config.config_service.plan_xml_compression` (#2171). `gzip` (default) keeps today’s write path; `none` stores plain text in `query_plan_xml` so direct-SQL readers get plans back — PostgreSQL exposes no inflate, so gzip bytes are unreadable without an untrusted-language UDF. Rides `config_service` like V58/V59 so the `config_version` trigger makes a flip visible to the next reload poll | +| **V63–V69** — PostgreSQL collector tables | `collect.pg_wait_stats`, `collect.pg_statement_stats`, `collect.pg_wraparound_stats`, `collect.pg_xmin_horizon`, `collect.pg_replication_slot_stats`, `collect.pg_autovacuum_stats`, and `collect.pg_io_stats`, each with its time index — one rung per PostgreSQL collector. Additive and view-less, exactly like V46/V47: a fresh store gets all seven from V1's generated schema, and these rungs are what an already-existing store gets. They add tables only, so a store that monitors no PostgreSQL target carries seven empty tables and nothing else changes | +| **V70** — monitored-server engine + port | `config.config_monitored_servers.engine` (`NOT NULL DEFAULT 'sqlserver'`) and `.port` (`NOT NULL DEFAULT 0` = the driver's default). The registry is authoritative for the server list once seeded, and these were the two `MonitoredServer` fields with no column — so a PostgreSQL target round-tripped as a SQL Server one and was connected to with `SqlConnection`. Every existing row means exactly what it meant before, and the SQL-Server-only writers keep inserting without naming either column | +| **V71** — PostgreSQL blocking edges | `collect.pg_blocking_edges` + its time index — the eighth PostgreSQL collector's store table. One row per (blocked, blocking) pair rather than a rendered tree, which is what lets the read layer compute root blocker, chain depth and fan-out in SQL instead of parsing a string. Additive and view-less exactly like V63–V69. **Sparse by design**: empty on a healthy instance, and because PostgreSQL has no engine-side blocked-process recorder, a gap means "not sampled" rather than "not blocked" — a count over this table measures how often blocking was *caught* | +| **V83** — PostgreSQL per-database counters | `collect.pg_database_stats` + its time index — the ninth PostgreSQL collector's store table (#2539): `pg_stat_database`'s temp-file, cache, deadlock and transaction counters. Additive and view-less exactly like V63–V69 and V71. Every counter column is nullable on purpose — they are cumulative counters differenced at read time, and a NOT NULL 0 default would turn "not reported" into a measurement; `database_name` is nullable because PostgreSQL genuinely emits a NULL-named row for shared relations | +| **V84** — PostgreSQL index usage | `collect.pg_index_usage_stats` + its time index — the tenth PostgreSQL collector's store table (#2541): per-index scan counts and sizes, plus the catalog facts that decide whether an unscanned index can be dropped at all. Most of the width is that second half, and it is stored rather than looked up on demand because the MCP has no ad-hoc path back to a monitored server — whatever is not captured here cannot be recovered later. `last_scan` is nullable with two distinct meanings the read separates (PostgreSQL 15 and below do not record it; on 16+ it is NULL for an index never scanned since the counters were reset), and `stats_reset` is nullable because it is genuinely NULL until a database's statistics are first reset | +| **V85** — PostgreSQL table bloat | `collect.pg_table_bloat_stats` + its time index — the eleventh PostgreSQL collector's store table (#2542): the statistics-based bloat ESTIMATE with the measured sizes and counter-based dead-tuple pair beside it. **Three tiers of certainty share the table and the column names carry the difference**: `heap_bytes`/`toast_bytes`/`index_bytes` are measured (`pg_relation_size` asks the filesystem), `live_tuples`/`dead_tuples` are the server's own counters, and `bloat_bytes_estimate`/`bloat_pct_estimate` are arithmetic over column-width statistics — suffixed `_estimate` in the store so the qualifier cannot be lost between here and a screen. `estimate_unavailable` is load-bearing rather than advisory: it is TRUE when the monitoring login cannot SELECT the table and `pg_stats` filtered every row out, a state in which the estimator does not fail but silently returns large numbers | +| **V86** — PostgreSQL session states | `collect.pg_session_states` + its time index — the twelfth PostgreSQL collector's store table (#2540): who is holding a transaction open and whether it pins the xmin horizon. **`horizon_age` is the column the table exists for**, and its `-1` is load-bearing rather than tidy: `pg_xmin_horizon` can say a session holds the horizon, and only this can say a session does NOT — which is the answer for two of the four idle-in-transaction shapes measured on a live instance. There is **no query-text column, deliberately**: `pg_stat_activity.query` carries literal parameter values and this table fills on a duration floor an ordinary application crosses, so `query_id` (joinable to `pg_statement_stats`, whose text is already `$1`-normalised) and a whitelisted `command_tag` carry the statement identity instead. `state_is_redacted` exists because without `pg_monitor` PostgreSQL does not refuse the read — it returns every row with the state columns NULL while leaving `backend_xmin` and `backend_xid` visible, so the horizon still reads as pinned and nothing can say by what | +| **V87** — PostgreSQL plan-capture readiness | `collect.pg_plan_capture_readiness` + its time index (#2564): whether a target could capture execution plans at all, and if not, WHICH step is missing. One row per facet rather than a single verdict, because “no plans” has several unrelated causes — the module not preloaded, a threshold that captures nothing, a log line prefix that cannot attribute a plan to a query — and each has a different remedy, carried on the row | +| **V88** — PostgreSQL write side | `collect.pg_write_stats` + its time index (#2544): checkpoints, background writing and WAL from `pg_stat_checkpointer`, `pg_stat_bgwriter` and `pg_stat_wal`, as ONE collector because they are one story. A union schema with version-conditional column expressions: PostgreSQL 17 moved checkpoint counters out of `pg_stat_bgwriter`, so a major that lacks a column reports NULL rather than the collector splitting into per-version twins. `wal_bytes` is `numeric(38,0)` because that is how upstream types it | +| **V89** — PostgreSQL extension availability | `collect.pg_extension_availability` + its time index (#2545): four states, not a boolean — `installed`, `outdated`, `available`, `absent`. The distinction is what makes it actionable: `available` is a `CREATE EXTENSION` away and `absent` is not. Absence is not a row in any catalog, so it is derived against an enumerated roster; `auto_explain` and `pg_wait_sampling` are deliberately off that roster, being preload-only modules that never appear in `pg_available_extensions` even where they are loaded and working. Runs per database as of V95 | +| **V90** — PostgreSQL lock states | `collect.pg_lock_stats` + its time index (#2544): lock state by mode, type and relation. Does NOT duplicate `pg_blocking_edges`, which stores blocked/blocker PAIRS — that answers who is stuck behind whom and cannot answer what lock, on what object, is being waited on | +| **V91** — PostgreSQL column statistics | `collect.pg_column_stats` + its time index (#2543): the planner inputs behind a misestimate — `n_distinct`, `null_frac`, `avg_width`, `correlation`. **The SHAPE of the skew, never the values**: `most_common_freqs[1]` and `cardinality(most_common_vals)` are stored and `most_common_vals` / `histogram_bounds` deliberately are not, because those hold raw customer data and every finding this table exists for survives without them. `n_distinct` is a floating type rather than a count because negatives are a RATIO of row count. **Needs `pg_read_all_data`**: `pg_stats` filters on `has_column_privilege` and `pg_monitor` confers no SELECT on user tables, so without that grant this collector succeeds and returns zero rows | +| **V92** — PostgreSQL replication | `collect.pg_replication_stats` + its time index (#2544): connected standbys and how far behind each one is — four byte distances measured from `pg_current_wal_lsn()` and three time lags. The lag columns keep growing during a stall rather than freezing, which is what makes them usable as an alerting signal | +| **V93** — PostgreSQL buffer residency | `collect.pg_buffer_usage` + its time index (#2544): what is resident in shared buffers, by relation — a hit ratio says how often the pool worked, this says what is IN it. Joins `pg_buffercache.relfilenode` to `pg_relation_filenode(c.oid)`, **not** to `c.oid`, which is the join every published example gets wrong, and scopes to the connected database because the pool is cluster-wide while `pg_class` is not | +| **V94** — PostgreSQL index bloat | `collect.pg_index_bloat` + its time index (#2561): b-tree index bloat MEASURED via `pgstatindex`, not estimated. The estimator route is blind under this product's permissions — it needs `pg_stats`, which returns nothing to a `pg_monitor` role — while `pgstatindex` runs, because pgstattuple grants EXECUTE to `pg_stat_scan_tables`. `avg_leaf_density` is stored RAW and never converted to a percentage, and `skipped_reason` exists so a size ceiling can never masquerade as an absence of bloat | +| **V95** — PostgreSQL per-database attribution | `database_name` on `collect.pg_column_stats`, `collect.pg_index_bloat` and `collect.pg_extension_availability` (#2599), and `pg_extension_availability` now runs per database. Two of these collectors ran per database and could not say WHICH database a row described, so on a cluster carrying one schema in two databases their rows collided; the third read the per-database `pg_extension` catalog through a single connection and reported one database's answer for the whole server. Nullable with no backfill — a catalog-only change in PostgreSQL, and rows collected before this rung genuinely do not know where they came from | +| **V72** — Query Store plan map | `collect.query_store_plan_map` — `(server_id, database_name, plan_id)` → digest, so Query Store facts can reference plan XML they no longer carry once that content moves into the shared `query_plan_dim`. Plan XML was stored INLINE on `query_store_stats` at roughly 5x redundancy. Not a hypertable: one row per distinct plan per database, so it is dimension-shaped and pruned on `last_seen` rather than by `drop_chunks`. Its `last_seen` is load-bearing — the dimension GC sweeps on timestamps rather than counting references, so ending the re-shipping also ends the liveness signal that used to keep those dim rows alive | +| **V73** — PostgreSQL statement text | `collect.pg_statement_text` — `(server_id, queryid)` → statement text, refreshed hourly, so `get_pg_top_queries` returns something readable (#2219). `pg_statement_stats` stores no text because `showtext` is a real per-collection cost and normalized text is highly repetitive; but `queryid` is NOT stable across a major version upgrade, so without this the stored history joins to nothing after one — a list of integers that used to be your slowest queries, unrecoverable because the live view no longer holds the old ids. Text is INLINE rather than a `query_text_dim` digest: the dimension route needs the GC liveness interlock whose failure mode is silently missing text, and inline cannot dangle. Not a hypertable and not a collector table, exactly like V72 — a bespoke upsert path, pruned on `last_seen` with a margin that makes text OUTLIVE the statistics referencing it | +| **V76** — Query Store health | `collect.query_store_health` + its index + the `v_query_store_health` passthrough view — the #2319 per-database `sys.database_query_store_options` collector's store table: actual vs desired state (the cap-hit READ_ONLY transition and its readonly_reason), current vs max storage, cleanup thresholds, and the runtime-stats interval length. A fresh store gets the table from V1's generated schema; V76 is what an already-existing store gets | +| **V77** — Activity-driven plan fetch | Three strokes behind #2312's reshape of the Query Store plan/text fetch: `query_store_plan_map.digest` goes **nullable** (a plan whose XML the engine cannot persist gets a NULL-digest map row — the content-less marker that stops the probe re-selecting it forever), `query_store_text` gains `query_hash` (the Query Store reset detector: an id whose stored hash differs from the live one names a DIFFERENT statement now and its text refetches within one cycle), and the retired `planwm:`/`textwm:` watermark state rows are deleted wholesale. The fetch itself no longer walks the plan catalog by watermark — the cycle's collected rows name their plans, the store answers which are missing, and only those are fetched | + +All timestamps in the store are **naive-UTC** `timestamp` columns — the product-wide cross-store contract (Lite's DuckDB does the same). + +### Reading the store directly (plan XML is compressed) + +The store is deliberately queryable — it is documented PostgreSQL with named tables, and people build +panels and reports straight off it. One thing will surprise you if you do that: **execution-plan XML is +stored gzip-compressed**, and has been since v3.4.0. + +`collect.query_plan_dim` holds plan content once, keyed by a content digest, in one of two columns: + +| Column | Meaning | +|---|---| +| `query_plan_gz` (`bytea`) | The plan XML, **gzip-compressed** (magic bytes `1f 8b`). This is where new plans go. | +| `query_plan_xml` (`text`) | Uncompressed plan XML. Nullable since v3.4.0; only rows written by older builds still carry it. | + +So a consumer that reads only `query_plan_xml` silently returns nothing for anything collected by a +current build. **`query_plan_xml IS NULL` does not mean "no plan" — it means look at `query_plan_gz`.** + +Both apps and every MCP tool decompress client-side, so nothing in the product is affected; this note +exists because the change altered the contract for direct SQL consumers and the v3.4.0 release notes did +not say so. That omission is on us. + +**Getting the XML back.** PostgreSQL has no built-in gunzip for arbitrary `bytea`, so a plain-SQL +consumer cannot decompress in the database without an extension. Practical options, in the order most +people should try them: + +1. **Ask the product for the plan** rather than the store — `get_plan_xml` over MCP, or the Viewer's + plan surfaces. Both hand back decompressed XML and neither cares how it is stored. +2. **Decompress in your client.** Any language's gzip library reads the bytes directly. Python: + `gzip.decompress(row['query_plan_gz']).decode('utf-8')`. PowerShell: a `GZipStream` over a + `MemoryStream` of the bytes. C#: the same, which is exactly what the apps do. +3. **Ship a UDF into your own store** if your tooling is SQL-only (Grafana, a reporting view). A + `plpython3u` function works and has been used in the field, at the cost of an untrusted-language + extension in a monitoring database — weigh that against how much you need it. + +Why compressed at all: plan XML dominates store size, and gzip took a production dim table from 885 GB +of raw text to 64 GB — a 14x reduction. That is the tradeoff being made on your behalf. + +### TimescaleDB (Optional, Auto-Adopted) + +At startup, right after migration, the service attempts `CREATE EXTENSION IF NOT EXISTS timescaledb` and checks `pg_extension`: + +- **Present** — every collector table is converted to a hypertable (partitioned on its own time column into **1-day chunks**, existing rows migrated) and gets a compression policy: chunks older than **1 day** compress automatically (segmented by `server_id`), checked **hourly**. The hourly tick is passed explicitly because TimescaleDB's own default is **12 hours** for 1-day chunks — that is a second, separate wait *after* a chunk is already eligible, and on a field store it left the newest closed chunk (always the least-compressed data on disk) uncompressed for most of a day. Stores created before this shipped are retuned automatically on the next service start. The short intervals matter at the 1-minute collection cadence — a chunk cannot compress until it closes and then ages, so TimescaleDB's 7-day default left the store fully uncompressed for ~2 weeks (a near-idle 5-server fleet still reached ~1 GB in a couple of days); 1-day chunks + 1-day compress keep it compact (measured ~16.7x on perfmon, ~6.4x on the plan-XML-heavy query_stats). Compressed chunks stay fully queryable — this is Darling's archival tier, the centralized-store answer to Lite's Parquet archive. Everything is idempotent and re-converges on every service start; a table that fails conversion stays a plain table and keeps working. +- **Absent** — the service logs one Information line and runs in plain-PostgreSQL mode, which is a fully supported configuration, not a degraded one. + +`IF NOT EXISTS` short-circuits before privilege checks, so a store whose administrator pre-created the extension works for a service login that could never create it. + +### Background workers: sizing an unmanaged store, and what happens if you don't + +**This section is for bring-your-own PostgreSQL only.** In managed mode the service sizes these itself on every start and there is nothing to do. + +Every TimescaleDB policy — compression, retention, continuous-aggregate refresh — runs in a **background worker**, and a policy that cannot get a worker does not run. PostgreSQL's stock `max_worker_processes = 8` is far below what this store needs, so an unmanaged store left at the defaults silently does very little compressing. + +Managed mode derives the two settings from the live hypertable count, and an unmanaged store wants the same numbers: + +``` +timescaledb.max_background_workers = + 2 +max_worker_processes = 3 + timescaledb.max_background_workers + 8 +``` + +Today that is **57** and **68** for 55 hypertables (the 54 collector tables plus `collection_log`). The `+ 2` is not slack — it is exactly TimescaleDB's own two built-in jobs, `policy_telemetry` and `policy_job_stat_history_retention`, so a fully migrated store holds precisely one job per worker: + +```sql +SELECT proc_name, count(*) FROM timescaledb_information.jobs GROUP BY proc_name; +``` + +Both settings need a **server restart** (`max_worker_processes` is restart-only — a reload leaves the old value serving), and the hypertable count grows as collectors are added, so re-check it after a major upgrade rather than pinning 57/68 forever. + +**One store per cluster is the assumption.** `timescaledb.max_background_workers` is a **cluster-wide** pool shared by every database, while the derivation above is **per-store**. Managed mode puts one store on one cluster so the two coincide, but if you run **N Darling stores on one PostgreSQL cluster** — or share the cluster with any other TimescaleDB database — multiply both numbers by N. Each database with the extension loaded also permanently holds a scheduler slot out of that same pool, so the sharing starts before any policy fires. + +**What under-provisioning looks like.** The postmaster log (`pg.log`, or wherever your cluster logs) is where it shows up, in one of two shapes: + +``` +WARNING: failed to launch job 1042 "Columnstore Policy [1042]": out of background workers +WARNING: ... failed to start a background worker +``` + +The first means TimescaleDB's own pool is full; the second means PostgreSQL's is. Neither is fatal and neither corrupts anything — the job is skipped and retried on its next schedule, so **light contention is benign** and you may see a couple of these without any consequence. It matters at scale: when the shortfall is persistent rather than momentary, compression falls behind the 1-day policy and the store grows at its uncompressed rate (measured compression is ~16.7x on perfmon and ~6.4x on the plan-XML-heavy `query_stats`, so the gap is large), retention stops reclaiming chunks, and the jobs that keep losing the race are the ones whose backlog is worst. `timescaledb_information.job_stats` is the check that settles it — a healthy store shows successes with no failures: + +```sql +SELECT sum(total_runs), sum(total_successes), sum(total_failures) FROM timescaledb_information.job_stats; +``` + +### Retention + +A purge runs on the first sweep after startup and then daily, driven by the same shared per-collector horizons Lite uses: + +| Horizon | Tables | +|---|---| +| 7 days | `query_snapshots`, `waiting_tasks`, `running_jobs` | +| 30 days | Most collector tables (wait/query/procedure/Query Store stats, CPU, memory, file I/O, tempdb, perfmon, deadlocks, blocking, sessions, config snapshots), plus `collection_log` and `analysis_findings` | +| 90 days | `database_size_stats`, `index_object_stats`, `pvs_stats` | +| 365 days | `server_properties` | + +On plain PostgreSQL the purge is DELETE-based. With TimescaleDB it switches to `drop_chunks` — a metadata-only detach of whole expired chunks (rows inside a partially-expired chunk survive until the whole chunk ages out; up to ~1 day of grace at the 1-day chunk width), with a per-table DELETE fallback for any table that is not a hypertable. Failure-isolated per table: one stuck purge is logged and retried the next day without stopping the sweep. + +#### The rollup tiers, on a TimescaleDB store + +The table above is the **collector** horizon, and for three tables it is not the binding one. `query_stats`, `procedure_stats` and `query_store_stats` are rolled up into hourly and daily continuous aggregates, and a separate tiered policy drops their raw chunks at **4 days** — the aggregates hold the history past that point, and a read is routed to whichever tier covers the window it asks for. On a store without TimescaleDB none of this exists and the collector horizons above are the whole story. + +| Tier | Horizon | +|---|---| +| Raw `query_stats`, `procedure_stats`, `query_store_stats` | 4 days | +| Hourly **history** rollups | 90 days | +| Daily **history** rollups | kept indefinitely (no policy) | +| Baseline aggregates | 35 days | +| `query_store_stats_interval_hourly`, `query_store_stats_interval_daily` | 7 days, 10 days | + +Every one of these is visible in `timescaledb_information.jobs`, and the last row is the one worth knowing before you look: those two are **internal dedup plumbing, not history**. The corrected Query Store rollups are built from them, nothing reads them directly, and each horizon is sized only to outlive whatever gates on it — 7 days has to exceed raw's 4, and the 10-day layer has to outlive the 7-day one it consumes. So a horizon SHORTER than the tier above it is correct there and costs no history, which is the opposite of how it reads at a glance. The service's startup summary line names all of these for the same reason. + +No raw tier is ever dropped before the aggregate that preserves it has caught up: each policy is created paused, and arms itself only once its rollup demonstrably covers what the tier below holds. + +### Logs + +The service's PRIMARY log is a **rolling file** under `%ProgramData%\PerformanceMonitorDarling\logs\darling-service_yyyyMMdd.log` — every collector run line, connect edge, reload notice, warning, and error lands there (buffered writes, one file per day, 14-day retention, and a logging failure can never crash the service). Console runs write the same file plus console output. + +Warnings and errors also go to the **Windows Application event log** (source `PerformanceMonitor Darling`) — but only if that event source exists. Registering an event source requires elevation, and the recommended `NT SERVICE` virtual account cannot do it, so run the `New-EventLog` line in the install steps above (or any elevated run of the exe) once; without it, Windows silently drops the events and the file log is your only surface. Collection outcomes are also queryable in the store itself — `collection_log` records every collector run per server with status and timings, and the viewer's Collection Health tab renders exactly that. + +### The Viewer + +`PerformanceMonitor.Darling.Viewer.exe` is a WPF app that talks **only to the PostgreSQL store** — it never connects to your monitored SQL Servers. It reads the same `darling.json` the service uses, but only the `postgres` section, resolved in the same order (explicit path, then `DARLING_CONFIG`, then `darling.json` next to the binary) plus one viewer-only fallback: the parent directory, so the release zip's layout — viewer in a `viewer\` subfolder, `darling.json` beside the service exe — works with no setup. A viewer seat on **another machine** is set up by exporting that config folder from the service host — see [Connect a Remote Viewer](#connect-a-remote-viewer). If the file is missing it shows a hint instead of crashing. + +At startup the viewer writes **which of those rules won**, the absolute path it produced, and whether that file exists to `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log` — before it tries to read the file, so a missing or malformed one still says where it looked. Once the file loads it adds a non-secret summary of what it parsed (host, port, username, database, SSL mode, search path, whether the connection string was read verbatim or derived from `postgres.managed`, and the certificate — the value as written, the absolute path it resolves to, the folder a relative one was anchored to, and whether that file exists). Credentials are never written. The same block appears in the connection-failure window with a **Copy details** button — see [Troubleshooting](#troubleshooting). + +The layout mirrors the Lite desktop app: a left sidebar lists the servers from the `servers` registry the service maintains, and the top tab strip holds three fixed **aggregate tabs** — Overview, Recommendations, and Alerts — alongside a closable **per-server tab** for each server you open. Overview (the all-servers server-cards grid) and Alerts (the all-servers alert history) span every server; Recommendations has its own server selector, independent of the sidebar. **Double-click a server** in the sidebar — or **double-click its Overview card** — to open (or focus) its tab, and close it with the × on the tab header; an empty-state panel is shown until the store has at least one server. + +Each per-server tab has fourteen inner tabs: + +| Inner tab | Contents | +|---|---| +| **Overview** | Five correlated, X-axis-synced timeline lanes over the last 24 hours — CPU % (SQL Server vs SQL+other Total), total wait ms/sec, blocking + deadlocking, buffer pool MB, and file-I/O latency — each with a ±2σ baseline band and anomaly markers, all sharing one crosshair so a spike in one lane lines up against the others | +| **Wait Stats** | A searchable wait-type picker (poison + usual-suspect + `PAGELATCH_` defaults, checked-to-top, a 30-type selection guide) beside a per-**type** trend chart for the checked types over the last 24 hours, with a Wait Time (ms/sec) ↔ Avg Wait Time (ms/wait) metric toggle — the per-type companion to the Overview's single total-wait lane | +| **Queries** | Six sub-tabs over the last 24 hours — **Performance Trends** (a 2×2 of per-second trend charts: query duration, procedure duration, Query Store duration, execution count), **Active Queries** (the ~26-column filterable snapshot grid of captured running queries with a time-range slicer, a **Latest Snapshot** button that re-reads the newest stored capture, and per-row Estimated / Actual plan buttons that open the stored plan in the Plan Viewer), **Top Queries by Duration** (the full query-stats grid with in-grid bar cells for executions/CPU/duration/reads and a CPU-by-database breakdown), **Top Procedures by Duration**, **Query Store by Duration**, and **Query Heatmap** (query counts per 5-minute bin × per-execution magnitude bucket, by a chosen metric; right-click a cell to drill into Active Queries for that window) — the three grids each carry a time-range slicer (drag to narrow the window) and a shared **Compare** control that overlays the current window against a baseline period (yesterday, last week, or same day last week), flagging new and vanished queries | +| **Plan Viewer** | Hosts execution plans as closable sub-tabs (the shared plan-viewer control, the same one Lite and the Dashboard use). Right-click a **Top Queries** or **Query Store** row and choose **View Plan** to open the plan the service captured for it (`query_stats.query_plan_xml` / `query_store_stats.query_plan_text`); Top Queries rows also carry a **Query Plan** column whose Download button saves the stored plan as a `.sqlplan` file (enabled only when a plan was captured). Top Procedures and the blocking / deadlock reports deliberately do **not** surface a plan here — procedure plans aren't stored, and blocked-process / deadlock rows carry only a `sql_handle` (not plan XML); resolving either to a plan needs a live SQL connection the viewer never makes. "Get Actual Plan" (a live re-execution) is likewise out | +| **CPU** | Raw per-sample CPU utilization (SQL Server vs other processes) over the last 24 hours — every ring-buffer sample, full-bleed as two series; the Overview's CPU lane plots the same raw samples compactly (SQL vs SQL+other Total) with a baseline | +| **Memory** | Four sub-tabs over the last 24 hours — **Overview** (a summary strip of physical / SQL Server / target / buffer pool / plan cache / page-file memory plus the system memory state and model, over a Total-vs-Target-vs-Buffer-Pool memory trend with a memory-grants overlay), **Memory Clerks** (a searchable clerk-type picker — top-5 default, checked-to-top, clear-only-the-filtered — beside a per-clerk memory trend for the checked clerks with a non-buffer-pool total and top-clerk summary), **Memory Grants** (per-resource-pool grant sizing — available / granted / used MB — and activity — grantees / waiters / timeouts / forced grants), and **Memory Pressure Events** (hour-bucketed stacked bars of `RING_BUFFER_RESOURCE_MONITOR` pressure, SQL Server vs OS, medium vs severe) | +| **File I/O** | Two sub-tabs over the last 24 hours — **Latency** (per-file read and write latency, with a dashed queued-I/O overlay) and **Throughput** (per-file read and write MB/s) — the top 10 files by activity | +| **tempdb** | Three stacked charts over the last 24 hours — space usage (user / internal objects / version store), total allocated size, and per-file I/O latency | +| **Blocking** | Four sub-tabs over the last 24 hours — **Trends** (lock-wait rate, blocking incidents, deadlocks), **Current Waits** (waiting-task duration by wait type, blocked sessions by database), **Blocked Process Reports** (the full ~25-column filterable grid — XE reports preferred with the always-on DMV blocking snapshot merged in as fallback, each row badged with its source, a time-range slicer, per-row report-XML save, and long-block highlighting; double-click or right-click **View Block Chain** to reconstruct and draw the blocking chain the row belongs to), and **Deadlocks** (one filterable row per process parsed from each deadlock graph, a slicer, per-row graph-XML save; double-click or right-click **View Deadlock Graph** to draw the deadlock graph) | +| **Perfmon** | A searchable counter picker with the shared counter packs (General Throughput, Memory Pressure, CPU / Compilation, I/O Pressure, TempDB Pressure, Lock / Blocking) beside a per-counter delta trend for the checked counters (up to 12) over the last 24 hours | +| **Running Jobs** | Latest snapshot of currently-running SQL Agent jobs — start time, current vs average vs p95 duration, % of average, and a highlighted row when a job is running past its p95 (a store-derived banner appears when the service's login lacks msdb access) | +| **Configuration** | Four column-filterable snapshot grids of the server's latest capture — server configuration (`sys.configurations`), database configuration (28 columns of `sys.databases`), database-scoped configuration, and trace flags | +| **Daily Summary** | A one-row roll-up of the selected day (default today, UTC, with a date picker) — total wait time, the top wait type, distinct query count, deadlock / blocking-event / high-CPU-sample counts, collector errors, and an overall health band | +| **Collection Health** | Three sub-tabs — **Health Summary** (a 7-day per-collector roll-up: run / success / error counts, failure rate, average duration, last success / run / error, and a health band of HEALTHY / WARNING / STALE / FAILING / NEVER_RUN / NO_PERMISSIONS — double-click a collector to open its full run history), **Collection Log** (the recent run log with per-run SQL and store-write timings and row counts), and **Duration Trends** (a per-collector success-duration scatter) | + +The three aggregate tabs — **Overview** and **Alerts** span every server; **Recommendations** has its own server selector, independent of the sidebar: + +| Tab | Contents | +|---|---| +| **Overview** | A card per registered server (all servers, not the sidebar selection): server name + status dot, CPU (total non-idle with the SQL-only number alongside), memory, blocking and deadlock counts over the last hour, and last-collection time, each colour-banded (CPU ≥ 80% red / ≥ 50% amber / green; blocking and deadlocks red-or-amber when present) with a red **Offline** overlay. Status is derived from **collection freshness** — the newest `collection_log` age — rather than a live ping (the viewer never connects to the monitored servers): fresh is Online, older than twice the fastest collector's one-minute cadence is a Warning, and no recent collection is Offline. **Double-click a card** to open that server's tab. Refreshes every 30 seconds | +| **Recommendations** | The latest analysis run's findings for the tab's **own selected server** — a server selector independent of the sidebar, a Refresh button, and a status line showing the last analysis time — re-skinned to Lite's advise-only **card** design: a scrollable list of collapsible **incident** sections, each holding severity-banded cards (a severity badge, the affected `[database]`, the title, and the advice). Every card offers **Ask AI** (copies an MCP investigation prompt referencing `analyze_server` / `get_analysis_findings`); a card whose stored remediation carries a copy-paste statement also offers **Copy fix** (copies the suggested T-SQL). Advise-only — the viewer never applies anything, and there is no mute affordance here (alert muting lives on the Alerts surface). There is no in-app "Generate now": the service runs analysis on its own 30-minute cadence, so the status line surfaces the last analysis time instead | +| **Alerts** | The full alert history from `config_alert_log` across **all servers** (newest first, selectable time range), with a Server column and a Server filter. Double-click a row (or **View Details**) for a modal detail window showing the alert's stored detail and structured advice / remediation / drill-down from its dedup-fingerprint context. **Dismiss Selected / Dismiss All** hide alerts from the view (a durable `dismissed` flag on `config_alert_log`); column filters, Copy Cell/Row/All, and Export to CSV match Lite's grid. Right-click to **Mute This Alert** or **Mute Similar** (metric-only), and a **Manage Mute Rules** button opens the mute-rule editor | + +Only the visible tab loads (Lite's visible-only rule). The Alerts tab and the visible server tab's active inner tab refresh every 60 seconds; the Overview refreshes on its own faster 30-second timer (Lite's Overview cadence); and **Recommendations** refreshes on tab activation, its Refresh button, and its own server-selector change only, never on the timer — its findings change on the service's 30-minute analysis cadence, so a 60-second auto-refresh would be pointless churn (and would reset the incident expanders under the reader), matching Lite. + +The viewer is read-only over collected data, but it does perform a small set of **user-initiated writes** — and those go straight to the PostgreSQL store, which is the coordination point (the service honors them on its next read; there is no viewer-to-service channel). From the Alerts tab, creating a mute rule from an alert (**Mute This Alert** / **Mute Similar**) or adding, editing, toggling, deleting, or purging one via **Manage Mute Rules** writes `config_mute_rules` (a rule scopes to a server by name, exactly as Lite's mute rules do); and **dismissing alerts** sets the `dismissed` flag on `config_alert_log` so they drop out of the Alert History view (a single atomic UPDATE — Darling has no parquet archive tier, so there is no dismissed-archive sidecar). The viewer never writes collector data. + +### Restart Semantics + +The service is built to restart cleanly, any time: + +- **Delta continuity** — delta-based collectors (wait stats, file I/O, perfmon, memory grants) re-seed their baselines from the store at startup, so the first cycle after a restart produces real deltas instead of zeroes. +- **Alert no-re-fire** — edge-trigger watermarks and the failed-job watermark persist in `config_edge_trigger_watermarks`, and per-alert cooldowns re-seed from `config_alert_log`, so a restart does not replay alerts you already received. +- **Idempotent store setup** — migrations are versioned and skip what is already applied; TimescaleDB conversion and compression policies re-converge as no-ops. +- **Per-connect snapshots** — the on-connect config snapshot collectors run once per (re)connect, mirroring Lite's server-open behavior. +- Mute rules (`config_mute_rules`) load once at service startup — restart the service after adding rows. + +A monitored server that is down is retried every 60 seconds forever; a collector that errors is logged and retried at its next scheduled time; a mid-cycle connection-level failure forces a clean reconnect and re-probe. The loop never dies for one bad cycle. + +--- + +## Connect a Remote Viewer + +For the person sitting at a machine with **nothing installed on it**, whose only goal is looking at a Darling service that already runs somewhere else. Three steps, nothing to hand-edit. + +**The one service-side prerequisite.** The store has to be reachable from your LAN — a `postgres.network` block on the service host, which the `--configure-network` wizard writes for you. A store still on its loopback default accepts no remote viewer at all, and no amount of viewer-side configuration changes that. See [Store endpoint (viewer over the LAN)](#store-endpoint-viewer-over-the-lan) for that side; everything below assumes it is done. + +### 1. Export the handoff folder (on the service host) + +``` +PerformanceMonitor.Darling.Service.exe --export-viewer-config +``` + +It writes the viewer machine's **whole configuration folder** — connection string resolved, certificate copied, every field documented in place: + +``` +viewer-config\darling.json the complete viewer config: the resolved connection string and + "managed": false already set, every field explained in comments + IN the file +viewer-config\server.crt the store's TLS certificate, the file the connection pins +viewer-config\README.txt the same field reference in plain text, including the valid + "Root Certificate=" values and the one-line install instruction +``` + +The folder lands beside the service's own `darling.json` by default. Pass a directory to put it elsewhere (`--export-viewer-config D:\handoff`), and `--config ` if `darling.json` is not where the service would resolve it. + +**The exported `darling.json` contains a live database password** — that is what the viewer authenticates with. The verb says so before it writes, ACLs the file to SYSTEM + Administrators + the account running it + INTERACTIVE (the Viewer reads it interactively, the same posture as the admin/viewer credentials), and confirms the ACL took: if the secret is still readable by ordinary users it says so and exits non-zero. Copy the folder over a channel you trust and keep it ACL'd on the viewer machine. + +The verb refuses rather than clobbers: it will not export into the **service's own config directory** (that would overwrite the service's `darling.json` with the viewer's, destroying its servers, encrypted passwords and tokens), will not overwrite a file it did not write, and will not follow a junction or symlink. A destination it cannot use is named in the refusal. + +### 2. Copy the folder to the viewer machine + +Put the three files **next to `PerformanceMonitor.Darling.Viewer.exe`** — that works with nothing edited. (The Viewer ships in the same release zip as the service, in its `viewer\` subfolder; from source it is `dotnet build Darling/PerformanceMonitor.Darling.Viewer/PerformanceMonitor.Darling.Viewer.csproj -c Release`.) + +To keep the folder somewhere else instead, point the `DARLING_CONFIG` environment variable at the exported `darling.json`. That works unedited too: a bare or relative `Root Certificate` resolves against **the folder holding `darling.json`**, so the `server.crt` beside it is found wherever you keep the folder ([#1970](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1970)). Keep the three files together and the folder can live anywhere. + +### 3. Start the Viewer + +That is the whole setup. Re-run the export after a credential or certificate rotation — the store's certificate regenerates when its bind IP changes — and copy the folder over again; it replaces its own previous output without ceremony. + +### If it does not connect + +The failure window carries a **Configuration this viewer used** block naming the `darling.json` it actually read, which rule picked it, and the host, port, username, database, SSL mode, search path and certificate path it parsed — with a **Copy details** button, and the same lines in `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log`. Read it before changing anything: it separates *the viewer read a different file than you edited* from *it read your file and a value in it is wrong*. It never contains a password. See [Troubleshooting](#troubleshooting) for the individual failures. + +### Manual configuration (fallback) + +Only for the case where you want the connection string itself — to paste into a config that already exists, or to check what the viewer will dial. The export above is the supported path; this one is the same values, assembled by hand. + +``` +PerformanceMonitor.Darling.Service.exe --print-viewer-connection +``` + +It decrypts the `network.role` credential and prints a paste-ready connection string plus the server certificate PEM. Every warning is printed **before** the payload, but the payload is still a **live database password on STDOUT** — redirect it to an ACL'd file or pipe it to the clipboard (`... --print-viewer-connection | clip`); do not leave it in shell scrollback, CI logs, or a screenshare. The minimal viewer `darling.json` it targets is bring-your-own mode with the string pasted in verbatim (the string is consumed as-is), and the emitted PEM saved where `Root Certificate` points: + +```json +{ + "postgres": { + "managed": false, + "connectionString": "Host=192.168.1.205;Port=5641;Username=viewer;Password=...;Database=darling;Search Path=collect,config,public;SSL Mode=VerifyFull;Root Certificate=server.crt" + } +} +``` + +`"managed": false` is not a typo next to the service's `"managed": true`: the flag says who **owns** the PostgreSQL, not who is connecting. A viewer left on `true` goes looking for a bundled local PostgreSQL that is not there. (The export sets it for you, which is the point.) + +**`Root Certificate=` — what the field accepts.** It is a path to the PEM the connection validates the store's certificate against, and under `SSL Mode=VerifyFull` it is what makes the check meaningful. A relative value anchors to **the folder holding the `darling.json` the viewer read**, never the process working directory, so how the Viewer was launched cannot change the answer: + +| Value | Resolves to | +|---|---| +| `server.crt` | that name in the folder holding `darling.json` — the exported layout, correct wherever the folder lives | +| `certs\server.crt` | same anchor, one level down | +| `C:\Darling\server.crt` | an absolute path, used exactly as written, for a certificate kept somewhere else | +| omitted | nothing viewer-side to pin against: the store's certificate must already chain to a root the machine trusts. A managed store's certificate is **self-signed**, so it never does — omitting the field there fails `VerifyFull` | + +**Where the certificate comes from.** In managed mode the service generates `server.crt` / `server.key` **beside the data directory** (`%ProgramData%\PerformanceMonitorDarling\pg\` unless you set `postgres.dataDirectory`), with an IP SAN for the `network.listen` address and a DNS SAN for the machine hostname. It **auto-regenerates if the bind IP changes**, so verify-full keeps working after a `listen` change — and every viewer must then re-copy the new certificate, because an old copy stops matching. To rotate on demand, delete the pair beside the data directory; the service regenerates it on its next start. + +**Bring-your-own PostgreSQL.** Darling generates no certificate — your PostgreSQL's TLS is yours to configure — so `Root Certificate` points at the PEM that signed **your** server's certificate (the CA certificate, or the server's own certificate if it is self-signed), exactly the file you would hand `psql` as `sslrootcert`. The same relative-path anchoring applies, so keeping it beside `darling.json` is still the simplest layout. + +**Plaintext at rest on the viewer machine.** However you get there, the connection string holds the role password in cleartext in that machine's `darling.json` (there is no client-side secret store yet). That is acceptable for the read-only `viewer` credential on a single-operator, ACL'd profile; if you use `role: "admin"`, treat that file as a secret and NTFS-ACL it to your account. DPAPI-encrypting the viewer's BYO connection string is future hardening, out of scope today. + +--- + +## Troubleshooting + +**"Cannot load configuration"** (critical, service idles) — no `darling.json` was found at the resolved path. The message names the path it tried; copy `darling.sample.json` there or point `DARLING_CONFIG` at your file. + +**"Configuration problem: ..."** (critical, service idles) — validation failed. The messages are literal and per-field, e.g. `postgres.connectionString is required.`, `servers must contain at least one entry.`, `server 'X': host is required.`, `server 'X': sql auth requires username.`, `server 'X': sql auth requires encryptedPassword (preferred; see --encrypt-password) or password.`, `server 'X': auth must be 'integrated' or 'sql'`. Fix the file and restart the service. + +**"Cannot reach or migrate the Postgres store"** (critical, service idles) — the store connection string is wrong, PostgreSQL is down/unreachable, or the login cannot create tables. Collection does not start until this succeeds; fix and restart. + +**"uses a plaintext password in darling.json"** (warning, every connect) — you set `"password"` instead of `"encryptedPassword"`. It works, but run `--encrypt-password` on the service machine and switch. + +**DPAPI decrypt fails after moving darling.json** — `encryptedPassword` blobs are machine-bound (DPAPI LocalMachine). Re-run `--encrypt-password` on the new machine. + +**"Failed to ensure XE sessions"** — the login lacks `ALTER ANY EVENT SESSION` (or the database-scoped equivalent on Azure SQL Database). Deadlock and blocked-process collection read zero rows until the sessions exist; grant the permission or have an administrator create/start `PerformanceMonitor_Deadlock` and `PerformanceMonitor_BlockedProcess`. "Already exists / already started" XE errors are logged as benign and mean the sessions are up. + +**Blocked-process reports empty** — the blocked-process threshold may still be 0. On AWS RDS set `blocked process threshold (s)` via a Parameter Group (the `sp_configure` bootstrap cannot run there); on Azure SQL Database the threshold is fixed at 20 seconds. Blocking stays visible either way through the always-on DMV blocking snapshot. + +**`PERMISSIONS` rows in `collection_log`** — that collector's reads were denied (SQL errors 229/297/300). Check the [permissions](#permissions-on-monitored-servers); the collector retries every cycle and recovers as soon as the grant lands. + +**"Skipping recently-failed-job check"** (info) — the login cannot read `msdb.dbo.sysjobs` / `sysjobhistory`, so failed-job alerts are skipped. Expected for minimal-privilege monitoring logins. If you want job alerts, add the direct msdb table `SELECT`s from the [permissions](#permissions-on-monitored-servers) section — **not** `SQLAgentReaderRole`, which gates the `sp_help_job*` procedures this product never calls and leaves the reads failing with error 229. + +**"TimescaleDB setup failed — continuing in plain-PostgreSQL mode"** (warning) — the extension exists but conversion hit a problem. Everything still works (DELETE-based retention, plain tables); conversion is retried on the next service start. + +**"out of background workers" / "failed to start a background worker" in the postmaster log, or the store keeps growing despite compression** — bring-your-own stores only: the cluster has fewer worker slots than the store has policies, so compression and retention jobs are being skipped. An occasional one is benign (the job retries on its next schedule); persistent ones mean the store is effectively uncompressed. Size the two settings and restart the server — see [Background workers](#background-workers-sizing-an-unmanaged-store-and-what-happens-if-you-dont), and multiply them if the cluster hosts more than one store. `timescaledb_information.job_stats` tells you whether jobs are actually succeeding. + +**"Why are there 40+ postgres.exe processes?"** — the count is three populations, and only one is client connections: (1) PostgreSQL's own system processes (postmaster, checkpointer, WAL/background writers, autovacuum, stats); (2) **TimescaleDB background workers** — the managed conf sizes `timescaledb.max_background_workers` to the hypertable count + 2 (≈57), and every RUNNING compression/retention policy job is its own process, so the count legitimately surges during checkpoint/compression waves and falls back when they finish; (3) client backends — the service's pools are capped at 24, the co-located viewer's at 10. Decompose it live with: `SELECT backend_type, count(*) FROM pg_stat_activity GROUP BY backend_type ORDER BY 2 DESC;` — and remember Windows charges the shared buffer segment to every attached process's working set, so per-process memory numbers cannot be summed. + +**query_store bursts every ~15 minutes** — two or three near-empty cycles, then one large one, is Query Store's own behavior, not a collector bug: the engine buffers in memory and flushes to its persisted tables on `DATA_FLUSH_INTERVAL_SECONDS` (default 900s), so the collector genuinely sees nothing new between flushes. Narrowing the collection interval will not smooth it. The per-database log lines show which database drove a burst. + +**MCP client cannot connect** — MCP defaults to off. Enable it live from the Viewer's Settings (the checkbox writes the control plane; the service starts the endpoint within seconds, no restart), or set `mcp.enabled: true` in `darling.json` for a file-seeded install. If the log says `Port 5152 is already in use — MCP server not started`, change `mcp.port`. The MCP server binds to `localhost` only unless you opt into a LAN endpoint (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan)); a remote client that gets 401 is missing or mismatching the required bearer token, and one that is refused before any response is outside the configured `allowFrom` CIDR. + +**Recommendations tab says no findings** — analysis runs every 30 minutes per server but only once the store holds at least 24 hours of collected data for that server; a fresh install simply has not earned findings yet. + +**The Viewer will not connect** — the failure window carries a **Configuration this viewer used** block naming the `darling.json` it read (and which rule picked it: an explicit command-line path, `DARLING_CONFIG`, beside the viewer, or the service root), plus the host, port, username, database, SSL mode, search path and certificate path it parsed. Read it before changing anything: the two faults it separates are *the viewer read a different file than you edited* and *it read your file and a value in it is wrong*. **Copy details** puts the whole block on the clipboard for a bug report, and the same lines are in `%APPDATA%\PerformanceMonitorDarling\logs\darling-viewer_yyyyMMdd.log`. It never contains a password. + +**"Root Certificate ... exists: NO"** — with `SSL Mode=VerifyFull`, a **relative** `Root Certificate` path resolves against **the folder holding the `darling.json` the viewer read** — not the working directory, so how the viewer was launched no longer changes the answer. The diagnostics block prints that folder and the absolute path it actually opened; either put `server.crt` beside the config or make the `Root Certificate` value an absolute path. If you no longer have the certificate, re-run `--export-viewer-config` on the store host and copy the folder again (see [Connect a Remote Viewer](#connect-a-remote-viewer)) — it regenerates if the bind IP changes, so an old copy stops matching. + +--- + +## How It Runs (Reference) + +Fixed cadences, hardcoded on purpose: + +| What | Cadence | +|---|---| +| Collector sweep loop | Every 15 seconds (each collector runs when its own shared schedule is due — most every 1 minute, some every 5, sizes hourly, index stats daily) | +| Alert evaluation | Every 30 seconds per connected server (Lite's overview cadence) | +| Scheduled analysis | Every 30 minutes per server, 120-second budget, analyzing the last 4 hours; findings persist to `analysis_findings` and high-severity ones notify through the configured channels | +| Retention purge | First sweep after startup, then daily | +| Reconnect attempts | Every 60 seconds while a server is unreachable | + +--- + +## Managed Bundled PostgreSQL + +With `postgres.managed = true` (the sample's default), the service runs its own bundled PostgreSQL 18 + TimescaleDB and a from-zero install needs no database provisioning at all. Windows only, like every DPAPI surface here. + +```json +{ + "postgres": { + "managed": true, + "port": 5641, + "dataDirectory": null + } +} +``` + +**What first run does.** The service looks for `pg-runtime\pgsql\` beside its binary, extracting it from `pg-runtime.zip` when only the zip is present (deleting the extracted directory is therefore always safe — it self-heals). If the data directory has no cluster, it generates a 32-character random password, protects it with DPAPI LocalMachine into `pg-credential.dpapi` beside the data directory (credential first, so a crash mid-initdb never strands a cluster nobody can log into), then runs `initdb` with `scram-sha-256` auth, data checksums, and UTF8/C locale. A marker-guarded block appended to `postgresql.conf` preloads TimescaleDB, sets the port, and restricts listening to `127.0.0.1`; a second versioned block sizes background workers up for the per-hypertable compression jobs, DERIVED from the live hypertable count so it cannot go stale as collectors are added (`timescaledb.max_background_workers = hypertables + 2`, `max_worker_processes = 3 + that + 8` — today 57 and 68 for 55 hypertables; PostgreSQL's default of 8 workers cannot launch them); a third versioned block sizes memory from the host's physical RAM for the up-to-500-servers case (`shared_buffers = min(25% RAM, 1GB)`, `effective_cache_size = 75% RAM`, `maintenance_work_mem = min(max(5% RAM, 1536MB), 25% RAM, 2048MB)`, and a deliberately-modest per-connection `work_mem = clamp(RAM/512, 16MB, 64MB)` — on an 8 GB box that is `shared_buffers 1024MB` / `work_mem 16MB`; the stock 128 MB / 4 MB defaults are fine at small scale but bottleneck at fleet scale). Later blocks re-state single settings that field measurement moved: a fifth caps `shared_buffers` for the co-located store, a sixth turns on the log-rotation ring, and a seventh carries the `maintenance_work_mem` floor that TimescaleDB's compression sort runs on (measured at ~+70% compression throughput on a 16 GB-class host, plateauing by 1536 MB). `postgresql.conf` takes the LAST assignment of a setting, so these override without rewriting anything. Every append is re-checked on every start, so a crash between initdb and the append heals itself instead of silently degrading — and clusters initialized before a given block existed gain it on their next start (effective at the next PostgreSQL restart). Then `pg_ctl start`, `CREATE DATABASE darling`, and the normal startup path (migrations, TimescaleDB adoption — you should see `N/N collector table(s) are hypertables`, both numbers equal and equal to the collector count; a converted count BELOW the total means some table stayed plain and the line above it says which) continues exactly as in bring-your-own mode. The connection string is derived from the stored credential; the Viewer and the MCP host on the same machine derive it the same way, so nothing needs configuring there either. + +**Why scram and not trust, even loopback-only.** Trust auth would hand superuser to any local code that can open a loopback socket — every other local user, and network-capable-but-not-filesystem-capable attack primitives like SSRF from a co-hosted app. With scram the credential travels on the wire, failed attempts are auditable, and access is confined to what can read the DPAPI-protected credential file. `listen_addresses = '127.0.0.1'` keeps the server unreachable off the machine on top — unless you deliberately opt into a LAN endpoint (see [Opt-in Network Endpoints (LAN)](#opt-in-network-endpoints-lan)), which reconciles `listen_addresses`, a `hostssl` pg_hba rule, and TLS on every start and is otherwise off. + +**Lifecycle.** On shutdown the service stops the server (`pg_ctl stop -m fast`) **only when it started it**. A server that was already running — an operator's own `pg_ctl`, or a postmaster that survived a service crash — is adopted for connections but never stopped: you'll see `already running … will not stop it` in the log, and the service keeps collecting into it. + +**The runtime zip.** `pg-runtime.zip` ships beside the service binary in packaged releases. Building from source, produce it once with `Darling\tools\fetch-pg-runtime.ps1` — it downloads the pinned EDB PostgreSQL 18 binaries and TimescaleDB, verifies their SHA256, prunes what the service doesn't need, and writes the zip to `Darling\artifacts\`; copy it next to the built service exe. + +**Server log.** The bundled server's own log is `pg.log` beside the data directory — that's where PostgreSQL explains a refused start; bootstrap errors in the service log quote its tail. + +## Security & Least-Privilege Roles + +The store is split into two schemas so that no consumer connects with more privilege than it needs: + +- **`collect`** — the collector hypertables (one per collector) plus the service-written, user-read metadata (`servers`, `collection_log`, `analysis_findings`, the `v_*` views). Read-only to everyone but the service. +- **`config`** — exactly the tables a human operator changes through the Viewer or MCP: `config_mute_rules`, `config_alert_log` (alert dismissals), `config_edge_trigger_watermarks`, and `analysis_muted`. + +Table names are unchanged — only their schema moved — and the shared SQL keeps using the bare, unqualified names, resolved through `search_path = collect, config, public` (set as the database default and carried on the managed connection strings). This is deliberate: Darling's SQL is byte-identical to Lite's DuckDB SQL, and re-qualifying it would fork that twin. + +**The roles.** The service still owns the store as the `darling` superuser (it does the DDL — migrations, hypertable conversion, retention). On top of that, **managed mode provisions three least-privilege login roles** (BYO provisions two — see below): + +| Role | Privileges | Used by | +|---|---|---| +| `darling` | superuser / owner | the service (collection, migration, provisioning) | +| `admin` | SELECT on both schemas — **including** the secret columns, which the Settings window reads — plus INSERT/UPDATE/DELETE on `config` only. No statement timeout | the Viewer, by default (`connectAs: "admin"`) | +| `viewer` | SELECT on all of `collect`, and on `config` **minus the secret columns** of `config_monitored_servers` / `config_command` / `config_notification` (carved fail-closed, below) + INSERT/UPDATE/DELETE on `config.custom_views` only (the web composer's saved views). Runs under `statement_timeout = 15s` | a locked-down Viewer (`connectAs: "viewer"`), and the web dashboard | +| `mcp` | `viewer`'s exact read surface + INSERT on `collect.analysis_findings` / `config.analysis_muted` + INSERT/UPDATE/DELETE on `config.custom_views` (the custom-view tools) + the alert-tuning writes (INSERT/UPDATE/DELETE on `config.config_mute_rules`, UPDATE on `config.config_alert_settings`, and the `config_service` reload-beacon columns) + the server-onboarding writes (INSERT/UPDATE/DELETE on `config.config_monitored_servers` — the credential column stays SELECT-carved, so it can WRITE a password blob but never READ one back) | the store identity the opt-in MCP **network** endpoint connects as (managed only); dormant until MCP is exposed on the LAN | + +`admin` cannot `DROP`, alter schema, touch `collect` data, or create objects — it can only do what the Viewer's mute-rule / alert-dismiss surfaces need. The `mcp` role is narrower still: it reads exactly what `viewer` reads (the secret config columns are carved out identically) and its writes are a small, enumerated set — the two analysis-table INSERTs (`analyze_server` + `mute_analysis_finding`), the single-table `config.custom_views` CRUD (the custom-view tools), the alert-tuning writes (`config.config_mute_rules` CRUD + a single-row `config.config_alert_settings` UPDATE, plus the two `config_service` beacon columns so a settings write's self-bump trigger can fire), and the server-onboarding writes (`config.config_monitored_servers` CRUD for `add_servers` / `remove_server` — its `config_monitored_servers` write fires the SAME `config_service` beacon trigger, already covered by that column grant) — so a token-holder on the network MCP endpoint can never reach the `config`-table service-credential pivot, the secret columns, or a service flag like `paused`. Even on `config_monitored_servers`, which it may write, the `encrypted_password` column stays in the fail-closed secret carve, so `mcp` can WRITE a credential blob (onboarding) but can never READ one back. `ALTER DEFAULT PRIVILEGES` means new collector tables auto-inherit SELECT for `admin`/`viewer`, so the model never drifts as collectors are added (every `mcp` write is an explicit single-table/single-column grant, deliberately not schema-wide). + +**Managed mode** provisions all of this automatically on every start (idempotent and self-healing), generating a per-role DPAPI-LocalMachine credential — `pg-admin-credential.dpapi`, `pg-viewer-credential.dpapi`, and `pg-mcp-credential.dpapi` beside the data directory, same posture as the owner's `pg-credential.dpapi`. Nothing to configure beyond `connectAs`. + +**Credential file protection.** DPAPI LocalMachine scope is deliberate (the service writes the credential, a *different* interactive user's Viewer reads it), which means the machine-bound blob is decryptable by anything that can *read* the file. So the credential files are locked down with an NTFS ACL that strips the inherited world-read `%ProgramData%` would give them: + +| File(s) | Readable by | +|---|---| +| `pg-credential.dpapi` (superuser), `pg-mcp-credential.dpapi` (the network MCP role) + the transient init pwfile | SYSTEM, Administrators, the service account — **not** interactive users | +| `pg-admin-credential.dpapi`, `pg-viewer-credential.dpapi` | the above **+ `NT AUTHORITY\INTERACTIVE`** (the operator's Viewer) | + +`pg-mcp-credential.dpapi` sits with the superuser (non-interactive) rather than with the Viewer's credentials because only the in-service MCP host reads it — never an interactive Viewer. + +The principal model assumes the **single-operator VM** this edition targets: `INTERACTIVE` == the operator, so the admin/viewer credentials are readable by the Viewer with zero configuration, while non-interactive local code (other services, sandboxed/SSRF socket primitives, scheduled tasks) and the superuser credential are excluded outright. On a shared machine where untrusted users log on interactively, tighten those two files to the specific operator account by hand. The service also refuses to trust a credential file that isn't owned by SYSTEM/Administrators/itself (closing a pre-plant attack), and regenerates an untrusted role credential. + +**A read-only (`viewer`) Viewer degrades gracefully.** It probes its own privileges on connect (`has_table_privilege`), so the mute-rule Add/Edit/Toggle/Delete/Purge buttons and the alert Dismiss / Dismiss All buttons are hidden or disabled, and any write that still slips through returns a clear "read-only connection" message instead of an error. + +**Bring-your-own PostgreSQL.** The schema split runs everywhere (it's a migration — the service applies it on startup and best-effort sets the database `search_path`; if your collection login can't `ALTER DATABASE`, run that one statement yourself as the owner). Role provisioning is managed-only, so for BYO you create the roles yourself, once, with the shipped script: + +``` +psql -h -U -d darling -f Darling/tools/provision-roles.sql +``` + +Edit the two password placeholders (and the database/owner names if yours differ) first. Then point a read-only Viewer's `connectionString` at the `viewer` role. **That script is the authoritative grant list for a BYO store** — it is what actually runs, the table above is its summary, and an `ALTER DEFAULT PRIVILEGES` in it means a store gaining collectors later needs no re-grant. Re-run it after a schema upgrade to cover new tables. **It creates two login roles — `admin` and `viewer`** — the two the Viewer connects as. Managed mode creates a third, `mcp`, but BYO deliberately does not: the MCP **network** endpoint (the only consumer of the `mcp` role) is managed-mode-only, and a BYO operator governs their own PostgreSQL's network exposure. If you expose MCP through your own reverse proxy against a BYO store, point it at whichever least-privilege role you choose (the `viewer` role covers the read tools; `analyze_server`'s finding persistence and `mute_analysis_finding` need INSERT on `collect.analysis_findings` / `config.analysis_muted`). + +## Opt-in Network Endpoints (LAN) + +By default all three network surfaces bind **loopback only** — the store to `127.0.0.1`, the MCP server and the web dashboard to `localhost` — exactly as they always have. Three optional, independent opt-ins let a remote viewer, MCP client, or browser on your **trusted LAN** reach them. This is a home-lab / trusted-subnet feature: **never expose any of these endpoints to the internet.** All three are **managed-mode only** (in bring-your-own mode your own PostgreSQL / reverse proxy governs exposure, and the config is ignored with a warning), and all three are **fail-closed** — any invalid or incomplete field degrades that endpoint back to loopback and logs a critical line rather than exposing it. Removing the config on the next restart closes the box again. + +### Guided setup (`--configure-network`) + +The fastest path is the interactive wizard — run it on the **service host**: + +``` +PerformanceMonitor.Darling.Service.exe --configure-network +``` + +It shows the current exposure (read from the service's own resolvers), then walks you through the **store**, **MCP**, the **web dashboard**, any comma combination (e.g. `1,3`), or all three at once (or a **disable** that removes all exposure). Every answer is validated **by delegation to the exact checks the running service fail-closes on**, so the wizard can never write a config the service would refuse — it re-prompts with the resolver's own reason. It generates the MCP bearer / web access tokens for you (DPAPI-protected; each plaintext is printed once, so save it then), edits `darling.json` **in place preserving every comment** behind a timestamped `darling.json.bak-` backup, prints the scoped firewall command(s), the `--export-viewer-config` handoff, and the web dashboard's browser login URL (`http://:/?token=...`), and offers to restart the service to apply. `install-darling.ps1 -Network` runs it automatically right after the install reaches Running. The manual field reference below documents exactly what it writes. + +### Firewall rules (`--configure-firewall`) + +The service runs as `NT SERVICE\PerformanceMonitor Darling`, an unprivileged virtual account that **cannot create Windows Firewall rules** — and should not be able to. So the rules are managed from the elevated install instead: `install-darling.ps1` runs `--configure-firewall` for you (before the first start, and again after `-Network`), and `uninstall-darling.ps1` removes them. Run it by hand after any edit to a `network` block: + +``` +PerformanceMonitor.Darling.Service.exe --configure-firewall +``` + +Run **elevated**. It reconciles all three scoped rules — store, MCP, web dashboard — against `darling.json` in one pass: it opens the port for every surface that really is exposed and removes the rule for every surface that is not, so it also cleans up after an exposure you turned back off. It is idempotent (safe on every upgrade). + +The **port** each rule is named for comes from the control plane, not from the file. `mcp.port` / `web.port` in `darling.json` are only a first-run seed; `config.config_service.mcp_port` / `web_port` is what the endpoint actually binds, so a rule named from the file after someone moved the port in the Viewer's Settings would open a port nothing is listening on and leave the served port shut. The verb therefore reads those two columns before it plans. That read is **best effort** — the verb still has to run at install time, before the store exists, where `darling.json`'s port is the value the store will be seeded *with* — and it is never silent: when the store cannot be reached it names the port it used, why it could not do better, and the one case in which that port is wrong (a port that has since been changed in the Viewer), so re-running it once the service is up moves the rule. `--enable-mcp` / `--enable-web` have no such window: they take the port back from the store write they already perform, so they cannot open a port on a guessed number at all. + +Because the port is part of the rule's **name**, changing a port does not update a rule — it makes a different one and leaves the old one standing as an inbound allow on a port nothing serves. Every removal here therefore sweeps `PerformanceMonitor Darling (port *)` rather than one exact name, so the stale rule goes with it. The wildcard is derived from a rule name this product built; a name that does not parse degrades to the exact-name removal, so a sweep can never reach a rule that is not ours. + +"Really exposed" is decided by the same resolvers the running service fail-closes on, not by reading `listen` at face value. A `network` block the service would degrade to loopback — an unparseable `listen`, a missing or invalid `allowFrom`, an address family that disagrees, a missing token, BYO mode — gets **no open port**, and the verb tells you why. + +The running service never touches these rules. It **checks** them on start and logs what it finds: nothing at all for the normal loopback-only install, one INFO line when an exposed endpoint's rule is present, and one WARN naming the exact command when an exposed endpoint's rule is missing or when a loopback-only endpoint still has a stale rule open. It states each verdict once, not once per retry. + +### Headless enable/disable + firewall (`--enable-mcp` / `--enable-web`) + +On a box with no Viewer, two things are otherwise awkward: the `enabled` flags in the `mcp` / `web` blocks below are only a **first-run seed** — after the first run the store (`config.config_service.mcp_enabled` / `web_enabled`) is authoritative and is normally flipped only from the Viewer's Settings — and the service account (`NT SERVICE\PerformanceMonitor Darling`) **cannot open the firewall itself**. Four verbs, run on the **service host**, close both in one elevated action: + +``` +PerformanceMonitor.Darling.Service.exe --enable-mcp +PerformanceMonitor.Darling.Service.exe --disable-mcp +PerformanceMonitor.Darling.Service.exe --enable-web +PerformanceMonitor.Darling.Service.exe --disable-web +``` + +Each flips only its endpoint's **live store flag** with a targeted `config_service` write; the service **hot-reloads within one collection sweep — no restart.** If that endpoint's `network` block opts into LAN exposure (a non-loopback `listen`), the verb also reconciles that endpoint's **scoped, idempotent-by-name firewall rule**: **run elevated**, it opens (or, on `--disable-*`, removes) the rule; **run non-elevated**, the store toggle still succeeds and it prints the exact elevated firewall command to run by hand (a loopback-only endpoint needs no rule and says so). Managed-mode only, Windows only. So the headless bring-up is: write the `network` block (the wizard above or the manual reference below), then `--enable-mcp` / `--enable-web` from an **elevated** shell. + +### Verify it's actually reachable (and the two failures that look like bugs) + +Enabling an endpoint is **not** the same as reaching it, and both common failures leave the store flag reading `true`, so "it says enabled" is not proof. After `--enable-mcp` / `--enable-web`, verify on the **service host**: + +1. **The listener is on the LAN address, not loopback.** `Get-NetTCPConnection -State Listen | Where-Object LocalPort -eq 5152` (or `5153` for web) must show the box's LAN IP, e.g. `10.0.0.5:5152` — **not** only `::1` / `127.0.0.1`. *Enabled but still loopback-bound* is the single most common failure: the store flag is on, but the service loaded `darling.json` **before** the `network` block existed. The block is read **once at service start** — the enable toggle stops/starts the endpoint with the already-loaded config and does **not** reload the file. **Restart the service** (`Restart-Service 'PerformanceMonitor Darling'`) so it re-reads the block, then re-check the listener; after the restart run `--configure-firewall` **elevated** if the firewall rule is missing (the service account cannot create it, so the service only tells you it is missing). +2. **The scoped firewall rule exists and covers the client.** `Get-NetFirewallRule -DisplayName 'PerformanceMonitor Darling MCP (port 5152)'` (or `... Web (port 5153)`) should be `Enabled=True, Action=Allow`, scoped to the `network.allowFrom` CIDR. If it is absent, the service's own start-up log already says so and names the command; `--configure-firewall` elevated is the one-step fix. Reading rules needs no elevation, so this check works from any shell. + +Then from the **client** host: + +3. **Connect to the box's LAN IP, never `localhost`.** Use `http://:5152/`. `localhost` / `127.0.0.1` only resolves *on the box itself*, so an off-box MCP client pointed at localhost fails silently — this is the number-one "MCP won't connect" cause. Send `Authorization: Bearer ` (the `network.token`), and do a **fresh** `initialize` + `tools/list` rather than trusting a cached tool list from a previous version. + +**After a reinstall:** the installer replaces binaries but does **not** touch `darling.json` (the zip ships only `darling.sample.json`) or the store, so the `network` block and both live flags survive the upgrade — and the reinstall restarts the service, which re-reads the block. If MCP stops connecting afterward it is almost always failure 3 (the client pointed at `localhost`) or a missing firewall rule, **not** lost config: run `--configure-firewall` **elevated** to re-open the rule if check 2 comes up empty (the installer already does this, so an in-place upgrade normally leaves the rules correct). A stale loopback bind is unlikely after a restart unless the block itself is invalid, in which case the endpoint fail-closes to loopback and logs a critical line saying why — fix the block and restart again. A full `--configure-network` re-run is only needed if the `network` block itself is gone. + +### Store endpoint (viewer over the LAN) + +Add a `network` block to `postgres` (managed mode): + +```json +"postgres": { + "managed": true, + "port": 5641, + "network": { + "listen": "192.168.1.205", + "allowFrom": "192.168.1.0/24", + "role": "viewer" + } +} +``` + +On every start the service reconciles this against the live cluster: it adds the bind IP to `listen_addresses`, generates a self-signed TLS certificate (`server.crt` / `server.key` beside the data directory, with both an IP SAN for `listen` and a DNS SAN for the machine hostname), writes a marked `hostssl darling scram-sha-256` rule into `pg_hba.conf` for each admitted role and reloads, and **checks** (never creates — see [Firewall rules](#firewall-rules---configure-firewall)) that the store's scoped firewall rule matches. + +- **`role`** — the pg_hba login role(s) the rule names: `"viewer"` (default, **read-only** — the secure default, covering a laptop reading every dashboard, chart, and finding), `"admin"` (full remote **writes**; the service logs a warning because `admin` holds the `config_command` / `config_monitored_servers` / `config_notification` service-credential pivot), or **both at once** — `"admin,viewer"`, separated by a comma, a `+`, or whitespace. Both writes one `hostssl` line per role inside the managed block, so an `admin` seat and any number of read-only ones reach the same store. Never the superuser, and never `all`: every generated line still names exactly one role and one CIDR, so narrowing `allowFrom` later tightens all of them together — unlike a second line hand-added outside the markers, which the reconciler preserves verbatim and therefore leaves on the old, wider CIDR. A value naming anything else is rejected **whole**, including a typo in one half of a pair: the store degrades to loopback with the reason logged rather than opening with the half that parsed. Absent or blank still means `viewer` alone. This is **distinct from `postgres.connectAs`** (the *local* VM viewer's loopback role, default `admin`): `network.role` is the *remote* role and defaults to `viewer`, so the two have opposite defaults — the local seat is writable, the remote seat is read-only, unless you say otherwise. +- **TLS is verify-full, not `require`.** Because Darling generates the cert, the client can pin it, so the connection string below uses `SSL Mode=VerifyFull` — which actually defends against an on-path MITM (`require` verifies nothing). The store's network pg_hba line is `hostssl`, so a non-TLS network client is refused. +- **The firewall is defense-in-depth, not the boundary** — pg_hba + TLS are. `--configure-firewall` (elevated) creates the store's scoped rule for you along with the other two; the equivalent by hand is: + + ``` + New-NetFirewallRule -DisplayName "PerformanceMonitor Darling store (port 5641)" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5641 -RemoteAddress 192.168.1.0/24 + ``` + +**That is the service side. The viewer side is [Connect a Remote Viewer](#connect-a-remote-viewer)** — `--export-viewer-config` on this host writes the viewer machine's whole configuration folder (config, certificate, and a plain-text field reference), and that section covers copying it over, the certificate's placement and rotation, and the manual `--print-viewer-connection` fallback. + +### MCP endpoint (assistant over the LAN) + +Add a `network` block to `mcp` (managed mode; `mcp.enabled` must be `true`): + +```json +"mcp": { + "enabled": true, + "port": 5152, + "network": { + "listen": "192.168.1.205", + "allowFrom": "192.168.1.0/24", + "encryptedToken": "" + } +} +``` + +When `listen` is a network address **and** a token is present **and** `allowFrom` is a valid CIDR, the MCP host binds that interface behind two gates: a **required bearer token** (checked first, constant-time, no loopback exemption) and an **in-app CIDR check** on the remote address (loopback is always allowed, so local clients keep working). Any missing precondition keeps MCP loopback-only. Prefer `encryptedToken` (a DPAPI blob from `--encrypt-password`); a plaintext `token` works for dev but is warned. Set the same scoped firewall rule for the MCP port: + +**Lost the token?** `--print-mcp-token` (elevated) reprints it from `darling.json` — `--print-web-token` does the same for the dashboard. Both write the live token to **STDOUT** with every warning on STDERR, so `... --print-mcp-token | clip` captures the value and still shows the warning. This discloses nothing new: the blob is DPAPI at `LocalMachine` scope with an entropy constant published in this repository, and `darling.json` grants `INTERACTIVE` read by design (see [Security & Least-Privilege Roles](#security--least-privilege-roles)), so anyone who can log on to the host interactively could already decrypt it. The elevation requirement makes reprinting a deliberate act; the token's actual protection is the file's ACL. If the token has **leaked** rather than been mislaid, reprinting is not the fix — `--configure-network` generates a new one, and every client configured against the old one must be updated. + +``` +New-NetFirewallRule -DisplayName "Darling MCP" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5152 -RemoteAddress 192.168.1.0/24 +``` + +**What a token-holder can — and cannot — do.** Start with the boundary: **no MCP tool runs SQL an AI client wrote against your monitored servers.** No such tool exists, and a stored custom view cannot become one either — a composed query names only `collect.*` collector tables in the monitoring store. The only live contact with a monitored SQL Server is `analyze_server`'s plan fetch and `add_servers`' one-time connection probe, and both run the product's own fixed, read-only queries under the same least-privilege monitoring login the collectors use — the ceiling on what they can see is the ceiling you granted that login, and it has no write grants to hit. Everything else answers from the monitoring store. + +What the token does gate is the monitor's own configuration and collected data: the entire read surface, `analyze_server`, the Custom Views tools (create / modify / delete the saved dashboards and notebooks in `config.custom_views`), the alert-tuning tools (`update_alert_settings` / `create_mute_rule` / `delete_mute_rule`), and the server-onboarding tools (`add_servers` / `remove_server`), which edit the monitored-server registry in `config.config_monitored_servers` — including storing a SQL-auth credential for a server they add. The store-side identity is still the least-privilege `mcp` role: read, the two analysis-table INSERTs, INSERT/UPDATE/DELETE on the single `config.custom_views` table (the same narrow write the web composer's `viewer` role has), the narrow alert-config writes (`config.config_mute_rules` CRUD + a single-row `config.config_alert_settings` UPDATE, plus the `config_service` reload-beacon columns), and the single-table `config.config_monitored_servers` CRUD. So a token-holder can read everything collected, trigger analysis, author custom views, tune alerting, and onboard/offboard servers — and can never reach the `config_command` service-credential pivot, the carved secret columns (SMTP/webhook credentials, and the monitored-server `encrypted_password` blob it can WRITE during onboarding but never READ back, all included), or a service flag like `paused`. Custom-view JSON and alert config carry no secrets. Guard the token like the keys to your monitoring configuration — that is what it opens; your SQL Servers are not behind it. + +**`add_servers` carries a credential in its request.** A SQL-auth `password` rides the request JSON; the service DPAPI-encrypts it at rest and never returns it, but on the wire it is only as protected as the endpoint — the same plaintext HTTP the token rides. On a segment you do not fully trust, front the MCP port with the TLS reverse proxy below, and prefer Windows/integrated auth for onboarded servers where you can — then no per-server secret crosses the wire at all. + +**MCP has no TLS — the MITM control is a TLS reverse proxy.** A self-signed cert breaks real MCP clients, so the MCP endpoint is plain HTTP and the bearer token travels **cleartext on the segment**; an active on-path attacker (ARP spoof, rogue DHCP, compromised switch) could capture and replay it. The in-app CIDR bounds *who can route to* the port; it does **not** protect the wire. If your segment is not fully trusted, put a **TLS-terminating reverse proxy** in front of the MCP port and point clients at that — the named MITM control for this endpoint. (The store endpoint needs no such proxy: it has verify-full TLS built in.) + +**Output format (GCF).** By default the MCP tools return JSON. Setting `DARLING_OUTPUT_FORMAT=gcf` makes the server return [Graph Compact Format](https://gcformat.com) instead: the repeated field names of the record arrays these tools return (blocking pairs, wait stats, index bloat, config rows, …) are factored into a single header and the indentation dropped. The saving is payload-shaped: about a quarter across a mixed real workload, more than half on the most uniform tools, and less on text-dominated ones (multi-line query text, plan or deadlock XML) where the field names are a smaller share of the bytes. In the wire a null field is written as `-` and an absent field (a key some rows omit) as `~`. It is opt-in and conservative: applied per result, and only when the GCF wire is both smaller than the JSON and decodes back to it exactly; a result carrying a number GCF cannot represent exactly (a non-integer beyond double precision) stays JSON, as does any result where the wire would be larger, so no result is ever grown, dropped, or garbled. `BlackwellSystems.Gcf` is a zero-dependency package. + +### Web endpoint (browser over the LAN) + +Add a `network` block to `web` (managed mode; `web.enabled` must be `true`): + +```json +"web": { + "enabled": true, + "port": 5153, + "network": { + "listen": "192.168.1.205", + "allowFrom": "192.168.1.0/24", + "encryptedToken": "" + } +} +``` + +When `listen` is a network address **and** a token is present **and** `allowFrom` is a valid CIDR, the web host binds that interface behind two gates: an **in-app CIDR check** on the remote address and an **access token**. A browser presents the token ONCE via `?token=` (open `http://192.168.1.205:5153/`, then paste it into the minimal login form, or append `?token=...` directly); the host validates it constant-time, sets an **HMAC-signed, HttpOnly, SameSite=Strict session cookie**, and 302-redirects to strip the token from the URL so it never lingers in history or a Referer header. Subsequent requests ride the cookie. **Loopback is exempt from the CIDR check only** — while the dashboard is exposed, a request from the box itself still has to present the token or a valid cookie. It used to pass tokenless on the grounds that the surface was read-only; the Custom Views composer's create / update / delete and `/api/compose/run` ended that, so in exposed mode the token *is* the loopback guard against SSRF and sandboxed local sockets ([#1649](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1649)). A dashboard with **no** `network` block registers no auth middleware at all and stays genuinely tokenless. An out-of-CIDR request is refused with **403**, cookie or no cookie; an in-CIDR request with no credential gets the login form back as a **200**, not a 401. The cookie signing key is per-process, so a service restart invalidates open sessions (just re-present the token). Prefer `encryptedToken` (a DPAPI blob from `--encrypt-password`); a plaintext `token` works for dev but is warned. Set the same scoped firewall rule for the web port: + +``` +New-NetFirewallRule -DisplayName "Darling Web" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 5153 -RemoteAddress 192.168.1.0/24 +``` + +**What a web token can reach.** The web dashboard is **read-only over the collected store** — it connects as the least-privilege `viewer` role and hosts no live-server queries (no `analyze_server`, no plan re-execution). Its one write path is the Custom Views composer: INSERT/UPDATE/DELETE on `config.custom_views`, and the web host maps no other write route. (The `viewer` role holds one further narrow grant — INSERT/UPDATE/DELETE on `config.database_state_expected`, for the Viewer's per-database state-override editor ([#1986](https://github.com/erikdarlingdata/PerformanceMonitor/issues/1986)) — which no web route reaches, and which the WPF editor itself gates on the `has_table_privilege` read-only probe a `viewer`-role seat fails.) So a token-holder can read everything collected and author saved views, and reach no monitored server — but that write is why an exposed dashboard authenticates loopback too, exactly as MCP does. + +**The web dashboard can serve HTTPS itself ([#2562](https://github.com/erikdarlingdata/PerformanceMonitor/issues/2562)).** Without it the token and the session cookie it mints travel cleartext on the segment, and the in-app CIDR bounds *who can route to* the port without protecting the wire — so an exposed dashboard with no `tls` block warns about exactly that at every start. Add a certificate to `web.network`: + +```jsonc +"network": { + "listen": "192.168.1.205", + "allowFrom": "192.168.1.0/24", + "encryptedToken": "", + "tls": { + "pfxPath": "C:\\ProgramData\\PerformanceMonitorDarling\\certs\\dashboard.pfx", + "encryptedPfxPassword": "" + // ...or a PEM pair instead: "certPath" + "keyPath" (what the compose distribution mounts) + } +} +``` + +Give **one** form, never both — a PKCS#12 bundle or a PEM pair — and the password may sit in `encryptedPfxPassword` (DPAPI), in `pfxPassword` as a `file:`/`env:` reference, or nowhere at all if the bundle is unprotected. TLS applies to the **network listener only**: the loopback listeners stay plain HTTP, because your certificate names the LAN address and not `localhost`, and loopback traffic never reaches the segment this protects. There is no redirect-from-HTTP, because one port cannot speak both schemes and adding a second HTTP port would reopen the cleartext surface; a plain-HTTP client simply fails the handshake. + +The product **consumes** a certificate and does not manage a PKI — no ACME, and deliberately no self-signed fallback, which would buy encryption without authentication and train you to click through the warning. An internal CA is the normal answer on a LAN. Loading **fails closed**: a certificate that is missing, unreadable, ambiguous, or **expired** keeps the dashboard loopback-only with a Critical log line rather than quietly serving the LAN over HTTP — and because an expired certificate takes the dashboard down, the service warns for the last 30 days before that happens. A TLS-terminating reverse proxy in front of the port remains a perfectly good alternative, and is still the only control for MCP, which has no TLS of its own. diff --git a/Darling/compose/darling.sample.json b/Darling/compose/darling.sample.json index 37b9fca8b..5e2d74346 100644 --- a/Darling/compose/darling.sample.json +++ b/Darling/compose/darling.sample.json @@ -18,6 +18,7 @@ "enabled": true, "port": 5153, "network": { + "// tls": "Without this the token above and the session cookie it mints cross the network in the clear, and allowFrom 0.0.0.0/0 bounds nothing. To enable: mount a PEM pair as the web_tls_cert / web_tls_key secrets (see docker-compose.yml) and add \"tls\": { \"certPath\": \"/run/secrets/web_tls_cert\", \"keyPath\": \"/run/secrets/web_tls_key\" } here. A bad or expired certificate keeps the dashboard loopback-only rather than falling back to HTTP, which in a container means unreachable — that is the intended failure.", "listen": "0.0.0.0", "allowFrom": "0.0.0.0/0", "token": "file:/run/secrets/web_token" diff --git a/Darling/compose/docker-compose.yml b/Darling/compose/docker-compose.yml index e6d0e73bb..d930d69a2 100644 --- a/Darling/compose/docker-compose.yml +++ b/Darling/compose/docker-compose.yml @@ -53,8 +53,13 @@ services: - sql_password - web_token - mcp_token + # Uncomment with the `secrets:` entries at the bottom to serve the dashboard over HTTPS (#2562), + # then point web.network.tls at these paths in darling.json. Without them the web token and its + # session cookie cross the network in the clear. + # - web_tls_cert + # - web_tls_key ports: - - "5153:5153" # web dashboard (token → cookie gate) + - "5153:5153" # web dashboard (token → cookie gate; plain HTTP unless web.network.tls is set) - "5152:5152" # MCP (bearer token) volumes: @@ -71,3 +76,8 @@ secrets: file: ./secrets/web_token.txt mcp_token: file: ./secrets/mcp_token.txt + # HTTPS for the dashboard (#2562) — uncomment together with the `secrets:` list above. + # web_tls_cert: + # file: ./secrets/web_tls_cert.pem + # web_tls_key: + # file: ./secrets/web_tls_key.pem diff --git a/Darling/compose/secrets/README.md b/Darling/compose/secrets/README.md index 09237545a..72fea2a62 100644 --- a/Darling/compose/secrets/README.md +++ b/Darling/compose/secrets/README.md @@ -4,5 +4,6 @@ One secret per file, no trailing content beyond the value (a trailing newline is - `store_password.txt` — the same store password on its own, for the TimescaleDB container's `POSTGRES_PASSWORD_FILE` - `sql_password.txt` — the monitoring login's SQL Server password - `web_token.txt` / `mcp_token.txt` — the dashboard/MCP access tokens (generate long random values) +- `web_tls_cert.pem` / `web_tls_key.pem` — optional, and the only thing that keeps `web_token.txt` off the wire ([#2562](https://github.com/erikdarlingdata/PerformanceMonitor/issues/2562)): the dashboard's PEM certificate and its PKCS#8 private key. Both files and the matching `web.network.tls` block are needed; the certificate must name the address browsers use, and the service refuses to expose the dashboard at all rather than fall back to HTTP if it is missing or expired Keep this directory out of version control and readable only by the deploying user (`chmod 700 secrets`, `chmod 600 secrets/*`). diff --git a/Darling/tools/fetch-pg-runtime.ps1 b/Darling/tools/fetch-pg-runtime.ps1 index a177eb8bf..0d2b5f003 100644 --- a/Darling/tools/fetch-pg-runtime.ps1 +++ b/Darling/tools/fetch-pg-runtime.ps1 @@ -51,7 +51,7 @@ #> # pwsh 7+ only, and not for style: under Windows PowerShell 5.1 this script runs on .NET -# Framework, whose ZipFile.CreateFromDirectory writes BACKSLASH entry separators — a +# Framework, whose ZipFile.CreateFromDirectory writes BACKSLASH entry separators - a # non-conformant zip that Info-ZIP/unzip mangles. CI runs this step under `shell: pwsh`; # requiring 7 here makes a local build byte-behave like the shipped one. #Requires -Version 7.0 diff --git a/Darling/tools/install-darling.ps1 b/Darling/tools/install-darling.ps1 index 805b35dc9..f87e77dfb 100644 --- a/Darling/tools/install-darling.ps1 +++ b/Darling/tools/install-darling.ps1 @@ -5,22 +5,26 @@ and creates Desktop + Start Menu shortcuts for the bundled viewer. .DESCRIPTION Run from an ELEVATED PowerShell, from the folder you extracted the Darling zip into (the script -installs the service pointing at THAT folder — extract to the final location first, e.g. +installs the service pointing at THAT folder - extract to the final location first, e.g. C:\PerformanceMonitorDarling). What it does, in order: 1. Verifies elevation, the service exe, and darling.json (offers to copy darling.sample.json - and stops so you can edit it — the service is not installed with an unedited sample). + and stops so you can edit it - the service is not installed with an unedited sample). 1b. REFUSES an install directory the service account can never read: anywhere under a user profile (C:\Users\...), or a UNC / mapped-drive path. The service runs as an unprivileged virtual account that is neither you nor an administrator, and a profile folder grants nothing to it, so the service installs cleanly and then dies at the bundled PostgreSQL's first step (#2185, #2187). Extract to a machine-scoped local path such as C:\PerformanceMonitorDarling instead. + 1c. REFUSES an install when the ASP.NET Core Runtime 10 is missing, and WARNS when the .NET Desktop + Runtime 10 is (#2479). Both shipped binaries are framework-dependent; a stock Windows Server image + has neither runtime, and the failure is the .NET host's own "You must install .NET" error with + nothing of ours on it. Part of the pre-flight, so -SkipPreflight skips it. 2. Optional pre-flight: runs `--test-connection` and shows the per-server PASS/FAIL lines (continue-or-abort prompt on failure; -SkipPreflight to skip). - 3. Registers the Windows Event Log source 'PerformanceMonitor Darling' (requires elevation — + 3. Registers the Windows Event Log source 'PerformanceMonitor Darling' (requires elevation - the service's own virtual account cannot; without it Event Log diagnostics are silently dropped. The file log under %ProgramData%\PerformanceMonitorDarling\logs works regardless). - 4. Creates the service under the NT SERVICE virtual account (NEVER LocalSystem — the bundled + 4. Creates the service under the NT SERVICE virtual account (NEVER LocalSystem - the bundled PostgreSQL refuses to run with administrative privileges), start=auto. If the service already exists this is an UPGRADE: it is stopped and its binPath updated in place; your darling.json, store data, and credentials are untouched. @@ -33,12 +37,13 @@ C:\PerformanceMonitorDarling). What it does, in order: 5. Starts the service and confirms it reaches Running. 6. Creates 'Darling Viewer' shortcuts on the Desktop and in the Start Menu pointing at viewer\PerformanceMonitor.Darling.Viewer.exe. (Taskbar pinning is deliberately not - attempted — Windows blocks programmatic pinning by design; pin from the Start Menu entry.) + attempted - Windows blocks programmatic pinning by design; pin from the Start Menu entry.) Uninstall with uninstall-darling.ps1 (same folder). .PARAMETER SkipPreflight -Skip the --test-connection pre-flight gate. +Skip BOTH pre-flight gates: the .NET runtime check (1c) and the --test-connection probe (2). The runtime +check is the escape hatch for a private or xcopy runtime layout this script cannot see. .PARAMETER NoShortcuts Do not create the viewer shortcuts. @@ -64,6 +69,12 @@ $viewerExe = Join-Path $root 'viewer\PerformanceMonitor.Darling.Viewer.exe' $configPath = Join-Path $root 'darling.json' $samplePath = Join-Path $root 'darling.sample.json' +# The .NET major both shipped binaries are built against, and the one page that offers every installer +# for it. Kept beside the other script-scope facts so a framework bump is one edit, and deliberately the +# SAME url the UAT onboarding prints - a tester who reads both should not be sent to two places. +$dotnetMajor = 10 +$dotnetDownloadUrl = 'https://dotnet.microsoft.com/download/dotnet/10.0' + function Fail([string]$message) { Write-Host "ERROR: $message" -ForegroundColor Red; exit 1 } # True when $candidate IS $parent or sits underneath it. @@ -166,6 +177,96 @@ function Get-NetworkPathKind([string]$path) { return $null } +# Every place the .NET host looks for a shared framework, in the order it looks. Used by the runtime gate +# at 1c, which has to answer "can these binaries start at all" on a box where dotnet.exe may not exist. +function Get-DotnetRootCandidates { + $roots = New-Object 'System.Collections.Generic.List[string]' + + # DOTNET_ROOT wins here because it wins for the host: an operator who redirected the runtime is not + # someone whose install should be refused for looking in the wrong place. + foreach ($fromEnv in @($env:DOTNET_ROOT, ${env:DOTNET_ROOT_X64})) { + if (-not [string]::IsNullOrWhiteSpace($fromEnv)) { $roots.Add($fromEnv) } + } + + # Where the official installers record themselves. The key is HKLM:\SOFTWARE\dotnet - NOT under + # Microsoft\ - and reading it is what finds an install that was pointed somewhere other than + # Program Files. Its absence proves nothing, so the default location below is checked regardless. + try { + $recorded = (Get-ItemProperty -Path 'HKLM:\SOFTWARE\dotnet\Setup\InstalledVersions\x64' -Name 'InstallLocation' -ErrorAction Stop).InstallLocation + if (-not [string]::IsNullOrWhiteSpace($recorded)) { $roots.Add($recorded) } + } + catch { + # No registry entry, or an unreadable one. Fall through. + } + + foreach ($programFiles in @($env:ProgramW6432, $env:ProgramFiles)) { + if (-not [string]::IsNullOrWhiteSpace($programFiles)) { $roots.Add((Join-Path $programFiles 'dotnet')) } + } + + return $roots +} + +# The version folder names present for one shared framework, e.g. '10.0.11' for Microsoft.AspNetCore.App. +# Returns an empty list when the framework is not installed anywhere this can see. +function Get-InstalledFrameworkVersions([string]$frameworkName) { + $versions = New-Object 'System.Collections.Generic.List[string]' + + foreach ($dotnetRoot in (Get-DotnetRootCandidates)) { + $frameworkDir = Join-Path (Join-Path $dotnetRoot 'shared') $frameworkName + if (-not (Test-Path -LiteralPath $frameworkDir)) { continue } + foreach ($child in @(Get-ChildItem -LiteralPath $frameworkDir -Directory -ErrorAction SilentlyContinue)) { + $versions.Add($child.Name) + } + } + + # `dotnet --list-runtimes` is the authoritative answer and sees layouts the three guesses above do + # not, but it is a SUPPLEMENT rather than the primary check: a box with neither runtime installed + # frequently has no dotnet.exe on it at all, and that is precisely the box this gate exists for. + $muxer = Get-Command 'dotnet' -CommandType Application -ErrorAction SilentlyContinue + if ($muxer) { + try { + foreach ($line in @(& $muxer.Source --list-runtimes 2>$null)) { + if ($line -match ('^' + [Regex]::Escape($frameworkName) + '\s+(\S+)\s')) { $versions.Add($Matches[1]) } + } + } + catch { + # A dotnet.exe that cannot list its own runtimes tells us nothing the directory scan did not. + } + } + + return $versions +} + +# True when at least one of $versions is a build of major version $major. +# +# The test is on the PARSED leading integer, never a text prefix. '1.10.0' and '110.0.0' both contain the +# characters '10.' and neither one is .NET 10, so a -like '10.*' or a bare -match '10\.' waves a box +# through with no runtime on it and hands the operator back the raw host error this gate exists to +# replace - the worse of the two failures, because it looks like the check ran. +# +# MAJOR is the right granularity because it is the granularity the host rolls forward at: a net10.0 app +# rolls forward across patch and minor by default but never across major, so any 10.x satisfies these +# binaries and an 11.x does not. +function Test-FrameworkMajorPresent([string[]]$versions, [int]$major) { + if ($null -eq $versions) { return $false } + + foreach ($version in $versions) { + if ([string]::IsNullOrWhiteSpace($version)) { continue } + if ($version -match '^\s*(\d+)(?:\.|$)') { + if ([int]$Matches[1] -eq $major) { return $true } + } + } + + return $false +} + +# What the gate says it found. 'nothing' rather than an empty string, because a blank there reads as a +# formatting bug and sends the operator looking in the wrong place. +function Format-FrameworkVersionList([string[]]$versions) { + if ($null -eq $versions -or $versions.Count -eq 0) { return 'nothing' } + return (($versions | Sort-Object -Unique) -join ', ') +} + # -- 1. Environment checks ------------------------------------------------------------------------ $identity = [Security.Principal.WindowsPrincipal][Security.Principal.WindowsIdentity]::GetCurrent() if (-not $identity.IsInRole([Security.Principal.WindowsBuiltInRole]::Administrator)) { @@ -206,6 +307,12 @@ if (-not (Test-Path $serviceExe)) { # This runs BEFORE the pre-flight, the Event Log source, and service creation, so a doomed location costs # nothing and leaves nothing behind. $existing = Get-Service -Name $serviceName -ErrorAction SilentlyContinue + +# The same fact as a boolean, for the post-start firewall reconcile at 5a (#2436). Captured here with +# $existing because it stops being observable the moment sc.exe create runs, and named rather than reusing +# $existing at the far end of the script because what 5a actually depends on is not "a service object was +# found" but "a STORE already exists to be asked" - which is what an existing service implies. +$isUpgrade = $null -ne $existing # Classify the NORMALIZED spelling, but keep installing to $root exactly as given (#2348). The \\?\ prefix # instructs the path parser and is not part of where the install lives, so stripping it for the decision # changes which rules see the path and nothing about where files land. @@ -283,6 +390,74 @@ Nothing was installed or changed. if ($answer -notmatch '^[Yy]') { exit 4 } } +# -- 1c. Refuse an install the .NET runtimes on this box cannot run (#2479) ------------------------ +# The first thing a new operator hits and the least self-explanatory. Both shipped binaries are +# FRAMEWORK-DEPENDENT publishes, so both name a shared framework in their runtimeconfig.json and neither +# carries one: +# +# PerformanceMonitor.Darling.Service.exe -> Microsoft.NETCore.App + Microsoft.AspNetCore.App +# viewer\PerformanceMonitor.Darling.Viewer.exe -> Microsoft.NETCore.App + Microsoft.WindowsDesktop.App +# +# ASP.NET Core is required UNCONDITIONALLY, which is the part nobody expects: the MCP tools reference +# ModelContextProtocol.AspNetCore, which brings the Microsoft.AspNetCore.App framework reference in +# transitively, so the framework is named in the runtimeconfig whether or not mcp.enabled and +# web.enabled are ever turned on. Setting them false does not make the requirement go away. +# +# A stock Windows Server image has neither runtime. Without this gate the operator's first signal is the +# .NET host's own error - "You must install .NET to run this application" - printed BY THE PRE-FLIGHT AT +# STEP 2, which then reports "the config is invalid or a server is unreachable" and offers to install +# anyway. That diagnosis is not merely unhelpful, it is wrong, and it points at the config file. So this +# has to run before step 2 touches the exe at all. +# +# Governed by -SkipPreflight rather than a switch of its own, because it is the same kind of gate: a +# question asked before anything is created, answered from the operator's machine, and worth bypassing +# only when the operator knows something this script cannot see. +if (-not $SkipPreflight) { + $aspNetVersions = @(Get-InstalledFrameworkVersions 'Microsoft.AspNetCore.App') + $desktopVersions = @(Get-InstalledFrameworkVersions 'Microsoft.WindowsDesktop.App') + + if (-not (Test-FrameworkMajorPresent $aspNetVersions $dotnetMajor)) { + Fail @" +The ASP.NET Core Runtime $dotnetMajor.0 is not installed, and the service cannot start without it. + + Need: ASP.NET Core Runtime $dotnetMajor.0 (x64). The Hosting Bundle contains it and also works. + Found: $(Format-FrameworkVersionList $aspNetVersions) + Get: $dotnetDownloadUrl + +This is required whether or not you ever enable MCP or the web dashboard - the MCP package brings the +ASP.NET Core framework reference in transitively, so the service names it at startup either way. + +Install it and run this script again. Nothing was installed or changed. + +If this machine DOES have it in a layout this check cannot see, re-run with -SkipPreflight, which skips +this gate and the --test-connection probe together. +"@ + } + + # A WARNING and not a refusal, and the asymmetry is deliberate. A missing ASP.NET Core runtime means + # the thing being installed cannot run; a missing Desktop runtime means only that the viewer cannot, + # and a headless collector host that nobody ever opens a window on is a legitimate deployment. What + # it must not do is stay quiet, because this script creates Desktop and Start Menu shortcuts a few + # steps from here and a tester will double-click one. + # + # Scoped to the co-located viewer\ folder on purpose: the remote-seat viewer installed from the + # Velopack Setup.exe is a SELF-CONTAINED publish and needs no Desktop runtime at all, so this is a + # statement about this box, not about every seat. + if (-not (Test-FrameworkMajorPresent $desktopVersions $dotnetMajor)) { + Write-Host '' + Write-Host "WARNING: the .NET Desktop Runtime $dotnetMajor.0 is not installed." -ForegroundColor Yellow + Write-Host " Found: $(Format-FrameworkVersionList $desktopVersions)" -ForegroundColor Yellow + Write-Host " Get: $dotnetDownloadUrl" -ForegroundColor Yellow + Write-Host '' + Write-Host 'The SERVICE does not need it and this install continues. The bundled viewer does: the Desktop and' -ForegroundColor Yellow + Write-Host 'Start Menu shortcuts this script creates will fail with the .NET host error until it is installed.' -ForegroundColor Yellow + Write-Host '' + } + else { + Write-Host "Runtime check passed (ASP.NET Core $dotnetMajor and .NET Desktop $dotnetMajor present)." -ForegroundColor Green + } +} + if (-not (Test-Path $configPath)) { if (Test-Path $samplePath) { Copy-Item $samplePath $configPath @@ -516,6 +691,12 @@ if ($failed.Count -gt 0) { # re-running it is a no-op, and it needs only darling.json - no store, no credentials - so it is safe here, # BEFORE the first start. # +# On a FRESH install that is not merely safe, it is the only correct time: config_service does not exist yet +# and is seeded FROM darling.json, so the file is the control plane's future answer and cannot be wrong. +# On an UPGRADE it can be: the store already holds a port that may have been moved in the Viewer, and the +# managed store is a child of the service, so with the service stopped there is nothing here to ask. That is +# what 5a below exists for - see the reasoning there. +# # Best-effort: a firewall failure warns rather than aborting an otherwise good install. Re-run it by hand. function Invoke-FirewallReconcile { & $serviceExe --configure-firewall @@ -536,6 +717,30 @@ Start-Service -Name $serviceName Write-Host "Service is Running. First start does real work (unpack pg-runtime, initdb, store migration, first collection cycle) - give it ~2 minutes." -ForegroundColor Green Write-Host "Primary log: %ProgramData%\PerformanceMonitorDarling\logs\darling-service_yyyyMMdd.log" +# -- 5a. Re-reconcile the firewall on an UPGRADE (#2436) ------------------------------------------- +# The 4c call ran with the service stopped, which on a managed install means the store was down too - it is +# a child of the service. So on an upgrade the verb fell back to darling.json's mcp.port / web.port, and on +# a box whose port was later moved in the Viewer's Settings that is the WRONG port: the rule lands where +# nothing listens while the served port stays shut. The verb says so when it falls back, and the running +# service's own start-up check then WARNs the exact command - so this always did self-heal with one operator +# action. This removes the need for that action in the case where it was never necessary, because the store +# is up now and can simply be asked. +# +# Only on an upgrade, and that is the point rather than an optimisation. On a fresh install the first start +# is doing initdb and the first migration for ~2 minutes, so a second call here could not read the store +# either - it would re-print the same fallback disclosure more loudly about a port that, on a fresh box, is +# by definition correct. Skipping it by construction beats skipping it by hoping the timing works out. +# +# Not a guarantee: the store answers when it answers, and the verb bounds its read to ten seconds. If it +# still cannot, this run says exactly what the 4c run said and the operator has the same remedy they had +# before - so the window narrows rather than closing, which is the honest claim. The verb is idempotent, so +# a redundant run costs nothing but output. +if ($isUpgrade) { + Write-Host '' + Write-Host 'Re-applying the firewall rules now that the store can answer (upgrade path)...' -ForegroundColor Cyan + Invoke-FirewallReconcile +} + # -- 5b. Optional guided network setup ------------------------------------------------------------ # Runs elevated (this whole script is), so the wizard's restart-to-apply works, and its restart is what # generates the store TLS cert on the first exposed start. Loopback-only stays the default without -Network. diff --git a/Darling/tools/upgrade-darling.ps1 b/Darling/tools/upgrade-darling.ps1 new file mode 100644 index 000000000..5a082acdb --- /dev/null +++ b/Darling/tools/upgrade-darling.ps1 @@ -0,0 +1,1469 @@ +<# +.SYNOPSIS +Upgrades an installed PerformanceMonitor Darling in place from a newer build, keeping a BOUNDED number of +rollback backups - or prunes the ones an earlier deploy left behind (-PruneOnly). + +.DESCRIPTION +This is the supported version of a procedure that had been living in people's heads and in ad-hoc SSM +scripts: stop the service, copy the install tree's root files aside as _rollback_manual_, lay the new +build over the top, start the service, verify. Every step of it was already documented somewhere. What was +missing was the step nobody remembered to do by hand, which is deleting the backups from the LAST twenty +deploys - a dogfood box was found carrying 46 of them, 5.48 GB, the oldest three weeks old, and the service +warning about every single one on every single start (#2525). + +Retention lives HERE, at deploy time, and not in the service. The service does not delete things it did not +create; this script created every one of these directories, so this script is the only thing entitled to +remove them. It keeps the newest -KeepRollbacks (3 by default) and removes the rest, after the new backup is +made - so a failed upgrade always has something to roll back to, including the copy it just took. + +What it does, in order: + + 1. Verifies elevation, resolves the install root from the REGISTERED service (or -InstallRoot), and + refuses if the service is not installed - this upgrades an install, it does not create one. Use + install-darling.ps1 for that. + 2. Refuses to copy a source that IS the install root. The upgrade would be overwriting this very script + while PowerShell is reading it, and Expand-Archive dying half-way through leaves the service stopped + with mixed binaries. Run the copy that came out of the NEW zip instead. + 3. Verifies the source zip's SHA256 (against -Sha256, or a SHA256SUMS.txt sitting beside it). Refuses an + unverified zip unless you say -SkipHashCheck out loud. + 4. Names any process running out of the install tree and STOPS - it never kills one. The bundled + PostgreSQL lives under pg-runtime and a blanket kill takes the store down with it; the last time this + guard fired, what it caught was an operator's own psql.exe sitting in the install directory from a + diagnostic query. + 5. Stops the service and waits for it to actually be Stopped. + 6. Copies the install root's FILES (not its subdirectories) to _rollback_manual_. A re-run inside + -BackupWindowMinutes reuses the existing backup instead of taking a second one, because a second one + would be a copy of the half-extracted tree written over your only good copy. + 7. Prunes the rollback backups past -KeepRollbacks, each in its own handler. + 8. Lays the new build over the install root, retrying once on the transient file lock that has bitten this + step before. + 9. Confirms darling.json is byte-identical to what it was. + 10. Names the files an EARLIER build shipped and this one does not, and removes them with + -RemoveStaleFiles. An in-place upgrade is an OVERLAY: it writes what the new build ships and deletes + nothing else, so a dropped dependency or a stranded runtimes\\lib\\ subtree stays in the + tree forever - in a directory .NET probes for assemblies. The authority is a manifest this script + writes into the install root after every copy (#2529). + 11. Starts the service, waits for Running, and prints the post-install health check - which is a STAGE of + the install, not a favour. + +Every step is safe to re-run. That is not a nicety: steps 5 through 9 leave the service DOWN if anything +between them fails, so "run it again" has to be the correct advice, and the script says so at the point of +failure rather than leaving you to guess. + +.PARAMETER Source +The new build: either the .zip as downloaded, or a folder you already extracted it into. Defaults to the +folder this script is in, which is what you get by extracting the new zip to a staging directory and running +ITS copy of this script. + +.PARAMETER InstallRoot +The install directory to upgrade. Defaults to the directory of the registered service's executable, which is +the one place that cannot be wrong about where the service is actually installed. + +.PARAMETER Sha256 +Expected SHA256 of the source zip. Without it the script looks for SHA256SUMS.txt beside the zip. + +.PARAMETER KeepRollbacks +How many rollback backups to keep, newest first, INCLUDING the one this run takes. Three is enough to roll +back a bad deploy; the fourth can only roll back to a version nobody wants. + +.PARAMETER BackupWindowMinutes +A backup newer than this is reused rather than replaced, so a re-run after an interrupted upgrade does not +overwrite the good pre-upgrade copy with a copy of the half-upgraded tree. + +.PARAMETER PruneOnly +Prune the rollback backups and exit. Nothing is stopped, nothing is copied, and the service keeps running. +This is what an existing box with a backlog needs, and it is the command the service's own layout report +tells operators to run. + +.PARAMETER ListRollbacks +Show which backups would be kept and which pruned, and exit. Changes nothing. + +.PARAMETER RemoveStaleFiles +Delete the files an earlier build shipped and this one does not, rather than only naming them. + +The check itself runs on every upgrade and reports either way; this switch is the difference between a +report and a delete. It is OFF for the first release on purpose. Everything this can name provably came out +of one of our own build payloads - the manifest is written from the payload, so darling.json, the DPAPI +credential blobs, the rollback backups and pg-runtime were never in it and cannot come out of it - but a +delete that runs inside a monitoring host's install directory should spend a few deploys showing operators +its answer before it starts acting on it. Flip the default once boxes have been reporting it and the lists +have been the ones people expected. + +.PARAMETER SkipHashCheck +Proceed with a source zip whose SHA256 could not be verified. + +.PARAMETER SkipStopGuard +Proceed even though processes are running out of the install tree. Almost always the wrong answer - the +copy will fail on a locked file and leave mixed binaries - but there is no way to be sure from here that +your case is not the exception. +#> +[CmdletBinding()] +param( + [string]$Source = $PSScriptRoot, + [string]$InstallRoot, + [string]$Sha256, + [ValidateRange(1, 100)] + [int]$KeepRollbacks = 3, + [ValidateRange(0, 10080)] + [int]$BackupWindowMinutes = 60, + [switch]$PruneOnly, + [switch]$ListRollbacks, + [switch]$RemoveStaleFiles, + [switch]$SkipHashCheck, + [switch]$SkipStopGuard +) + +$ErrorActionPreference = 'Stop' +$serviceName = 'PerformanceMonitor Darling' +$serviceExeName = 'PerformanceMonitor.Darling.Service.exe' +$configName = 'darling.json' +$manifestName = 'darling-install-manifest.txt' + +function Fail([string]$message) { Write-Host "ERROR: $message" -ForegroundColor Red; exit 1 } +function Note([string]$message) { Write-Host $message } +function Good([string]$message) { Write-Host $message -ForegroundColor Green } +function Warn([string]$message) { Write-Host "WARNING: $message" -ForegroundColor Yellow } + +# ============================ the rollback-backup convention ============================ +# +# The C# twin is DarlingRollbackBackups and the two must stay identical. That is not a style preference: +# the service RECOGNISES these directories so it can report the whole set on one line instead of one +# warning each, and this script CREATES and PRUNES them. If the two spellings ever drift apart the service +# goes back to naming 46 directories individually and nobody notices, because each of those lines is +# perfectly true. DarlingDeployRollbackRetentionTests runs the function below against the C# predicate over +# a shared case table to make the drift impossible to ship. + +# What a rollback backup is called. Prefix plus at least one more character, matched case-insensitively +# because Windows paths are - a matcher stricter than the filesystem is a matcher that misses a directory +# the operator can see with their own eyes. +# +# No stamp parsing. The backlog this has to recognise was made over months by a procedure that has spelled +# its stamp more than one way, and a rule that only accepted today's spelling would leave yesterday's +# backups unrecognised - which is the entire complaint in #2525. The prefix is a namespace: everything +# inside it belongs to this procedure. +function Test-DarlingRollbackBackupName([string]$name) { + $prefix = '_rollback_manual_' + if ([string]::IsNullOrEmpty($name)) { return $false } + if ($name.Length -le $prefix.Length) { return $false } + return $name.StartsWith($prefix, [StringComparison]::OrdinalIgnoreCase) +} + +# The name for a backup taken now. Seconds are in it because two deploys in one minute is a rehearsal, not +# a hypothetical, and a stamp that collides silently merges two builds into one directory. +function New-DarlingRollbackBackupName([datetime]$whenUtc) { + return '_rollback_manual_' + $whenUtc.ToString('yyyyMMdd-HHmmss') +} + +# Every rollback backup in the install root, NEWEST FIRST. That ordering is a contract - the prune below +# keeps a prefix of this list - so it is produced in exactly one place. +# +# Ordered by LastWriteTimeUtc rather than by the stamps in the names, because the filesystem knows when a +# directory was written and the name only carries whatever spelling the procedure used that month. The name +# is the tiebreak so the ordering is still total when two backups share a timestamp. +# +# Note there is deliberately NO -Filter '_rollback_manual_*' here. Get-ChildItem's filter is handed to the +# filesystem, which matches a directory's Windows 8.3 SHORT name as well as its real one - so a wildcard +# can hand a DELETE a directory whose real name looks nothing like the pattern. DarlingStoreUpgrade guards +# that by re-checking the real name after the wildcard; enumerating everything and testing the real name is +# the same guard with the trap removed, and on a directory holding tens of entries it costs nothing. +function Get-DarlingRollbackBackups([string]$installRoot) { + if ([string]::IsNullOrWhiteSpace($installRoot)) { return @() } + if (-not (Test-Path -LiteralPath $installRoot -PathType Container)) { return @() } + + $all = @(Get-ChildItem -LiteralPath $installRoot -Directory -Force -ErrorAction SilentlyContinue) + $mine = @($all | Where-Object { Test-DarlingRollbackBackupName $_.Name }) + return @($mine | Sort-Object -Property LastWriteTimeUtc, Name -Descending) +} + +# The backups past retention, given the newest-first list Get-DarlingRollbackBackups produces. +# +# $keep counts the backup this run just took, so -KeepRollbacks 3 means "this one and the two before it". +# The floor of 1 is not defensive clutter: this function's output is fed straight to a recursive delete in +# an install directory, and the one input that must never be possible is the one that selects everything. +function Select-DarlingRollbackBackupsToPrune($backups, [int]$keep) { + $ordered = @($backups) + if ($keep -lt 1) { $keep = 1 } + if ($ordered.Count -le $keep) { return @() } + return @($ordered[$keep..($ordered.Count - 1)]) +} + +# True when the newest backup is recent enough that this run is a RE-RUN of an interrupted upgrade rather +# than a new deploy. +# +# The failure this prevents is specific and expensive. Step 8 dies on a locked DLL, leaving the tree half +# extracted; the operator does the right thing and runs the script again; without this, step 6 backs up the +# HALF-EXTRACTED tree, and now the newest rollback copy - the one anybody would reach for - is a mixture of +# two builds. Reusing the existing backup keeps the copy that was taken while the tree was still coherent. +# +# $nowUtc is a parameter rather than a call to Get-Date so the rule can be tested at a known instant. +# +# The elapsed time has to be non-negative as well as small, and that is not pedantry - it was a live bug +# caught by running this function against a planted tree. A backup whose timestamp is in the FUTURE (a +# clock that stepped backwards, a directory restored from elsewhere with its metadata) produces a negative +# elapsed time, which is less than any window, which reads as "recent" - and the upgrade would then skip +# taking a backup at all. The two ways to be wrong here are not symmetric: an extra backup costs 120 MB +# that the prune reclaims on the next deploy, and a missing one costs the rollback this whole procedure +# exists to provide. So anything the clock cannot vouch for falls through to taking a new backup. +function Test-DarlingRollbackBackupIsRecent($backups, [int]$withinMinutes, [datetime]$nowUtc) { + $ordered = @($backups) + if ($ordered.Count -eq 0) { return $false } + if ($withinMinutes -le 0) { return $false } + + $elapsed = ($nowUtc - $ordered[0].LastWriteTimeUtc).TotalMinutes + return ($elapsed -ge 0) -and ($elapsed -lt $withinMinutes) +} + +# ============================ the install tree ============================ + +# True when two paths name the same directory. One spelling of path equality for the whole script: the +# comparison decides whether an upgrade is allowed to touch a tree, and two hand-rolled copies of it is +# the same "two spellings of one rule" that #2525 is a case study in. +# +# Separator from the runtime rather than a hardcoded backslash, for the same reason as the process filter +# below - a rule that gates a destructive operation should be verifiable off the platform it ships to. +function Test-DarlingSamePath([string]$left, [string]$right) { + if ([string]::IsNullOrWhiteSpace($left) -or [string]::IsNullOrWhiteSpace($right)) { return $false } + + $sep = [IO.Path]::DirectorySeparatorChar + $alt = [IO.Path]::AltDirectorySeparatorChar + + try { + $l = [IO.Path]::GetFullPath($left).TrimEnd($sep, $alt) + $r = [IO.Path]::GetFullPath($right).TrimEnd($sep, $alt) + } + catch { + return $false + } + + return $l.Equals($r, [StringComparison]::OrdinalIgnoreCase) +} + +# Where the service is ACTUALLY installed, read from the registered ImagePath rather than guessed from +# where this script happens to be sitting. Returns $null when the service is not installed. +# +# The ImagePath is quoted when it contains spaces and 'C:\PerformanceMonitorDarling' usually does not, so +# both spellings are handled; a path that cannot be parsed returns $null and the caller asks for +# -InstallRoot rather than upgrading a directory it guessed at. +function Get-DarlingInstallRootFromService([string]$name) { + try { + $service = Get-CimInstance -ClassName Win32_Service -Filter "Name='$name'" -ErrorAction Stop + } + catch { + return $null + } + + if (-not $service) { return $null } + + $imagePath = $service.PathName + if ([string]::IsNullOrWhiteSpace($imagePath)) { return $null } + + $imagePath = $imagePath.Trim() + if ($imagePath.StartsWith('"')) { + $close = $imagePath.IndexOf('"', 1) + if ($close -gt 1) { $imagePath = $imagePath.Substring(1, $close - 1) } + } + else { + # An unquoted path with arguments after it: everything up to the .exe is the executable. + $exe = $imagePath.IndexOf('.exe', [StringComparison]::OrdinalIgnoreCase) + if ($exe -ge 0) { $imagePath = $imagePath.Substring(0, $exe + 4) } + } + + try { return [IO.Path]::GetDirectoryName([IO.Path]::GetFullPath($imagePath)) } + catch { return $null } +} + +# Every process whose executable lives under $root, so the copy can refuse instead of failing half-way. +# +# It NAMES them. It never kills one, and no version of this script ever should: the bundled PostgreSQL runs +# out of pg-runtime under this very directory, and a sweep that force-kills "everything under the install +# dir" takes the monitoring store down as its first act. Stopping the service stops the store properly; +# anything still holding the tree after that is a person's session, and a person can close it. +function Get-DarlingProcessesUnderPath([string]$root) { + $hits = @() + if ([string]::IsNullOrWhiteSpace($root)) { return @($hits) } + + try { $prefix = [IO.Path]::GetFullPath($root).TrimEnd('\') + '\' } + catch { return @($hits) } + + foreach ($process in @(Get-Process -ErrorAction SilentlyContinue)) { + $path = $null + # A process owned by another account throws on .Path rather than returning empty, and an + # inaccessible process is not evidence of anything - skip it and keep looking. + try { $path = $process.Path } catch { continue } + if ([string]::IsNullOrEmpty($path)) { continue } + if ($path.StartsWith($prefix, [StringComparison]::OrdinalIgnoreCase)) { $hits += $process } + } + + return @($hits) +} + +# True for a process that stopping the service will take with it: the service's own executable, and +# anything under pg-runtime (the bundled PostgreSQL, which the service starts and stops). +# +# This exists because the stop guard has to run TWICE, and the two runs are asking different questions. +# Before the service is stopped, every install is holding its own tree - the service exe is right there and +# the store's postmaster is under pg-runtime - so an unfiltered check refuses every real upgrade there has +# ever been. Filtering those two out leaves exactly the processes a service stop will NOT clear: an +# operator's own psql.exe, a shell whose working directory is the install folder, a Darling Viewer someone +# left open holding viewer\*.dll. Catching those BEFORE the stop costs nothing but a re-run; catching them +# after costs an outage. +# +# The Viewer is deliberately NOT on this list even though we ship it. It is a separate process that the +# service does not own and stopping the service does not close, so it holds viewer\ exactly as hard as any +# other application would. +# The separator comes from the runtime rather than being typed as '\'. This script only ever RUNS on +# Windows, but this particular predicate decides which processes are EXCUSED from a guard, and a rule that +# excuses things is the one worth being able to test on the machine it was written on. A hardcoded +# backslash makes it verifiable only on the box it already shipped to. +function Test-DarlingProcessStopsWithTheService([string]$processPath, [string]$installRoot) { + if ([string]::IsNullOrEmpty($processPath)) { return $false } + + $sep = [IO.Path]::DirectorySeparatorChar + try { $root = [IO.Path]::GetFullPath($installRoot).TrimEnd($sep, [IO.Path]::AltDirectorySeparatorChar) } + catch { return $false } + + if ($processPath.Equals(($root + $sep + 'PerformanceMonitor.Darling.Service.exe'), [StringComparison]::OrdinalIgnoreCase)) { + return $true + } + + # The trailing separator is what keeps pg-runtime-prev - the rescued previous runtime, a real directory + # in this layout - from matching a pg-runtime prefix test. + return $processPath.StartsWith(($root + $sep + 'pg-runtime' + $sep), [StringComparison]::OrdinalIgnoreCase) +} + +function Get-DarlingDirectoryBytes([string]$path) { + try { + $files = @(Get-ChildItem -LiteralPath $path -File -Recurse -Force -ErrorAction SilentlyContinue) + if ($files.Count -eq 0) { return [long]0 } + return [long](($files | Measure-Object -Property Length -Sum).Sum) + } + catch { + return [long]0 + } +} + +function Format-DarlingBytes([long]$bytes) { + if ($bytes -ge 1GB) { return ('{0:N2} GB' -f ($bytes / 1GB)) } + if ($bytes -ge 1MB) { return ('{0:N1} MB' -f ($bytes / 1MB)) } + if ($bytes -ge 1KB) { return ('{0:N0} KB' -f ($bytes / 1KB)) } + return "$bytes bytes" +} + +# Deletes the selected backups, EACH IN ITS OWN HANDLER. +# +# That is #1775's lesson paid for once already: the store's retained-copy sweep had its failure handling +# outside the loop, so one directory an antivirus scan still held abandoned the sweep for every other +# directory too - and kept abandoning it for as long as the condition lasted, which is how nothing aged out +# at all. One directory that cannot be deleted must cost exactly that directory. +# +# The try wraps ONLY the measure and the delete, and the reporting happens after it on a success flag. That +# is not tidiness. With the write-up inside the try, anything that went wrong while composing a LINE OF TEXT +# landed in the catch and was recorded as a delete failure - so the log said "could not remove X" about a +# directory that was already gone, and the returned failure count disagreed with the disk. A run of this +# function against a planted tree produced exactly that: two directories removed and two failures reported, +# for the same two directories. What the caller does with the answer (exit codes, "re-running is safe") is +# built on those counts, so they have to mean what they say. +function Remove-DarlingRollbackBackups($prunable) { + $removed = 0 + $reclaimed = [long]0 + $failures = @() + + foreach ($backup in @($prunable)) { + $bytes = [long]0 + $gone = $false + $reason = '' + + try { + $bytes = Get-DarlingDirectoryBytes $backup.FullName + Remove-Item -LiteralPath $backup.FullName -Recurse -Force -ErrorAction Stop + $gone = $true + } + catch { + $reason = $_.Exception.Message + } + + if ($gone) { + $removed++ + $reclaimed += $bytes + Note (" removed {0} ({1})" -f $backup.Name, (Format-DarlingBytes $bytes)) + } + else { + $failures += $backup.Name + Warn ("could not remove {0}: {1}. The other backups were still swept; delete this one by hand when whatever is holding it lets go." -f $backup.Name, $reason) + } + } + + return [pscustomobject]@{ + Removed = $removed + Reclaimed = $reclaimed + Failures = @($failures) + } +} + +# ============================ the install manifest, and files a build stopped shipping ============================ +# +# THE PROBLEM (#2529). An in-place upgrade is an OVERLAY, not a replacement. Expand-Archive -Force and +# Copy-Item -Recurse -Force write what the new build ships and delete nothing else, so a file the old +# version had and the new one dropped stays in the install tree forever: a dependency that went away, an +# assembly that changed name, a satellite-resource directory for a culture nothing localizes into any more, +# a whole runtimes\\lib\\ subtree stranded by a target-framework move. The last one is the one +# that bites. Those are .NET PROBING directories, so a stale assembly sitting in one is not inert clutter - +# it is a candidate for loading. +# +# IT IS MEASURED, not feared. The file lists of consecutive release zips, diffed: the Lite package dropped +# 44 shipped files across twelve consecutive releases - 43 of them in a single step, where a +# target-framework move stranded every runtimes\*\lib\net8.0\ assembly including two copies of +# Microsoft.Data.SqlClient.dll, plus one lone assembly in a later step. The Darling package has dropped +# none in the four steps it has existed for, which is the whole of its history. So the shape of the thing +# is: rare, bursty, tied to a packaging or framework change rather than to ordinary development - and when +# it does happen it arrives forty files at a time, in a probing path. +# +# THE AUTHORITY, and why it cannot produce a false positive. After every successful copy this script writes +# a manifest of the files that copy laid down. The next upgrade diffs its own payload against that manifest, +# and the difference is exactly "files one of OUR builds put here that this build does not ship". Every path +# it can name provably came out of one of our own payloads, which is what disposes of the whole class of +# false positives the obvious approach has. It cannot warn about darling.json, the DPAPI credential blobs, +# the .bak- config copies, the _rollback_manual_* backups, pg-runtime, or whatever an operator legitimately +# put here, because none of those was ever in a payload and so none of them is ever in the manifest. That is +# a property of where the list comes from, not a list of exceptions somebody has to keep up to date - which +# matters, because a list of exceptions that goes one entry out of date is #2525 again with a new subject. +# +# FAIL SAFE, AND LOUDLY. Two things have to be KNOWN before anything is removed: what an earlier build +# shipped, and what this one ships. If either cannot be determined - no manifest yet (every install that +# predates this change), a manifest that will not parse, one that disagrees with its own count, a source +# this cannot read - the answer is to remove NOTHING and say so on its own line. Guessing is not on the +# table: the wrong guess here deletes the store. + +# One spelling of a manifest path for the whole script. Manifest paths are relative to the install root and +# written with backslashes, because that is what the tree they describe uses; zip entries arrive with +# forward slashes and get converted here rather than at six call sites. +function ConvertTo-DarlingManifestPath([string]$path) { + if ([string]::IsNullOrWhiteSpace($path)) { return '' } + + $normalized = $path.Trim().Replace('/', '\') + while ($normalized.StartsWith('\')) { $normalized = $normalized.Substring(1) } + return $normalized.TrimEnd('\') +} + +# Resolves a manifest path against the install root, in the separator the RUNTIME uses. +# +# Not a hardcoded backslash, for the same reason Test-DarlingSamePath is not: this composes the absolute +# path that a delete is handed, and a rule that decides what gets deleted should be verifiable on a machine +# other than the one it already shipped to. Returns '' for anything that will not resolve, and the callers +# treat '' as "skip", never as "the install root". +function Join-DarlingInstallPath([string]$installRoot, [string]$relativePath) { + $normalized = ConvertTo-DarlingManifestPath $relativePath + if ([string]::IsNullOrWhiteSpace($normalized)) { return '' } + if ([string]::IsNullOrWhiteSpace($installRoot)) { return '' } + + $separator = [string][IO.Path]::DirectorySeparatorChar + try { return [IO.Path]::GetFullPath([IO.Path]::Combine($installRoot, $normalized.Replace('\', $separator))) } + catch { return '' } +} + +# True for RAW text that names an absolute location rather than something inside the install root. +# +# It exists as its own function because of where it has to be CALLED, which review had to point out twice. +# The first attempt put this test inside Test-DarlingManifestPathIsPreserved, ahead of normalization, which +# read correctly and could not fire: every caller of that predicate normalizes first, and normalization +# strips leading separators - so '\\fileserver\share\x.dll' had already become the ordinary-looking +# relative 'fileserver\share\x.dll' before the predicate ever saw it. A guard belongs where the raw value +# IS, not where it reads best, so this is called at the two points raw text enters the script: a line of a +# manifest, and an entry of a payload. Both refuse the WHOLE input rather than dropping the line, on the +# same reasoning as the file-count check - this script never writes an absolute path, so one coming back +# means the input was edited or is not ours, and a partly-believed record is worse than none. +# +# Test-DarlingManifestPathIsPreserved calls it too, on its own raw argument, for the caller that hands it +# one directly. +function Test-DarlingPathIsAbsolute([string]$rawPath) { + if ([string]::IsNullOrWhiteSpace($rawPath)) { return $false } + + $raw = $rawPath.Trim() + + # A colon is a drive qualifier or an alternate data stream, neither of which is a name a Windows file + # can otherwise carry. A leading separator is root-relative, and two of them are UNC. + if ($raw.Contains(':')) { return $true } + if ($raw.StartsWith('\') -or $raw.StartsWith('/')) { return $true } + + return [IO.Path]::IsPathRooted($raw) +} + +# THE EXCLUSION RULE. True for a path this procedure must never remove and must never record. +# +# It is applied THREE times and that is deliberate, because the three are different questions. On the way +# INTO a manifest, so a preserved path cannot be recorded and therefore cannot be nominated later. In +# Select-DarlingStaleFiles, so a manifest written by some older version of this script cannot nominate one +# either. And one statement before the delete in Remove-DarlingStaleFiles, because that is the function a +# future caller will reach for, and a guard that lives only in the caller upstream is a guard the next +# caller does not get. +# +# Everything here is a string test on a relative path, with no filesystem access at all, so it is the same +# answer on any machine and can be run against a table of cases rather than reasoned about. +function Test-DarlingManifestPathIsPreserved([string]$relativePath) { + if ([string]::IsNullOrWhiteSpace($relativePath)) { return $true } + + # A manifest path is RELATIVE and stays inside the install root. Asked of THIS function's raw argument, + # before normalization strips the leading separators that would answer it. It only fires for a caller + # that hands over un-normalized text - Remove-DarlingStaleFiles is the one that does - because the + # readers refuse an absolute path at the point they read it, which is where it can still be seen. + if (Test-DarlingPathIsAbsolute $relativePath) { return $true } + + $path = ConvertTo-DarlingManifestPath $relativePath + if ([string]::IsNullOrWhiteSpace($path)) { return $true } + + $segments = @($path.Split('\') | Where-Object { $_ }) + if ($segments.Count -eq 0) { return $true } + + foreach ($segment in $segments) { + if ($segment -eq '.' -or $segment -eq '..') { return $true } + } + + # THE STORE, and the single most important line in this file. pg-runtime holds the bundled PostgreSQL + # and pg-runtime-prev the rescued previous one; deleting either destroys the monitoring store this + # product exists to keep, and no backup in this procedure holds it. + # + # Matched as a PREFIX on the top-level name rather than as the two names we ship today. An enumeration + # is a list somebody has to maintain, and the cost of it being one entry out of date HERE is the store - + # so the rule is "the pg-runtime namespace is off limits" and a future pg-runtime- is covered + # before it exists. Nothing legitimate is given up: pg-runtime.zip is shipped by every build, so it is + # in both sides of every diff and could never have been in a difference anyway. + if ($segments[0].StartsWith('pg-runtime', [StringComparison]::OrdinalIgnoreCase)) { return $true } + + # This script's own rollback backups. The prune above owns those, and two different pieces of one + # script deleting the same directory is how you get a delete nobody can account for. + if (Test-DarlingRollbackBackupName $segments[0]) { return $true } + + $leaf = $segments[-1] + + # Operator config and its backups. The zip ships darling.sample.json and never darling.json, so on a + # manifest this script wrote these cannot fire - which is exactly the point of having them. They are + # the assertion that the manifest is what it claims to be, placed where being wrong costs a monitoring + # host its configuration. + if ($leaf.Equals('darling.json', [StringComparison]::OrdinalIgnoreCase)) { return $true } + if ($leaf.StartsWith('darling.json.bak-', [StringComparison]::OrdinalIgnoreCase)) { return $true } + + # DPAPI credential blobs. LocalMachine-scoped with an entropy constant that ships in an open-source + # repo, so READ access is the secret - but they are also not recoverable from anywhere else, and a + # deleted one is a credential gone rather than a file gone. + if ($leaf.EndsWith('.dpapi', [StringComparison]::OrdinalIgnoreCase)) { return $true } + + # The manifest itself. It is written after the copy and is therefore in no payload, so it can never + # appear in a difference either - but the file that decides what gets deleted should not be deletable + # by the thing it decides for. Kept as a literal rather than reaching for $manifestName so this + # function answers the same way when it is lifted out of the script and run on its own; the two are + # pinned together by DarlingDeployStaleFileTests. + if ($leaf.Equals('darling-install-manifest.txt', [StringComparison]::OrdinalIgnoreCase)) { return $true } + + return $false +} + +# The file list of the build being installed, read from the PAYLOAD rather than from anything we maintain. +# +# Deriving it is what makes the manifest trustworthy. A hand-kept list of "what we ship" would be one more +# thing to forget to update, and forgetting it here does not produce a stale entry in a document - it +# produces a file the next upgrade believes was dropped. +# +# Returns an Ok/Files/Reason object rather than a bare array on purpose. PowerShell unrolls an empty array +# returned from a function into $null, so "read nothing" and "failed to read" would arrive at the caller +# looking identical - and those two have to be told apart here more than anywhere else in this script. +function Get-DarlingPayloadFiles([string]$source, [bool]$sourceIsZip) { + if ([string]::IsNullOrWhiteSpace($source)) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = 'no source path was given' } + } + + $paths = @() + + try { + if ($sourceIsZip) { + Add-Type -AssemblyName 'System.IO.Compression.FileSystem' -ErrorAction SilentlyContinue + $archive = [IO.Compression.ZipFile]::OpenRead($source) + try { + foreach ($entry in $archive.Entries) { + # A directory entry has an empty Name and a FullName ending in '/'. Only files. + if ([string]::IsNullOrEmpty($entry.Name)) { continue } + + # Checked HERE, on the entry as the archive spells it, because this is the last moment + # an absolute path is still recognisable as one - ConvertTo-DarlingManifestPath on the + # next line strips the leading separators that say so. No build of ours produces such an + # entry, so one is evidence about the archive rather than about the file. + if (Test-DarlingPathIsAbsolute $entry.FullName) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the zip holds an entry with an absolute path ('$($entry.FullName)'), which no build of ours produces" } + } + + $paths += (ConvertTo-DarlingManifestPath $entry.FullName) + } + } + finally { + $archive.Dispose() + } + } + else { + $prefix = [IO.Path]::GetFullPath($source).TrimEnd('\', '/') + [IO.Path]::DirectorySeparatorChar + foreach ($file in @(Get-ChildItem -LiteralPath $source -File -Recurse -Force -ErrorAction Stop)) { + if (-not $file.FullName.StartsWith($prefix, [StringComparison]::OrdinalIgnoreCase)) { continue } + + # The prefix carries a trailing separator, so this substring cannot begin with one and the + # test below cannot fire. It is here anyway, so the two payload branches answer the same + # question in the same words: a reader who has to work out which of two branches checks what + # is a reader who will eventually get it wrong. + $relative = $file.FullName.Substring($prefix.Length) + if (Test-DarlingPathIsAbsolute $relative) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the source folder produced an absolute path ('$relative')" } + } + + $paths += (ConvertTo-DarlingManifestPath $relative) + } + } + } + catch { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = $_.Exception.Message } + } + + $files = @($paths | Where-Object { $_ }) + + # An EMPTY payload is not a build that ships nothing, it is a build this did not read. Everything + # downstream treats "the new build does not ship X" as grounds for removing X, so an empty list is the + # one answer that must never be handed on as a success: it makes every file an earlier build shipped + # stale in one step, which is the largest delete this code can express. + if ($files.Count -eq 0) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = 'the source listed no files at all, which is not a build' } + } + + return [pscustomobject]@{ Ok = $true; Files = $files; Reason = '' } +} + +# What the last build to be installed here laid down, or a REASON this cannot be answered. +# +# Every failure path returns Ok = $false with something an operator can read, and none of them returns a +# partial list. A manifest that half-parsed is not a smaller manifest, it is a manifest of unknown content, +# and the caller's whole contract is that it removes nothing when it does not know. +function Read-DarlingInstallManifest([string]$path) { + if ([string]::IsNullOrWhiteSpace($path) -or -not (Test-Path -LiteralPath $path -PathType Leaf)) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = 'this install has no manifest yet, so there is no record of what the build before this one shipped' } + } + + try { $lines = @(Get-Content -LiteralPath $path -ErrorAction Stop) } + catch { return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the manifest could not be read ($($_.Exception.Message))" } } + + $version = '' + $declared = -1 + $sawMarker = $false + $files = @() + + foreach ($line in $lines) { + if ($sawMarker) { + # THE PLACE THIS CHECK HAS TO BE. A manifest line is raw text off a disk this script does not + # control, and one line further on it has been normalized into something that looks relative + # whatever it was. Write-DarlingInstallManifest never emits an absolute path, so a manifest + # holding one has been hand-edited, corrupted, or written by something else - and the answer to + # a record that is not ours is to believe none of it. + if (Test-DarlingPathIsAbsolute $line) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the manifest holds an absolute path ('$($line.Trim())'), which this script never writes - it has been edited, corrupted, or written by something else" } + } + + $normalized = ConvertTo-DarlingManifestPath $line + if ($normalized) { $files += $normalized } + continue + } + + $trimmed = "$line".Trim() + if ($trimmed.Length -eq 0 -or $trimmed.StartsWith('#')) { continue } + if ($trimmed -eq '--- files ---') { $sawMarker = $true; continue } + + if ($trimmed.StartsWith('manifest-version ')) { + $version = $trimmed.Substring('manifest-version '.Length).Trim() + continue + } + + if ($trimmed.StartsWith('file-count ')) { + $parsed = 0 + if ([int]::TryParse($trimmed.Substring('file-count '.Length).Trim(), [ref]$parsed)) { $declared = $parsed } + continue + } + } + + if (-not $sawMarker) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = 'the manifest carries no file list' } + } + + if ($version -ne '1') { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the manifest is format version '$version' and this script only understands 1" } + } + + if ($declared -lt 0) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = 'the manifest does not say how many files it holds' } + } + + # The count is the TRUNCATION check. Writing goes temp-file-then-move, so half a manifest should not be + # observable at all; this is what catches the case where it was observed anyway - an interrupted + # restore, a hand edit, a copy that died. The direction a truncated manifest fails in is the safe one + # (fewer files nominated, never more), so this is not a safety guard - it is the rule that a record this + # script writes and later acts on does not get trusted while it disagrees with itself. + if ($declared -ne $files.Count) { + return [pscustomobject]@{ Ok = $false; Files = @(); Reason = "the manifest says it holds $declared file(s) but carries $($files.Count), so it is truncated or edited" } + } + + return [pscustomobject]@{ Ok = $true; Files = @($files); Reason = '' } +} + +# Records what this install now holds that one of our builds put there. +# +# Written LAST, after any removal, so an interrupted run leaves the old manifest in place and a re-run +# re-decides the same question from the same evidence. +function Write-DarlingInstallManifest([string]$path, $relativePaths, [string]$sourceName, [datetime]$whenUtc) { + $files = @() + $seen = New-Object 'System.Collections.Generic.HashSet[string]' ([StringComparer]::OrdinalIgnoreCase) + + foreach ($relative in @($relativePaths)) { + $normalized = ConvertTo-DarlingManifestPath $relative + if ([string]::IsNullOrWhiteSpace($normalized)) { continue } + + # Preserved paths never enter a manifest, and this is not the same guard as the one at the diff. + # The ONE route by which pg-runtime could ever reach a manifest is a -Source FOLDER that is a copy + # of a live install tree - somebody's staging directory made with Copy-Item from a box that has + # run - and that folder walk happens right here, in Get-DarlingPayloadFiles. Filtering on the way + # in means a path that must never be deleted cannot be written down, so it cannot be read back and + # acted on by a later run of a later version of this script. + # + # Handed $relative and NOT $normalized. Review caught this one: the predicate's absolute-path half + # only works on raw text, because normalization has already stripped the leading separators that a + # UNC or root-relative path is recognised by. Passing the normalized value left that half doing + # nothing here - safe only because the readers upstream refuse an absolute path first, which is an + # invariant living in the CALLERS rather than in this function. A guard that holds only because of + # who happens to call it today is a guard with an expiry date nobody wrote down. + if (Test-DarlingManifestPathIsPreserved $relative) { continue } + if (-not $seen.Add($normalized)) { continue } + $files += $normalized + } + + # SORTED, so two manifests can be diffed by eye and a rebuild of the same payload produces the + # same file rather than the same set in a different order. Neither a zip's entry order nor a + # directory walk's is stable enough to hand an operator as a record. + $files = @($files | Sort-Object) + + $header = @( + '# PerformanceMonitor Darling install manifest.', + '#', + '# Written by upgrade-darling.ps1 after every successful copy. Every line below the marker is one', + '# file a build of ours laid down in this directory, spelled relative to it. The next upgrade diffs', + '# its own payload against this list to find the files an earlier build shipped and the new one does', + '# not, which is the only way an in-place upgrade can ever remove one (#2529).', + '#', + '# Paths this procedure may never remove are left out of the list ENTIRELY, so a manifest cannot', + '# nominate one: anything under pg-runtime, darling.json and its .bak- copies, *.dpapi credential', + '# blobs, the _rollback_manual_* backups, and this file.', + '#', + '# Deleting this file is safe. The next upgrade will say it cannot tell what the previous build', + '# shipped, remove nothing, and write a fresh one.', + 'manifest-version 1', + "build-source $sourceName", + ('written-utc ' + $whenUtc.ToString('yyyy-MM-ddTHH:mm:ssZ')), + "file-count $($files.Count)", + '--- files ---' + ) + + # Temp file then move, so an interrupted write never leaves a manifest that is half a manifest. The + # file-count header is the backstop for the day it happens anyway. + $temporary = $path + '.tmp' + + try { + Set-Content -LiteralPath $temporary -Value (@($header) + @($files)) -Encoding UTF8 -Force -ErrorAction Stop + Move-Item -LiteralPath $temporary -Destination $path -Force -ErrorAction Stop + return [pscustomobject]@{ Ok = $true; Count = $files.Count; Reason = '' } + } + catch { + $reason = $_.Exception.Message + try { + if (Test-Path -LiteralPath $temporary) { Remove-Item -LiteralPath $temporary -Force -ErrorAction SilentlyContinue } + } + catch { + # A leftover .tmp is worth nobody's error. The next write overwrites it. + } + + return [pscustomobject]@{ Ok = $false; Count = 0; Reason = $reason } + } +} + +# The files an earlier build shipped and this one does not. Pure: two lists in, one list out, no clock and +# no filesystem, so the rule that decides what a delete is handed can be run against a table of cases. +function Select-DarlingStaleFiles($previousPaths, $currentPaths) { + $shipped = New-Object 'System.Collections.Generic.HashSet[string]' ([StringComparer]::OrdinalIgnoreCase) + + foreach ($path in @($currentPaths)) { + $normalized = ConvertTo-DarlingManifestPath $path + if ($normalized) { [void]$shipped.Add($normalized) } + } + + # THE ONE INPUT THAT MUST NOT BE EXPRESSIBLE, the same shape as Select-DarlingRollbackBackupsToPrune's + # floor of 1. An empty "what this build ships" makes EVERY file the previous build shipped stale at + # once. It cannot come from a real payload - Get-DarlingPayloadFiles refuses to call an empty list a + # success, and the script has already refused a source with no service exe in it - so arriving here + # with an empty set means something upstream failed and reported success. Selecting nothing is the + # answer that is right either way. + if ($shipped.Count -eq 0) { return @() } + + $stale = @() + $seen = New-Object 'System.Collections.Generic.HashSet[string]' ([StringComparer]::OrdinalIgnoreCase) + + foreach ($path in @($previousPaths)) { + $normalized = ConvertTo-DarlingManifestPath $path + if ([string]::IsNullOrWhiteSpace($normalized)) { continue } + + # A file the NEW build ships is never removable, whatever a manifest says. Case-insensitively, + # because Windows filenames are: a build that respelled wwwroot\js\App.js as app.js ships the same + # file under a name an ordinal comparison calls absent, and removing it would delete what the copy + # had just written. + if ($shipped.Contains($normalized)) { continue } + + # $path, not $normalized, for the reason spelled out in Write-DarlingInstallManifest: the + # predicate's absolute-path half can only see a UNC or root-relative path in RAW text, and + # normalization has already removed what makes it one. The membership test above is a different + # question and correctly uses the normalized form, which is why the two are not the same argument. + if (Test-DarlingManifestPathIsPreserved $path) { continue } + if (-not $seen.Add($normalized)) { continue } + $stale += $normalized + } + + return @($stale | Sort-Object) +} + +# Deletes the selected files, EACH IN ITS OWN HANDLER - #1775's lesson, the same one +# Remove-DarlingRollbackBackups is built around. One file an antivirus scan still holds must cost exactly +# that file, and the write-up happens outside the try on a success flag so that a failure while composing a +# LINE OF TEXT cannot be recorded as a delete failure for a file that is already gone. +# +# Returns four lists rather than a count, because they mean four different things to an operator: what went, +# what was refused by a guard, what was tried and would not go, and which directories emptied out. +# +# $shippedPaths is what the build being installed ships, and it is a PARAMETER rather than something the +# caller is trusted to have filtered for. Select-DarlingStaleFiles already excludes them, and that is +# exactly the argument for asking again here: this is the function that performs the delete, so the +# property "a file the new build ships is never deletable" has to hold at THIS call and not only at the one +# that happens to precede it today. An empty $shippedPaths therefore refuses everything, on the same +# reasoning as the empty-payload guard in Select-DarlingStaleFiles: a caller that cannot say what the build +# ships has not earned a delete. +function Remove-DarlingStaleFiles([string]$installRoot, $relativePaths, $shippedPaths) { + $removed = @() + $refused = @() + $failures = @() + $emptied = @() + $parents = @() + $reclaimed = [long]0 + + try { $root = [IO.Path]::GetFullPath($installRoot).TrimEnd('\', '/') } + catch { + return [pscustomobject]@{ Removed = @(); Refused = @($relativePaths); Failures = @(); Emptied = @(); Reclaimed = [long]0 } + } + + $shipped = New-Object 'System.Collections.Generic.HashSet[string]' ([StringComparer]::OrdinalIgnoreCase) + foreach ($path in @($shippedPaths)) { + $normalized = ConvertTo-DarlingManifestPath $path + if ($normalized) { [void]$shipped.Add($normalized) } + } + + if ($shipped.Count -eq 0) { + return [pscustomobject]@{ Removed = @(); Refused = @($relativePaths); Failures = @(); Emptied = @(); Reclaimed = [long]0 } + } + + $prefix = $root + [IO.Path]::DirectorySeparatorChar + + foreach ($relative in @($relativePaths)) { + if (Test-DarlingManifestPathIsPreserved $relative) { $refused += $relative; continue } + + # A file the NEW build ships, whatever the caller thinks. Case-insensitively, because the copy that + # just ran wrote it under whichever spelling this build uses and Windows would hand a delete the + # same file under either. + if ($shipped.Contains((ConvertTo-DarlingManifestPath $relative))) { $refused += $relative; continue } + + # CONTAINMENT, and it is a different question from the predicate above rather than a repeat of it. + # That one reads a STRING; this resolves the path the filesystem would actually open, so it is what + # catches a spelling the string rule did not anticipate. Anything resolving outside the install + # root is refused and REPORTED - a delete that quietly did not happen is worse than one that says + # so, and on a manifest this script wrote neither of these can fire at all. + $full = Join-DarlingInstallPath $root $relative + if ([string]::IsNullOrWhiteSpace($full) -or -not $full.StartsWith($prefix, [StringComparison]::OrdinalIgnoreCase)) { + $refused += $relative + continue + } + + if (-not (Test-Path -LiteralPath $full -PathType Leaf)) { continue } + + $bytes = [long]0 + $gone = $false + $reason = '' + + try { + $bytes = [long]((Get-Item -LiteralPath $full -Force).Length) + Remove-Item -LiteralPath $full -Force -ErrorAction Stop + $gone = $true + } + catch { + $reason = $_.Exception.Message + } + + if ($gone) { + $removed += $relative + $reclaimed += $bytes + $parents += [IO.Path]::GetDirectoryName($full) + Note (" removed {0} ({1})" -f $relative, (Format-DarlingBytes $bytes)) + } + else { + $failures += $relative + Warn ("could not remove the stale file {0}: {1}. The others were still swept; re-running the upgrade retries this one." -f $relative, $reason) + } + } + + # The directories that went empty because everything in them was stale. + # + # This is not tidiness, it is the difference between fixing a problem and moving it. The service's + # install-layout report classifies a satellite-resource directory STRUCTURALLY - one holding nothing but + # *.resources.dll - and an EMPTY directory deliberately fails that test, so it is reported as "not part + # of the product's layout" on every single service start. Removing the last file out of de\ and leaving + # the shell behind would convert one stale file into a permanent startup warning: the same + # guard-goes-too-loud outcome #2525 was about, this time caused by our own cleanup. + # + # NON-RECURSIVE deletes, deepest first, each walking up only while the parent is empty too. A directory + # delete that cannot recurse cannot destroy anything - the worst it can do is fail - which is what makes + # this safe to do at all. Deepest first because a shallow directory still holding a soon-to-go child + # would stop the walk and never be revisited. + $unique = @($parents | Where-Object { $_ } | Sort-Object -Unique) + $deepestFirst = @($unique | Sort-Object -Property Length -Descending) + + foreach ($directory in $deepestFirst) { + $current = $directory + while ($current -and $current.StartsWith($prefix, [StringComparison]::OrdinalIgnoreCase)) { + $next = [IO.Path]::GetDirectoryName($current) + try { + if (@(Get-ChildItem -LiteralPath $current -Force -ErrorAction Stop).Count -gt 0) { break } + [IO.Directory]::Delete($current) + } + catch { + break + } + + $emptied += $current.Substring($prefix.Length) + $current = $next + } + } + + return [pscustomobject]@{ + Removed = @($removed) + Refused = @($refused) + Failures = @($failures) + Emptied = @($emptied) + Reclaimed = $reclaimed + } +} + +# ============================ preamble ============================ + +if (-not ([Security.Principal.WindowsPrincipal][Security.Principal.WindowsIdentity]::GetCurrent()).IsInRole([Security.Principal.WindowsBuiltInRole]::Administrator)) { + Fail "Run this from an ELEVATED PowerShell. Stopping the service, writing to the install directory, and starting it again all need it." +} + +if ([string]::IsNullOrWhiteSpace($InstallRoot)) { + $InstallRoot = Get-DarlingInstallRootFromService $serviceName + if ([string]::IsNullOrWhiteSpace($InstallRoot)) { + Fail "The '$serviceName' service is not installed (or its ImagePath could not be read), so there is nothing here to upgrade. Install with install-darling.ps1, or pass -InstallRoot to point at the tree you mean." + } +} + +try { $InstallRoot = [IO.Path]::GetFullPath($InstallRoot).TrimEnd('\') } +catch { Fail "'$InstallRoot' is not a path this script can resolve." } + +if (-not (Test-Path -LiteralPath $InstallRoot -PathType Container)) { + Fail "The install directory '$InstallRoot' does not exist." +} + +if (-not (Test-Path -LiteralPath (Join-Path $InstallRoot $serviceExeName))) { + Fail "'$InstallRoot' does not hold $serviceExeName, so it is not a Darling install. Refusing to copy a build over it." +} + +Note "Install directory: $InstallRoot" + +# ============================ -ListRollbacks / -PruneOnly ============================ +# +# Both run against a LIVE service on purpose. Neither touches a binary, and the backups are inert copies +# that nothing has open - so making an operator stop their monitoring host to reclaim 5 GB would be asking +# them to take an outage to clean up after us. -PruneOnly is what the service's own layout report tells +# people to run, and it has to be something they can run at 3pm on a Tuesday. + +$backups = Get-DarlingRollbackBackups $InstallRoot + +if ($ListRollbacks -or $PruneOnly) { + if ($backups.Count -eq 0) { + Good "No rollback backups in $InstallRoot." + exit 0 + } + + $prunable = Select-DarlingRollbackBackupsToPrune $backups $KeepRollbacks + $total = [long]0 + foreach ($backup in $backups) { $total += (Get-DarlingDirectoryBytes $backup.FullName) } + + Note ("{0} rollback backup(s), {1} total. Keeping the newest {2}:" -f $backups.Count, (Format-DarlingBytes $total), $KeepRollbacks) + + $prunableNames = @($prunable | ForEach-Object { $_.Name }) + foreach ($backup in $backups) { + $verdict = if ($prunableNames -contains $backup.Name) { 'prune' } else { 'KEEP ' } + Note (" [{0}] {1} ({2}, last written {3:yyyy-MM-dd HH:mm} UTC)" -f $verdict, $backup.Name, (Format-DarlingBytes (Get-DarlingDirectoryBytes $backup.FullName)), $backup.LastWriteTimeUtc) + } + + if ($ListRollbacks) { + Note "Nothing was changed (-ListRollbacks). Re-run with -PruneOnly to remove the ones marked prune." + exit 0 + } + + if ($prunable.Count -eq 0) { + Good "Nothing to prune." + exit 0 + } + + $result = Remove-DarlingRollbackBackups $prunable + Good ("Removed {0} rollback backup(s), reclaiming {1}." -f $result.Removed, (Format-DarlingBytes $result.Reclaimed)) + if ($result.Failures.Count -gt 0) { + Warn ("{0} could not be removed: {1}. Re-running is safe and will retry them." -f $result.Failures.Count, ($result.Failures -join ', ')) + exit 1 + } + + exit 0 +} + +# ============================ the source build ============================ + +if ([string]::IsNullOrWhiteSpace($Source)) { + Fail "No -Source given and this script is not running from a file, so there is nothing to install from." +} + +try { $Source = [IO.Path]::GetFullPath($Source).TrimEnd('\') } +catch { Fail "'$Source' is not a path this script can resolve." } + +if (-not (Test-Path -LiteralPath $Source)) { + Fail "The source '$Source' does not exist." +} + +$sourceIsZip = (Test-Path -LiteralPath $Source -PathType Leaf) -and $Source.EndsWith('.zip', [StringComparison]::OrdinalIgnoreCase) +$sourceRoot = if ($sourceIsZip) { [IO.Path]::GetDirectoryName($Source) } else { $Source } + +# THE SELF-OVERWRITE REFUSAL. +# +# This script ships inside the zip and therefore also lives in the install root, so an upgrade run from the +# INSTALLED copy would write the new build over the .ps1 that PowerShell is reading line by line. Best case +# the copy fails on the lock and the tree is half written with the service stopped; worst case the script's +# own remaining lines change underneath it. Neither is worth a clever workaround like relaunching from a +# temp copy: a refusal that names the fix is understood in one read and cannot go subtly wrong. +# +# The condition is WHERE THIS SCRIPT IS, not where the source is. Keying it off the source was the first +# spelling and it had a hole big enough to drive the whole failure through: a zip sitting inside the install +# directory - which is exactly where someone downloads it - has a source root equal to the install root and +# would have been waved past by any rule about folders. The hazard is "the file being executed is about to +# be overwritten", so that is the thing to ask about. +# +# Scoped to the copy path. -PruneOnly and -ListRollbacks exit above this and write no binaries, so the +# installed copy is exactly the right thing to run for those - and it is the one the service's report names. +$runningFrom = $PSScriptRoot +if (Test-DarlingSamePath $runningFrom $InstallRoot) { + Fail "This script is running from the install directory, so the upgrade would write the new build over the copy of itself that PowerShell is currently reading. Extract the new zip to a staging folder (e.g. C:\staging\) and run ITS upgrade-darling.ps1 instead. To prune rollback backups from the installed copy, use -PruneOnly, which copies nothing." +} + +# And the degenerate case the rule above does not cover: a source folder that IS the install directory, +# handed in from a script running somewhere else. Copying a tree over itself is not an upgrade. +if (-not $sourceIsZip -and (Test-DarlingSamePath $sourceRoot $InstallRoot)) { + Fail "The source folder is the install directory itself, so there is nothing to upgrade from. Point -Source at the new build." +} + +if ($sourceIsZip) { + $expected = $Sha256 + + if ([string]::IsNullOrWhiteSpace($expected)) { + $sums = Join-Path ([IO.Path]::GetDirectoryName($Source)) 'SHA256SUMS.txt' + if (Test-Path -LiteralPath $sums) { + $leaf = [IO.Path]::GetFileName($Source) + foreach ($line in (Get-Content -LiteralPath $sums)) { + # ' ' and ' *' are both in the wild; splitting on whitespace and + # comparing the leaf handles either without a regex nobody can read. + $parts = @($line -split '\s+' | Where-Object { $_ }) + if ($parts.Count -ge 2 -and $parts[-1].TrimStart('*').Equals($leaf, [StringComparison]::OrdinalIgnoreCase)) { + $expected = $parts[0] + break + } + } + } + } + + if ([string]::IsNullOrWhiteSpace($expected)) { + if (-not $SkipHashCheck) { + Fail "No SHA256 for '$Source' - pass -Sha256 , put SHA256SUMS.txt beside the zip, or say -SkipHashCheck. This overwrites the binaries of a running monitoring host; an unverified zip is not something to find out about afterwards." + } + Warn "Proceeding with an UNVERIFIED source zip (-SkipHashCheck)." + } + else { + $actual = (Get-FileHash -LiteralPath $Source -Algorithm SHA256).Hash + if (-not $actual.Equals($expected.Trim(), [StringComparison]::OrdinalIgnoreCase)) { + Fail "SHA256 mismatch on '$Source'. Expected $expected, got $actual. Nothing has been stopped or copied." + } + Good "Source zip SHA256 verified." + } +} +else { + if (-not (Test-Path -LiteralPath (Join-Path $Source $serviceExeName))) { + Fail "'$Source' does not hold $serviceExeName, so it is not an extracted Darling build." + } + Warn "The source is a folder, so this script cannot verify it - verify the zip's SHA256 before you extract it." +} + +# ============================ the service has to exist ============================ +# +# The auto-resolve path cannot reach here without a registered service - it reads the install root out of +# the ImagePath - but an explicit -InstallRoot skips that check entirely. A tree holding the binaries of a +# service that was renamed, removed, or never registered then sails through the stop guard, the backup, the +# prune and the copy, and falls over at Start-Service with a raw terminating error instead of one of this +# script's own messages. Failing HERE costs nothing and says what to do; failing there costs a completed +# copy, a stopped-that-was-never-running service, and an error nobody can interpret. +# +# Deliberately not applied to -PruneOnly, which exits above: reclaiming disk from a tree whose service is +# gone is a perfectly reasonable thing to want, and it copies nothing. +if (-not (Get-Service -Name $serviceName -ErrorAction SilentlyContinue)) { + Fail "The '$serviceName' service is not registered on this machine, so there is nothing for this copy to stop and start around it. NOTHING has been stopped or copied. If '$InstallRoot' is a staging tree rather than an install, you want install-darling.ps1; if you only meant to reclaim disk, re-run with -PruneOnly, which needs no service." +} + +# ...and it has to be THIS install's service. +# +# "A service by that name exists" and "that service runs from the directory I am about to overwrite" are +# different claims, and the gap between them is a SILENT SUCCESS - the worst shape a failure can take here. +# Stop-Service and Start-Service act on the service BY NAME, never by path, so an -InstallRoot pointing at +# a stale copy of the tree (an old runbook, a paste from another box, a leftover directory that still holds +# the exe and therefore passes every check above) produces a run in which nothing is detectably wrong: the +# REAL service is stopped - a real outage on a monitoring host - the new build is laid down in a directory +# nothing reads, the real service is restarted on its old untouched binaries, darling.json is unchanged +# because it was never touched, and the script reports success end to end. The operator believes they +# upgraded. Nothing did, and the next person to look will be debugging a version that never shipped. +# +# Not gated on whether -InstallRoot was passed. When it was not, the two are equal by construction and this +# costs a registry read; gating it would make the check absent for precisely the one caller who needs it. +$registeredRoot = Get-DarlingInstallRootFromService $serviceName + +if ([string]::IsNullOrWhiteSpace($registeredRoot)) { + # Registered but unreadable ImagePath. Not a refusal - we know the service exists and the auto-resolve + # path would have failed earlier - but the operator should know this was not confirmed rather than + # assume it was. + Warn "Could not read the '$serviceName' service's ImagePath, so this could not confirm that '$InstallRoot' is the directory the service actually runs from. Check that before trusting the result." +} +elseif (-not (Test-DarlingSamePath $registeredRoot $InstallRoot)) { + Fail "The '$serviceName' service runs from '$registeredRoot', not from '$InstallRoot'. NOTHING has been stopped or copied. Stopping and starting act on the service by NAME, so upgrading '$InstallRoot' would have taken the real service down, written the new build somewhere it does not read, and brought it back up on its old binaries - reporting success the whole way. Drop -InstallRoot to upgrade the registered install, or pass -InstallRoot '$registeredRoot' if that is really the one you meant." +} + +# ============================ the stop guard ============================ + +# PHASE ONE, before anything is stopped: the processes a service stop will NOT clear. +# +# The service's own exe and the bundled PostgreSQL under pg-runtime are filtered out here because they are +# about to be stopped on purpose. Without that filter this guard refuses EVERY upgrade of a running install +# - the service is always holding its own tree - which is a guard that fails closed on the happy path and +# trains people to pass -SkipStopGuard, i.e. a guard that has stopped guarding by being unusable. +$holders = @(Get-DarlingProcessesUnderPath $InstallRoot | + Where-Object { -not (Test-DarlingProcessStopsWithTheService $_.Path $InstallRoot) }) + +if ($holders.Count -gt 0 -and -not $SkipStopGuard) { + # Parenthesised before -join on purpose: `$x | ForEach-Object { ... } -join ', '` binds -join to + # ForEach-Object as a parameter and throws, which is a fine way to lose a deploy to a formatting bug. + $names = @($holders | ForEach-Object { "$($_.ProcessName) (pid $($_.Id))" }) + Note "Processes are running out of the install tree that stopping the service will not close:" + foreach ($name in $names) { Note " $name" } + Fail "Close them and re-run. NOTHING has been stopped or copied - the service is still running and the install is untouched. Do NOT kill them blindly. If these are your own psql.exe or a shell sitting in the install directory, or a Darling Viewer you left open, just exit them. Use -SkipStopGuard only if you are certain the copy will not hit a locked file." +} + +# ============================ stop, back up, prune, copy, start ============================ +# +# From here on a failure leaves the service DOWN. Every step below is idempotent and the failure messages +# say so, because "run it again" is the correct advice and an operator staring at a stopped monitoring +# service should not have to work that out. + +$configPath = Join-Path $InstallRoot $configName +$configHashBefore = if (Test-Path -LiteralPath $configPath) { (Get-FileHash -LiteralPath $configPath -Algorithm SHA256).Hash } else { $null } + +# The previous build's manifest, read BEFORE the copy for the same reason the config hash above is taken +# before it - and it took review to see why that reason applies here too (#2529). +# +# A -Source FOLDER is copied wholesale: Copy-Item -Path "$Source\*" -Recurse -Force takes everything in it, +# unfiltered. A staging directory made from a live install therefore carries THAT install's manifest, and +# -Force lays it over this one. Reading afterwards would compute this run's stale-file list against some +# other box's history - the one scenario the manifest's own write-side filter already names as expected, +# arriving from the other direction. It self-heals on the next upgrade, because the manifest written at the +# end comes from the real payload, but one cycle of a wrong answer is one too many for a list that feeds a +# delete. +# +# The copy cannot change the answer to "what did the build before this one lay down here", so before it is +# not merely safe, it is the only correct time to ask. +$manifestPath = Join-Path $InstallRoot $manifestName +$previousManifest = Read-DarlingInstallManifest $manifestPath + +# Status is re-read here rather than carried down from the existence check above: it is a snapshot, and +# between the two the service can legitimately have been stopped by someone else or crashed on its own. +if ((Get-Service -Name $serviceName).Status -ne 'Stopped') { + Note "Stopping '$serviceName'..." + Stop-Service -Name $serviceName -Force + try { (Get-Service -Name $serviceName).WaitForStatus('Stopped', [TimeSpan]::FromMinutes(2)) } + catch { Fail "'$serviceName' did not reach Stopped within two minutes. Nothing has been copied. Check what it is waiting on and re-run." } +} +Good "Service is stopped." + +# PHASE TWO, now that it is down: anything STILL holding the tree, with no exclusions at all. +# +# Everything phase one filtered out should be gone by now, so a hit here is the interesting case rather +# than the normal one - most often a postmaster under pg-runtime that outlived the service stop, which is +# precisely the process nothing may kill. Phase one cannot see this and phase two cannot see phase one's +# cases without an outage, which is why there are two. +$stillHolding = Get-DarlingProcessesUnderPath $InstallRoot +if ($stillHolding.Count -gt 0 -and -not $SkipStopGuard) { + $names = @($stillHolding | ForEach-Object { "$($_.ProcessName) (pid $($_.Id))" }) + Note "The service is stopped, but processes are STILL running out of the install tree:" + foreach ($name in $names) { Note " $name" } + Fail "Nothing has been copied, so the install is intact - but the service is now STOPPED. Either close these and re-run (safe, and it will reuse the backup it is about to take), or abandon the upgrade with: Start-Service '$serviceName'. Do NOT kill anything under $InstallRoot\pg-runtime - that is the bundled PostgreSQL and killing it takes the store down; give a postmaster that outlived the stop a few seconds and re-run." +} + +$backups = Get-DarlingRollbackBackups $InstallRoot +$nowUtc = [datetime]::UtcNow + +if (Test-DarlingRollbackBackupIsRecent $backups $BackupWindowMinutes $nowUtc) { + Note ("Reusing the rollback backup {0}, taken {1:N0} minute(s) ago - this looks like a re-run of an interrupted upgrade, and a second backup now would copy a half-upgraded tree over the good one." -f $backups[0].Name, ($nowUtc - $backups[0].LastWriteTimeUtc).TotalMinutes) +} +else { + $backupPath = Join-Path $InstallRoot (New-DarlingRollbackBackupName $nowUtc) + Note "Backing up the install root's files to $backupPath ..." + try { + New-Item -ItemType Directory -Path $backupPath -Force | Out-Null + # FILES only, not subdirectories - the shape the documented procedure has always used, and the + # reason a backup is ~120 MB rather than ~1 GB. pg-runtime is the bundled PostgreSQL (extracted on + # first run, and hundreds of megabytes of it), and viewer\ / wwwroot\ / runtimes\ all come back + # from the zip. + # + # SO A BACKUP IS NOT A WHOLE-TREE SNAPSHOT, and the difference matters in exactly one scenario: + # a copy that dies PARTWAY can leave viewer\ / wwwroot\ / runtimes\ mixed old-and-new, and + # restoring these files over the top does not unmix them. The complete revert is these files PLUS + # the previous version's zip re-extracted. Backing those directories up instead was considered and + # rejected: it multiplies what #2525 is about - retained disk - to cover a case whose real fix is + # re-extracting a zip you still have, and it still would not cover pg-runtime. The failure paths + # below say this rather than leaving an operator to discover it while recovering. + Get-ChildItem -LiteralPath $InstallRoot -File -Force | Copy-Item -Destination $backupPath -Force + + # HARDEN THE SECRETS WE JUST COPIED (#2574). darling.json holds every monitored server's + # encryptedPassword plus the MCP and web tokens, all DPAPI LocalMachine scope with an entropy + # constant published in this open-source repo - so anything that can READ a copy can decrypt the + # lot, which is #1647's finding and the reason the LIVE file is hardened. + # + # Copy-Item into a new directory takes the DESTINATION's inherited DACL, not the source file's, so + # the copy lands with whatever the install root grants. Measured on a real install root: BUILTIN\Users + # ReadAndExecute - the inherited-from-C:\ DACL #1647 called out. Every retained backup was therefore + # a config readable by any local user, on a box where the live file is locked down. + # + # No INTERACTIVE grant here, deliberately, and unlike the live file: the Viewer and the CLI verbs + # read darling.json, and nothing reads a backup copy except a human recovering, who is an + # administrator by then. install-darling.ps1 draws the same line for the .bak-* copies. + # + # Best-effort and non-fatal: a backup that is taken but not hardened is strictly better than an + # upgrade that refuses to proceed, and the operator is told which. + $secretCopies = @(Get-ChildItem -LiteralPath $backupPath -File -Force -ErrorAction SilentlyContinue | + Where-Object { $_.Name -ieq 'darling.json' -or $_.Extension -ieq '.dpapi' }) + + # An enumeration that fails - a locked handle, an AV scan mid-copy - yields an EMPTY set, not an + # error, so without this the loop below would simply not run and nothing would be said. That is the + # silent-exposure shape this whole change is about, so the absence is checked against what the + # install root actually holds rather than assumed benign. + if ($secretCopies.Count -eq 0 -and (Test-Path -LiteralPath (Join-Path $InstallRoot 'darling.json'))) { + Warn "Found no darling.json in the rollback backup to restrict, though the install root has one. The backup may hold an unprotected copy of your credentials - check $backupPath by hand, or delete it." + } + + foreach ($copy in $secretCopies) { + try { + $wk = [System.Security.Principal.WellKnownSidType] + $systemSid = New-Object System.Security.Principal.SecurityIdentifier($wk::LocalSystemSid, $null) + $adminsSid = New-Object System.Security.Principal.SecurityIdentifier($wk::BuiltinAdministratorsSid, $null) + $acl = New-Object System.Security.AccessControl.FileSecurity + $acl.SetAccessRuleProtection($true, $false) + $acl.AddAccessRule((New-Object System.Security.AccessControl.FileSystemAccessRule($systemSid, 'FullControl', 'Allow'))) + $acl.AddAccessRule((New-Object System.Security.AccessControl.FileSystemAccessRule($adminsSid, 'FullControl', 'Allow'))) + Set-Acl -LiteralPath $copy.FullName -AclObject $acl + + # VERIFY, do not assume (#1957). A Set-Acl that returns without throwing is not proof the + # file is protected: install-darling.ps1 carries the same re-read for the same files, added + # after a permissions call that appeared to succeed left them readable on three consecutive + # field installs. Trusting the absence of an exception is precisely how that went unnoticed + # for months, and the stake here is every monitored server's encrypted credentials. + $after = Get-Acl -LiteralPath $copy.FullName + $stillOpen = @($after.Access | Where-Object { $_.IdentityReference -match 'Users|Everyone|Authenticated' }) + if ((-not $after.AreAccessRulesProtected) -or $stillOpen.Count -gt 0) { + Warn "Restricting $($copy.Name) in the rollback backup reported success but the file is STILL readable by ordinary users (protected=$($after.AreAccessRulesProtected)). It holds encrypted credentials - restrict it by hand, or delete the backup." + } + } + catch { + Warn "Could not restrict $($copy.Name) in the rollback backup ($($_.Exception.Message)). It holds encrypted credentials and inherited the install root's permissions - restrict it by hand, or delete the backup." + } + } + } + catch { + Fail "Could not take the rollback backup ($($_.Exception.Message)). The service is STOPPED and NOTHING has been overwritten, so the install is intact: start it with 'Start-Service ''$serviceName''', or free up disk and re-run." + } + + $backups = Get-DarlingRollbackBackups $InstallRoot + Good ("Rollback backup taken ({0})." -f (Format-DarlingBytes (Get-DarlingDirectoryBytes $backupPath))) +} + +# Pruning comes AFTER the backup, so the copy this run just took is counted among the ones kept and the +# tree never passes through a moment with fewer rollback points than retention promises. +$prunable = Select-DarlingRollbackBackupsToPrune $backups $KeepRollbacks +if ($prunable.Count -gt 0) { + Note ("Pruning {0} rollback backup(s) past the newest {1}:" -f $prunable.Count, $KeepRollbacks) + $result = Remove-DarlingRollbackBackups $prunable + Good ("Reclaimed {0}." -f (Format-DarlingBytes $result.Reclaimed)) +} + +# THIS COPY IS AN OVERLAY, not a replacement. Expand-Archive -Force (and the Copy-Item -Recurse -Force +# folder path) overwrite what the new build ships and delete nothing else, so a file the old version had +# and the new one dropped survives this step by construction. DarlingInstallDirectoryReport cannot see it +# either: it walks top-level DIRECTORIES, so a stale DLL in the root, or in viewer\ or runtimes\, is +# structurally invisible to it. +# +# That is dealt with AFTER the copy rather than here, against the manifest the previous upgrade wrote +# (#2529) - search this file for "the files this build no longer ships". After, because the difference +# being looked for is between what an earlier build put in this tree and what THIS payload ships, and the +# second half of that only exists once the copy has landed. +Note "Laying the new build over $InstallRoot ..." +$copied = $false +foreach ($attempt in 1, 2) { + try { + if ($sourceIsZip) { + Expand-Archive -LiteralPath $Source -DestinationPath $InstallRoot -Force + } + else { + Copy-Item -Path (Join-Path $Source '*') -Destination $InstallRoot -Recurse -Force + } + $copied = $true + break + } + catch { + # The transient one is a DLL an antivirus scan or a not-yet-exited process still holds, and a retry + # a moment later has worked more than once. Two attempts, then stop: a third would just be a longer + # way to arrive at the same half-written tree. + if ($attempt -eq 1) { + Warn "The copy failed ($($_.Exception.Message)). Retrying in 10 seconds - this step has lost to a transiently locked DLL before." + Start-Sleep -Seconds 10 + } + else { + Fail "The copy failed twice ($($_.Exception.Message)). The service is STOPPED and the install tree may be HALF WRITTEN - do not start it. Re-run this script with the same arguments: it will reuse the rollback backup it already took rather than replacing it, and finish the copy, which is the FIRST thing to try. To go back to the old version instead, note that a half-written tree needs BOTH halves: re-extract the PREVIOUS version's zip over $InstallRoot (that restores viewer\, wwwroot\ and runtimes\, which the backup does not hold), then copy the files from the newest _rollback_manual_* directory over the top. Restoring only the backup leaves old root binaries paired with partly-new subdirectories." + } + } +} + +if (-not $copied) { Fail "The copy did not complete." } +Good "New build in place." + +if ($configHashBefore) { + $configHashAfter = if (Test-Path -LiteralPath $configPath) { (Get-FileHash -LiteralPath $configPath -Algorithm SHA256).Hash } else { $null } + if ($configHashAfter -ne $configHashBefore) { + # The zip ships darling.sample.json and never darling.json, so this should be impossible - which is + # exactly why it is checked rather than trusted. A config replaced by a deploy is a monitoring host + # that comes back up watching nothing, and the newest rollback backup still holds the original. + Warn "darling.json CHANGED during the copy. The zip ships only darling.sample.json, so this should not happen. The original is in the newest _rollback_manual_* directory - compare them before starting the service." + } + else { + Good "darling.json is unchanged." + } +} + +# ============================ the files this build no longer ships (#2529) ============================ +# +# AFTER the copy - but only the half that has to be. The tree is now everything the new build says it +# should be plus whatever an earlier one left in it, so the DIFFERENCE, the existence check against the +# tree, and the removal all belong here. The previous build's manifest does NOT: it is read further up, +# before the copy, because a folder source can overwrite it. Review caught the two being conflated, and +# they are worth keeping apart in the reader's head as well as in the script. +# +# NOTHING HERE CALLS Fail, and that is a decision rather than an oversight. A monitoring host that is up +# with one stale DLL in it is in far better shape than one this script refused to start over a cleanup, so +# every failure below is a warning and a re-run. The service is still stopped at this point and starting it +# is what the operator came here for. + +$payload = Get-DarlingPayloadFiles $Source $sourceIsZip +$carryForward = @() + +if (-not $payload.Ok) { + Warn "Could not read the new build's own file list ($($payload.Reason)), so this run cannot tell which files an earlier build left behind. NOTHING was removed, and no manifest was written - the next upgrade will be in the same position until a run gets a source it can read." +} +elseif (-not $previousManifest.Ok) { + Note "Stale-file check skipped: $($previousManifest.Reason). Nothing was removed. This run writes the manifest, so the NEXT upgrade can answer the question." +} +else { + $candidates = Select-DarlingStaleFiles $previousManifest.Files $payload.Files + + # Only the ones actually on disk. A path an operator already deleted by hand is not news, and carrying + # it forward would keep it in the manifest, and in this report, for the rest of the install's life. + $stale = @($candidates | Where-Object { $file = Join-DarlingInstallPath $InstallRoot $_; $file -and (Test-Path -LiteralPath $file -PathType Leaf) }) + + if ($stale.Count -eq 0) { + Good "No files from the previous build were dropped by this one." + } + elseif (-not $RemoveStaleFiles) { + # Every path, uncapped. The service's layout report caps and summarises because it repeats on every + # start; this prints once, at the moment of the deploy, with the operator reading it - and the case + # that matters most is the forty-file one, where a count tells you nothing and the list tells you a + # whole framework directory was stranded. + Warn ("{0} file(s) in the install tree were shipped by an earlier build and are NOT in this one. Nothing was removed - re-run with -RemoveStaleFiles to delete them, or delete them by hand:" -f $stale.Count) + foreach ($relative in $stale) { Note " $relative" } + Note "Every path above came out of one of our own build payloads: this check reads the manifest THIS SCRIPT wrote on the previous upgrade, so it can only ever name files one of our zips laid down - never darling.json, never the credential blobs, never pg-runtime." + $carryForward = @($stale) + } + else { + Note ("Removing {0} file(s) shipped by an earlier build and not by this one:" -f $stale.Count) + $staleResult = Remove-DarlingStaleFiles $InstallRoot $stale $payload.Files + + Good ("Removed {0} stale file(s), reclaiming {1}." -f $staleResult.Removed.Count, (Format-DarlingBytes $staleResult.Reclaimed)) + + if ($staleResult.Emptied.Count -gt 0) { + Note ("Also removed {0} director(ies) that held nothing but those files: {1}" -f $staleResult.Emptied.Count, ($staleResult.Emptied -join ', ')) + } + + if ($staleResult.Refused.Count -gt 0) { + Warn ("{0} path(s) were REFUSED rather than removed, because they name something this procedure never deletes: {1}. Nothing needs doing about it - it is reported because a delete that quietly did not happen is worse than one that says so." -f $staleResult.Refused.Count, ($staleResult.Refused -join ', ')) + } + + if ($staleResult.Failures.Count -gt 0) { + Warn ("{0} stale file(s) could not be removed: {1}. Re-running the upgrade retries them." -f $staleResult.Failures.Count, ($staleResult.Failures -join ', ')) + } + + if ($staleResult.Removed.Count -gt 0) { + Note "A stale file in a SUBDIRECTORY was not in the rollback backup, which holds the install root's files only. The complete revert is unchanged: re-extract the previous version's zip over the install root, then copy the newest _rollback_manual_* directory's files over the top." + } + + # Whatever is still on disk stays NOMINATED. Without this the manifest written below would record + # only what this build ships, a file that was reported and not removed would drop out of the record, + # and no later upgrade would ever mention it again - a check that forgets, which is a check nobody + # can act on at their own pace. + $carryForward = @($staleResult.Failures) + @($staleResult.Refused) + } +} + +if ($payload.Ok) { + $manifestResult = Write-DarlingInstallManifest $manifestPath (@($payload.Files) + @($carryForward)) ([IO.Path]::GetFileName($Source)) ([datetime]::UtcNow) + + if ($manifestResult.Ok) { + Good ("Install manifest written ({0} file(s))." -f $manifestResult.Count) + } + else { + Warn "Could not write the install manifest ($($manifestResult.Reason)). The upgrade itself is fine and the service will start; the next upgrade will not be able to tell which files this build shipped, and will remove nothing." + } +} + +Note "Starting '$serviceName'..." +Start-Service -Name $serviceName +try { (Get-Service -Name $serviceName).WaitForStatus('Running', [TimeSpan]::FromMinutes(2)) } +catch { Fail "'$serviceName' did not reach Running within two minutes. The new build IS in place - read %ProgramData%\PerformanceMonitorDarling\logs before rolling back." } + +Good "Service is Running." +Note "" +Note "The install is NOT verified yet. 'Running' means the process started, not that it collects." +Note "In 10-15 minutes, against the store, confirm all of:" +Note " 1. MAX(version) FROM darling_schema_version equals this build's expected schema rung." +Note " 2. COUNT(DISTINCT server_id) FROM collect.collection_log over the last 15 minutes equals the fleet size." +Note " 3. COUNT(DISTINCT collector_name) over the same window is in the mid-30s, not single digits." +Note " 4. Any non-SUCCESS rows since the restart are READ, not just counted - YIELDED is the lock-timeout guard working; anything else is a finding." +Note " 5. If this build added a collector or a migration, one targeted probe that ITS table moved." diff --git a/Directory.Packages.props b/Directory.Packages.props index e6fc27315..df7daa601 100644 --- a/Directory.Packages.props +++ b/Directory.Packages.props @@ -1,32 +1,34 @@ - + - - true - - - - - - - - - - - - - - - - - - - - - - - - + tools/CompactionRepro does, to keep reproducing against its original DuckDB). --> + + true + + + + + + + + + + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/Lite.Tests/AgAlertEvaluatorTests.cs b/Lite.Tests/AgAlertEvaluatorTests.cs index aee0c8154..f6f7c2b56 100644 --- a/Lite.Tests/AgAlertEvaluatorTests.cs +++ b/Lite.Tests/AgAlertEvaluatorTests.cs @@ -101,6 +101,137 @@ public void Disconnected_FiresOnTheEdge_DoesNotRepeat_ThenReconnectResolves() Assert.True(back.IsResolution); } + /* ---------------- #2426: the disconnect re-fire ---------------- */ + + private static readonly DateTime Noon = new(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc); + private static readonly TimeSpan Refire = TimeSpan.FromMinutes(10); + + [Fact] + public void DisconnectRefire_ReAnnouncesUnderTheSameMetricName_AndSaysThatItIsStill() + { + var now = Noon; + var e = new AgAlertEvaluator(() => now); + + e.EvaluateReplicas(ServerId, new[] { Replica(connected: "CONNECTED") }, Refire); + + var lost = Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + Assert.Equal(AgAlertPolicy.ReplicaDisconnectedMetric, lost.MetricName); + Assert.DoesNotContain("STILL", lost.DetailText, StringComparison.Ordinal); + e.NoteDelivered(lost); + + /* Inside the window there is nothing new to say. */ + now = now.AddMinutes(5); + Assert.Empty(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + + /* Past it: the SAME metric name, because webhook automation keyed on it is what a re-fire exists to + re-trigger — and a detail that tells an operator reading the history this is hour two, not a + second outage. */ + now = now.AddMinutes(6); + var again = Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + Assert.Equal(AgAlertPolicy.ReplicaDisconnectedMetric, again.MetricName); + Assert.False(again.IsResolution); + Assert.Contains("STILL DISCONNECTED", again.DetailText, StringComparison.Ordinal); + Assert.Contains("re-alerting every 10 min", again.DetailText, StringComparison.Ordinal); + } + + [Fact] + public void DisconnectRefire_TheWindowOpensOnDeliveryOnly_NotOnTheDecision() + { + var now = Noon; + var e = new AgAlertEvaluator(() => now); + + e.EvaluateReplicas(ServerId, new[] { Replica(connected: "CONNECTED") }, Refire); + + /* Evaluated and NOT delivered — exactly what MainWindow does every sweep while a server is + acknowledged or silenced. */ + Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + + /* So nothing consumed the window, and the next sweep still has something to say. Stamping on the + decision instead would spend window after window on alerts nobody received, and the operator + would come back from an acknowledgement to silence. */ + now = now.AddMinutes(1); + var again = Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + Assert.Contains("STILL DISCONNECTED", again.DetailText, StringComparison.Ordinal); + + /* Delivered this time, so the clock finally starts. */ + e.NoteDelivered(again); + now = now.AddMinutes(1); + Assert.Empty(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + } + + [Fact] + public void DisconnectRefire_ReconnectClearsTheClock_SoTheNextOutageIsNotHeldQuietByTheLastOnes() + { + var now = Noon; + var e = new AgAlertEvaluator(() => now); + + e.EvaluateReplicas(ServerId, new[] { Replica(connected: "CONNECTED") }, Refire); + e.NoteDelivered(Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire))); + + now = now.AddMinutes(1); + Assert.True(Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "CONNECTED") }, Refire)).IsResolution); + + /* A second outage a minute later whose opening edge is suppressed. With the clock cleared the next + sweep re-announces; with the first episode's stamp still on it, this replica would sit silent for + the remaining eight minutes of a window that belongs to an outage already over. */ + now = now.AddMinutes(1); + Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + + now = now.AddMinutes(1); + var again = Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + Assert.Contains("STILL DISCONNECTED", again.DetailText, StringComparison.Ordinal); + } + + [Fact] + public void DisconnectRefire_EachReplicaKeepsItsOwnClock() + { + var now = Noon; + var e = new AgAlertEvaluator(() => now); + + e.EvaluateReplicas(ServerId, new[] + { + Replica(connected: "CONNECTED", replica: "NODE2"), + Replica(connected: "CONNECTED", replica: "NODE3"), + }, Refire); + + var lost = e.EvaluateReplicas(ServerId, new[] + { + Replica(connected: "DISCONNECTED", replica: "NODE2"), + Replica(connected: "DISCONNECTED", replica: "NODE3"), + }, Refire); + Assert.Equal(2, lost.Count); + + /* Only NODE2's alert was delivered; NODE3's window is still open. */ + Assert.Contains("NODE2", lost[0].DetailText, StringComparison.Ordinal); + e.NoteDelivered(lost[0]); + + now = now.AddMinutes(1); + var again = Assert.Single(e.EvaluateReplicas(ServerId, new[] + { + Replica(connected: "DISCONNECTED", replica: "NODE2"), + Replica(connected: "DISCONNECTED", replica: "NODE3"), + }, Refire)); + Assert.Contains("NODE3", again.DetailText, StringComparison.Ordinal); + } + + [Fact] + public void DisconnectRefire_Forget_DropsTheClockWithTheRestOfTheState() + { + var now = Noon; + var e = new AgAlertEvaluator(() => now); + + e.EvaluateReplicas(ServerId, new[] { Replica(connected: "CONNECTED") }, Refire); + e.NoteDelivered(Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire))); + + e.Forget(ServerId); + + /* Re-added a minute later: a first sighting again, and with re-fire on that announces rather than + baselining silently. A clock that survived Forget would hold it quiet for nine more minutes. */ + now = now.AddMinutes(1); + var again = Assert.Single(e.EvaluateReplicas(ServerId, new[] { Replica(connected: "DISCONNECTED") }, Refire)); + Assert.Contains("STILL DISCONNECTED", again.DetailText, StringComparison.Ordinal); + } + /* ---------------- suspended ---------------- */ [Fact] diff --git a/Lite.Tests/AgAlertPolicyTests.cs b/Lite.Tests/AgAlertPolicyTests.cs index b70fb4e21..d52f82dc5 100644 --- a/Lite.Tests/AgAlertPolicyTests.cs +++ b/Lite.Tests/AgAlertPolicyTests.cs @@ -6,6 +6,7 @@ * Licensed under the MIT License. See LICENSE file in the project root for full license information. */ +using System; using PerformanceMonitor.Common; using Xunit; @@ -81,6 +82,102 @@ public void DecideConnection_AlreadyDisconnectedAtFirstSighting_StaysSilent_ButI Assert.Equal(AgConnectionDecision.Reconnected, AgAlertPolicy.DecideConnection("DISCONNECTED", "CONNECTED")); } + /* ---------------- #2426: the disconnect re-fire ---------------- */ + + private static readonly DateTime Noon = new(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc); + + [Theory] + [InlineData("CONNECTED", "DISCONNECTED", AgConnectionDecision.Disconnected)] + [InlineData("DISCONNECTED", "CONNECTED", AgConnectionDecision.Reconnected)] + [InlineData("DISCONNECTED", "DISCONNECTED", AgConnectionDecision.None)] + [InlineData("CONNECTED", "CONNECTED", AgConnectionDecision.None)] + [InlineData(null, "DISCONNECTED", AgConnectionDecision.None)] + [InlineData("CONNECTED", null, AgConnectionDecision.None)] + public void DecideConnection_RefireOff_IsTheEdgeOnlyOverloadExactly( + string? previous, string? current, AgConnectionDecision expected) + { + /* The shipped default, and the whole matrix rather than one case: null, zero and a negative all + mean OFF, and off has to be byte-for-byte the two-argument behavior or an upgrade would start + re-alerting on a knob nobody set. */ + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, null, null, Noon)); + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, TimeSpan.Zero, null, Noon)); + Assert.Equal(expected, AgAlertPolicy.DecideConnection(previous, current, TimeSpan.FromMinutes(-5), null, Noon)); + } + + [Fact] + public void DecideConnection_StillDisconnected_WaitsOutTheWindow_ThenSaysItAgain() + { + var refire = TimeSpan.FromMinutes(10); + + /* Announced at noon: inside the window there is nothing new to say. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddMinutes(9))); + + /* At the boundary, and still hours later — the point of the knob is that a week-long outage does + not read like a blip. */ + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddMinutes(10))); + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, Noon, Noon.AddHours(8))); + } + + [Fact] + public void DecideConnection_ARealEdgeOutranksARefire() + { + var refire = TimeSpan.FromMinutes(10); + + /* The edge that OPENS the outage announces as Disconnected even though the window is trivially due + on it, or the caller would have two reasons to announce the same sweep. */ + Assert.Equal( + AgConnectionDecision.Disconnected, + AgAlertPolicy.DecideConnection("CONNECTED", "DISCONNECTED", refire, null, Noon)); + + /* And a recovery is a recovery whatever the clock says. */ + Assert.Equal( + AgConnectionDecision.Reconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "CONNECTED", refire, Noon.AddHours(-9), Noon)); + } + + [Fact] + public void DecideConnection_NoStampIsDueNow_SoARestartMidOutageStillReAnnounces() + { + var refire = TimeSpan.FromMinutes(10); + + /* Both apps hold this edge state in memory, so a restart during a week-long outage sees a replica + already DISCONNECTED with no record of it ever having been announced. Rule 1's silent baseline + would make that silence permanent, which is the exact failure the knob exists to prevent — so + with re-fire ON, and only then, a first sighting announces. */ + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection(null, "DISCONNECTED", refire, null, Noon)); + Assert.Equal( + AgConnectionDecision.StillDisconnected, + AgAlertPolicy.DecideConnection("DISCONNECTED", "DISCONNECTED", refire, null, Noon)); + + /* With it off, rule 1 governs unchanged. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection(null, "DISCONNECTED", null, null, Noon)); + } + + [Fact] + public void DecideConnection_ARefireStillNeedsAnExactDisconnected() + { + var refire = TimeSpan.FromMinutes(10); + + /* Same rule the edge follows, and it matters more here: a re-fire pages repeatedly, so a state + string the product never learned to interpret must not become a standing page. */ + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("SOMETHING_NEW", "SOMETHING_NEW", refire, null, Noon)); + Assert.Equal( + AgConnectionDecision.None, + AgAlertPolicy.DecideConnection("CONNECTED", null, refire, null, Noon)); + } + /* ---------------- suspension ---------------- */ [Theory] diff --git a/Lite.Tests/AgentStatusCollectorDefinitionTests.cs b/Lite.Tests/AgentStatusCollectorDefinitionTests.cs index 4cc16cf60..47a4f6500 100644 --- a/Lite.Tests/AgentStatusCollectorDefinitionTests.cs +++ b/Lite.Tests/AgentStatusCollectorDefinitionTests.cs @@ -71,9 +71,12 @@ public void AppliesTo_CollectsEverywhereExceptAzureSqlDbRdsAndNoMsdb() empty snapshot that would drive a false "Agent Not Running" alert. */ Assert.False(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAzureSqlDb = true })); Assert.False(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAwsRds = true })); - /* No msdb access → the sysjobschedules next-run decode can't run; skip (gate collapsed from - IsCollectorSupported into the shared AppliesTo — #1: one authoritative gate surface). */ - Assert.False(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); + /* NOT gated on msdb access (#2559). It is a GRANT rather than an engine capability, and it was + probed once and cached for the connection's life - so running the GRANT we advise did nothing + until a restart. This now attempts and fails into PERMISSIONS, which error 916 already maps to, + and CollectorHealthClassifier bands a never-permitted collector as NO_PERMISSIONS ahead of + FAILING, so it does not read as broken. */ + Assert.True(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); /* Managed Instance has Agent; on-prem collects too. Bare target = msdb assumed accessible. */ Assert.True(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAzureManagedInstance = true })); Assert.True(AgentStatusCollector.Instance.AppliesTo(new CollectorTargetInfo())); diff --git a/Lite.Tests/AlertContextBuildersTests.cs b/Lite.Tests/AlertContextBuildersTests.cs index 823a79511..a48170211 100644 --- a/Lite.Tests/AlertContextBuildersTests.cs +++ b/Lite.Tests/AlertContextBuildersTests.cs @@ -276,6 +276,10 @@ public void BuildTempDbSpaceContext_RendersAllFields_WithTopConsumer() { ("Total Reserved", $"{800d:F0} MB"), ("Unallocated", $"{200d:F0} MB"), + /* #2515: no MaxSizeMb on this fixture, so the ceiling was never measured and the + percentage above is still against the allocation — which the detail now says out loud + rather than leaving the reader to assume a denominator. */ + ("Max Size", "Unknown"), ("User Objects", $"{500d:F0} MB"), ("Internal Objects", $"{250d:F0} MB"), ("Version Store", $"{50d:F0} MB"), @@ -284,6 +288,27 @@ public void BuildTempDbSpaceContext_RendersAllFields_WithTopConsumer() item.Fields); } + /// + /// #2515: the three states of the ceiling render as three different words, because they are three + /// different facts. "Unlimited" and "Unknown" take the same denominator but they do not mean the same + /// thing, and printing either as a number would claim a measurement that was never taken. + /// + [Theory] + [InlineData(65536d, "65536 MB")] + [InlineData(-1d, "Unlimited")] + [InlineData(0d, "Unknown")] + public void BuildTempDbSpaceContext_RendersTheCeilingsThreeStates(double maxSizeMb, string expected) + { + var context = AlertContextBuilders.BuildTempDbSpaceContext(new TempDbSpaceInfo + { + TotalReservedMb = 800, + UnallocatedMb = 200, + MaxSizeMb = maxSizeMb + }); + + Assert.Equal(("Max Size", expected), context!.Details[0].Fields[2]); + } + [Fact] public void BuildTempDbSpaceContext_NoTopConsumer_RendersNone() { @@ -294,7 +319,7 @@ public void BuildTempDbSpaceContext_NoTopConsumer_RendersNone() TopConsumerSessionId = 0 }); - Assert.Equal(("Top Consumer", "None"), context!.Details[0].Fields[5]); + Assert.Equal(("Top Consumer", "None"), context!.Details[0].Fields[6]); } /* ---------------- anomalous jobs ---------------- */ diff --git a/Lite.Tests/AlertSettingsControlWiringTests.cs b/Lite.Tests/AlertSettingsControlWiringTests.cs index 4896e964d..571ac67ea 100644 --- a/Lite.Tests/AlertSettingsControlWiringTests.cs +++ b/Lite.Tests/AlertSettingsControlWiringTests.cs @@ -85,9 +85,57 @@ private static HashSet ThresholdBoxesIn(string source, string methodName .ToHashSet(StringComparer.Ordinal); } + /// + /// Lite's other half of the same "adding a knob is more than one edit" problem, and the half the test + /// above cannot see. Lite persists to settings.json rather than a store, and SaveAlertSettings + /// does both jobs in one method: it copies each control into an App.Alert* static, then writes + /// those statics into the root[...] JSON document. Forgetting the second line is silent and + /// survives every manual test — the setting takes effect immediately and only reverts on the NEXT + /// launch, by which point nobody connects the two. #2391 added four file-growth controls to this + /// method; this pins that all four (and every sibling) actually reach disk. + /// + /// Note this does NOT mean loader/writer key symmetry: the writer parses the existing + /// settings.json into root and mutates it, so hand-edited keys it never touches are preserved + /// on save. The invariant is narrower — whatever the UI can CHANGE, the UI must WRITE. + /// + [Fact] + public void LiteSaveAlertSettings_PersistsEveryStaticItAssigns() + { + var source = File.ReadAllText(FindRepoFile(Path.Combine("Lite", "Windows", "SettingsWindow.xaml.cs"))); + var body = MethodBody(source, "SaveAlertSettings", "Lite"); + + /* "App.Foo =" but not "App.Foo ==" — the assignments the Save button makes. */ + var assigned = Regex.Matches(body, @"App\.(\w+)\s*=(?!=)") + .Select(m => m.Groups[1].Value) + .ToHashSet(StringComparer.Ordinal); + var persisted = Regex.Matches(body, @"root\[""[^""]+""\]\s*=\s*App\.(\w+)") + .Select(m => m.Groups[1].Value) + .ToHashSet(StringComparer.Ordinal); + + Assert.NotEmpty(assigned); + + /* AlertExcludedDatabases is the one legitimate exception: it is a collection, so it reaches the + document as a JsonArray built a few lines earlier rather than as a bare "= App.X" assignment. */ + var unpersisted = assigned.Except(persisted) + .Except(new[] { "AlertExcludedDatabases" }) + .OrderBy(n => n, StringComparer.Ordinal) + .ToList(); + + Assert.True(unpersisted.Count == 0, + "Lite: SaveAlertSettings copies these controls into App statics but never writes them to " + + "settings.json, so the setting applies for this session and silently reverts on next launch: " + + string.Join(", ", unpersisted)); + } + private static string MethodBody(string source, string methodName, string app) { + /* Not every anchor returns void — SaveAlertSettings returns bool to report whether it succeeded. */ var signature = source.IndexOf("void " + methodName, StringComparison.Ordinal); + if (signature < 0) + { + signature = source.IndexOf("bool " + methodName, StringComparison.Ordinal); + } + Assert.True(signature >= 0, $"{app}: no method named {methodName} — this guard's anchor moved and it is testing nothing."); var open = source.IndexOf('{', signature); diff --git a/Lite.Tests/AnalysisAsOfAnchorTests.cs b/Lite.Tests/AnalysisAsOfAnchorTests.cs new file mode 100644 index 000000000..d3b3f0fda --- /dev/null +++ b/Lite.Tests/AnalysisAsOfAnchorTests.cs @@ -0,0 +1,354 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitor.Analysis; +using PerformanceMonitorLite.Analysis; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2506: the analysis family's as_of anchor on Lite. #2495 anchored the read surface and left +/// these out because their window is built inside rather than in the tool, +/// and because analyze_server PERSISTS what it finds. +/// +/// Three claims, each of which failed before this change: a past window is reachable at all; the +/// anchor survives the trip from the tool into the engine (the seam #2495's review found eight broken +/// tools on, one level up); and an anchored analysis writes NOTHING, which is what makes anchoring the +/// persisting tool safe rather than merely possible. +/// +public sealed class AnalysisAsOfAnchorTests : IClassFixture, IDisposable +{ + private const string HistoricHash = "an2506-lite-historic"; + private const string RecentHash = "an2506-lite-recent"; + private const string PastWait = "AN2506_LITE_PAST_WAIT"; + + private readonly string _tempDir; + private readonly DuckDbInitializer _duckDb; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private long _nextId = -8_260_600; + private DuckDBConnection? _seedConn; + + public AnalysisAsOfAnchorTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _tempDir = Path.Combine(Path.GetTempPath(), "AnalysisAsOfTests_" + Guid.NewGuid().ToString("N")[..8]); + var configDir = Path.Combine(_tempDir, "config"); + Directory.CreateDirectory(configDir); + + /* Windows auth so AddServer never touches the credential store — no DPAPI side effects. */ + _serverManager = new ServerManager(configDir); + var server = new ServerConnection { ServerName = "TestServer", DisplayName = "TestServer" }; + _serverManager.AddServer(server); + + /* The id the TOOLS resolve to. Seeding under any other one would make every assertion below + pass or fail for a reason that has nothing to do with the anchor. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { if (Directory.Exists(_tempDir)) Directory.Delete(_tempDir, recursive: true); } + catch { /* best-effort cleanup */ } + } + + /// + /// The findings read windows on ANALYSIS TIME. A pass that ran 30 hours ago is outside every default + /// window (the tool's default is 24) and inside an anchored one — and, in the other direction, the + /// run from an hour ago is inside the default and outside the anchored one. Both directions are + /// asserted because only the second can fail if the upper bound is missing, which is exactly the + /// half-open-window mistake this parameter makes so easy. + /// + [Fact] + public async Task TheFindingsRead_SeesAPastRun_AndTheDefaultAnchorDoesNot() + { + var historicRun = DateTime.UtcNow.AddHours(-30); + await PlantFindingAsync(historicRun, HistoricHash); + await PlantFindingAsync(DateTime.UtcNow.AddHours(-1), RecentHash); + + var service = CreateTestService(); + + var defaultWindow = await McpAnalysisTools.GetAnalysisFindings(service, _serverManager); + Assert.Contains(RecentHash, defaultWindow, StringComparison.Ordinal); + Assert.DoesNotContain(HistoricHash, defaultWindow, StringComparison.Ordinal); + + var anchored = await McpAnalysisTools.GetAnalysisFindings( + service, _serverManager, null, 4, false, historicRun.AddMinutes(30).ToString("o")); + Assert.Contains(HistoricHash, anchored, StringComparison.Ordinal); + Assert.DoesNotContain(RecentHash, anchored, StringComparison.Ordinal); + Assert.Equal(1, JsonDocument.Parse(anchored).RootElement.GetProperty("finding_count").GetInt32()); + } + + /// + /// The UNANCHORED findings read keeps its half-open window, and that is a decision rather than an + /// omission: bounding it at "now" broke a live test on this change's first attempt, and the + /// mechanism behind that is real. analysis_time is stamped by the WRITER and filtered by the + /// READER, so a default read bounded at the reader's clock drops a run written a moment earlier the + /// day those two clocks stop being the same one — a findings read that "sometimes misses the + /// analysis that just finished", with nothing in it to point at a clock. A row stamped ahead of now + /// is the cheap, deterministic stand-in for that skew. + /// + [Fact] + public async Task TheUnanchoredFindingsRead_StillHasNoUpperBound_AndTheAnchoredOneDoes() + { + await PlantFindingAsync(DateTime.UtcNow.AddMinutes(30), RecentHash); + await PlantFindingAsync(DateTime.UtcNow.AddHours(-30), HistoricHash); + + var service = CreateTestService(); + + Assert.Contains( + RecentHash, + await McpAnalysisTools.GetAnalysisFindings(service, _serverManager), + StringComparison.Ordinal); + + /* Present exactly where it is needed and absent exactly where it would do harm. */ + Assert.DoesNotContain( + RecentHash, + await McpAnalysisTools.GetAnalysisFindings( + service, _serverManager, null, 4, false, DateTime.UtcNow.AddHours(-29.5).ToString("o")), + StringComparison.Ordinal); + } + + /// + /// The tool-to-engine seam. get_analysis_facts hands the anchor to + /// , which is where the window is actually + /// built — a tool that resolved the anchor and then called the engine without it would validate, + /// refuse bad input correctly, and score the present. + /// + [Fact] + public async Task TheFactsRead_ScoresAPastWindow_AndTheDefaultAnchorFindsNothingThere() + { + /* 30 hours ago, i.e. outside every default window on this surface (the widest default is 24). */ + var incident = DateTime.UtcNow.AddHours(-30); + await PlantWaitAsync(incident, 900_000L); + await PlantWaitAsync(incident.AddMinutes(10), 800_000L); + + var service = CreateTestService(); + + var anchored = await McpAnalysisTools.GetAnalysisFacts( + service, _serverManager, null, 4, null, 0, incident.AddMinutes(30).ToString("o")); + Assert.Contains(PastWait, anchored, StringComparison.Ordinal); + + var defaultWindow = await McpAnalysisTools.GetAnalysisFacts(service, _serverManager); + Assert.Equal("unavailable", StatusOf(defaultWindow)); + } + + /// + /// The persistence decision, A/B'd on the SAME window so nothing but the anchor can explain the + /// difference. An anchored pass returns its findings in full and writes none of them; the identical + /// window unanchored returns the same findings and writes them. + /// + /// Run against the engine rather than the tool because the rule lives in the engine on purpose: + /// there is no legitimate caller for "anchored AND persist", so there is no way to ask for it, and a + /// future caller that never read the argument cannot get it wrong. + /// + [Fact] + public async Task AnAnchoredPass_ReturnsItsFindings_AndWritesNoneOfThem() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.SeedMemoryStarvedServerAsync(); + + var service = CreateTestService(); + + var anchored = TestDataSeeder.CreateTestContext(); + anchored.AsOfUtc = anchored.TimeRangeEnd; + + var exploratory = await service.AnalyzeAsync(anchored); + Assert.NotEmpty(exploratory); + Assert.Empty(await service.GetRecentFindingsAsync(TestDataSeeder.TestServerId)); + + /* The control, and the reason the assertion above is not vacuous: the same window with no anchor + produces findings AND rows. Without this, "nothing was written" would also be the reading of a + fixture that had nothing to write. */ + var persisting = await service.AnalyzeAsync(TestDataSeeder.CreateTestContext()); + Assert.NotEmpty(persisting); + Assert.NotEmpty(await service.GetRecentFindingsAsync(TestDataSeeder.TestServerId)); + } + + /// + /// And the tool says so. A payload that quietly omitted the difference would leave an agent telling + /// someone to "check the persisted findings" for a run that never wrote any. + /// + [Fact] + public async Task TheAnalyzeTool_DisclosesThatAnAnchoredRunWasNotPersisted() + { + await PlantWaitAsync(DateTime.UtcNow.AddHours(-30), 900_000L); + await PlantWaitAsync(DateTime.UtcNow.AddHours(-29), 800_000L); + + var service = CreateTestService(); + + var anchored = await McpAnalysisTools.AnalyzeServer( + service, _serverManager, null, 4, DateTime.UtcNow.AddHours(-29).ToString("o")); + var (anchoredPersisted, anchoredNote) = PersistenceOf(anchored); + Assert.False(anchoredPersisted); + Assert.Contains("NOT written to the store", anchoredNote!, StringComparison.Ordinal); + + Assert.Equal(0, await CountFindingsAsync()); + + /* Unanchored, the same tool reports the ordinary answer and says nothing extra. */ + var (defaultPersisted, defaultNote) = PersistenceOf(await McpAnalysisTools.AnalyzeServer(service, _serverManager)); + Assert.True(defaultPersisted); + Assert.Null(defaultNote); + } + + /// + /// A bad anchor is REFUSED on the persisting tool too. Falling back to "now" here would not merely + /// answer a different question — it would run a real, persisting analysis the caller did not ask for. + /// + [Fact] + public async Task TheAnalyzeTool_RefusesAnUnusableAnchor_RatherThanAnalysingThePresent() + { + var service = CreateTestService(); + + Assert.StartsWith( + "Invalid as_of", + await McpAnalysisTools.AnalyzeServer(service, _serverManager, null, 4, "last tuesday"), + StringComparison.Ordinal); + + Assert.Contains( + "future", + await McpAnalysisTools.AnalyzeServer(service, _serverManager, null, 4, DateTime.UtcNow.AddDays(1).ToString("o")), + StringComparison.Ordinal); + + Assert.Equal(0, await CountFindingsAsync()); + } + + /// + /// compare_analysis moves BOTH windows, because baseline_hours_back has always been + /// measured from the comparison window's end. Moving only the comparison end would silently change + /// what the two windows are relative to each other. + /// + [Fact] + public async Task CompareAnalysis_HangsBothWindowsOffTheAnchor() + { + var service = CreateTestService(); + var anchor = DateTime.UtcNow.AddHours(-30); + await PlantWaitAsync(anchor.AddMinutes(-30), 900_000L); + await PlantWaitAsync(anchor.AddHours(-28), 500_000L); + + var compared = await McpAnalysisTools.CompareAnalysis( + service, _serverManager, null, 4, 28, anchor.ToString("o")); + + Assert.Contains(anchor.ToString("o"), compared, StringComparison.Ordinal); + Assert.Contains(anchor.AddHours(-28).ToString("o"), compared, StringComparison.Ordinal); + + /* The control: unanchored, both windows hang off now and neither instant appears. */ + var comparedNow = await McpAnalysisTools.CompareAnalysis(service, _serverManager, null, 4, 28); + Assert.DoesNotContain(anchor.ToString("o"), comparedNow, StringComparison.Ordinal); + } + + // ── helpers ── + + private AnalysisService CreateTestService() => new(_duckDb) { MinimumDataHours = 0 }; + + private static string StatusOf(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("status").GetString()!; + } + + /// + /// Reads persisted / persistence_note from either shape analyze_server can + /// return. A data-bearing result carries them at the root; a miss carries them under hints, + /// beside the analysis_time that already lived there — so the test asserts the CLAIM rather + /// than a path, and does not quietly start passing if the run happens to find nothing. + /// + private static (bool Persisted, string? Note) PersistenceOf(string json) + { + using var doc = JsonDocument.Parse(json); + var root = doc.RootElement; + var carrier = root.TryGetProperty("persisted", out _) ? root : root.GetProperty("hints"); + + var note = carrier.GetProperty("persistence_note"); + return (carrier.GetProperty("persisted").GetBoolean(), + note.ValueKind == JsonValueKind.Null ? null : note.GetString()); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task CountFindingsAsync() + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + using var cmd = conn.CreateCommand(); + cmd.CommandText = "SELECT COUNT(*) FROM analysis_findings WHERE server_id = $1"; + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + return Convert.ToInt32(await cmd.ExecuteScalarAsync() ?? 0); + } + + private async Task PlantWaitAsync(DateTime at, long deltaWaitMs) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + using var cmd = conn.CreateCommand(); + cmd.CommandText = @" +INSERT INTO wait_stats + (collection_id, collection_time, server_id, server_name, wait_type, + waiting_tasks_count, wait_time_ms, signal_wait_time_ms, + delta_waiting_tasks, delta_wait_time_ms, delta_signal_wait_time_ms) +VALUES ($1, $2, $3, 'TestServer', $4, 5000, $5, 0, 5000, $5, 0)"; + void P(object v) => cmd.Parameters.Add(new DuckDBParameter { Value = v }); + P(_nextId--); + P(at); + P(_serverId); + P(PastWait); + P(deltaWaitMs); + await cmd.ExecuteNonQueryAsync(); + } + + /// + /// One finding row as a scheduled pass would have left it. Planted rather than produced, because + /// producing it would stamp analysis_time with NOW — the very thing an anchored run refuses. + /// + private async Task PlantFindingAsync(DateTime analysisTime, string storyPathHash) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + using var cmd = conn.CreateCommand(); + cmd.CommandText = @" +INSERT INTO analysis_findings + (finding_id, analysis_time, server_id, server_name, database_name, + time_range_start, time_range_end, severity, confidence, category, + story_path, story_path_hash, story_text, + root_fact_key, root_fact_value, leaf_fact_key, leaf_fact_value, fact_count, incident_id, + remediation_action_json, drill_down_json) +VALUES ($1, $2, $3, 'TestServer', NULL, $4, $5, 0.9, 0.8, 'waits', + 'AN2506_LITE', $6, 'planted for the #2506 anchored findings read', + 'AN2506_LITE', 1, NULL, NULL, 1, NULL, NULL, NULL)"; + void P(object v) => cmd.Parameters.Add(new DuckDBParameter { Value = v }); + P(_nextId--); + P(analysisTime); + P(_serverId); + P(analysisTime.AddHours(-4)); + P(analysisTime); + P(storyPathHash); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/AnalysisBudgetTests.cs b/Lite.Tests/AnalysisBudgetTests.cs new file mode 100644 index 000000000..66935f4d8 --- /dev/null +++ b/Lite.Tests/AnalysisBudgetTests.cs @@ -0,0 +1,176 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.RegularExpressions; +using System.Threading; +using System.Threading.Tasks; +using PerformanceMonitorLite.Analysis; +using PerformanceMonitorLite.Database; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2412: App.AnalysisTimeoutSeconds cancelled nothing. The scheduler raced +/// AnalyzeAsync against a Task.Delay and, on losing, logged one warning and moved +/// on while the pass kept running — so the setting bounded how long the LOOP waited and nothing +/// about the analysis itself. +/// +/// The consequence was worse than the mislabelled knob. The in-flight guard is released only +/// on true completion (correctly — that is what stops a hung server piling up tasks), so a pass +/// that never completed left its marker set for the life of the process and that server was +/// skipped by continue on every later cycle, silently, with the single original warning the +/// only trace it ever left. +/// +/// Two layers guard the repair. The behavioural pair proves a pass carrying a cancelled +/// budget ABANDONS rather than running to completion — and abandons for that reason rather than +/// being turned away by the 24-hour data gate, which would make the assertion vacuous. The source +/// pin holds the three properties of the scheduler that no test in this suite can reach, because +/// CollectionBackgroundService needs a whole host to instantiate. +/// +public sealed class AnalysisBudgetTests : IClassFixture +{ + private readonly DuckDbInitializer _duckDb; + + public AnalysisBudgetTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + } + + /// + /// The discriminator is LastAnalysisTime. Every path that RUNS — a completed pass, and + /// the insufficient-data gate — stamps it. Only the abandon path returns without stamping, so + /// a null there cannot be produced by a pass that merely found nothing to say. + /// + /// The control runs first, on the same seed and the same context, and is what stops this + /// passing vacuously: it proves the pipeline reaches the end of a pass over this data, so the + /// null that follows is the cancellation and not a broken fixture. + /// + [Fact] + public async Task AnalyzeAsync_WithACancelledBudget_AbandonsInsteadOfCompleting() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.SeedCleanServerAsync(); + + var control = new AnalysisService(_duckDb) { MinimumDataHours = 0 }; + await control.AnalyzeAsync(TestDataSeeder.CreateTestContext()); + Assert.NotNull(control.LastAnalysisTime); + + var service = new AnalysisService(_duckDb) { MinimumDataHours = 0 }; + using var cts = new CancellationTokenSource(); + cts.Cancel(); + + var context = TestDataSeeder.CreateTestContext(); + context.CancellationToken = cts.Token; + + var findings = await service.AnalyzeAsync(context); + + Assert.Empty(findings); + Assert.Null(service.InsufficientDataMessage); + Assert.Null(service.LastAnalysisTime); + Assert.False(service.IsAnalyzing); + } + + /// + /// The scheduler calls the four-argument overload, so the token has to survive the hop from + /// that signature onto the context the pipeline actually reads. Drop that one assignment and + /// this pass would run the whole pipeline over an empty window and stamp + /// LastAnalysisTime — which is exactly what the assertion below refuses. + /// + [Fact] + public async Task AnalyzeAsync_CarriesTheCallersBudgetOntoTheContext() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.SeedCleanServerAsync(); + + var service = new AnalysisService(_duckDb) { MinimumDataHours = 0 }; + using var cts = new CancellationTokenSource(); + cts.Cancel(); + + var findings = await service.AnalyzeAsync( + TestDataSeeder.TestServerId, TestDataSeeder.TestServerName, hoursBack: 4, cts.Token); + + Assert.Empty(findings); + Assert.Null(service.LastAnalysisTime); + } + + /// + /// Three properties, each of which the fix is worthless without. + /// + /// The pass must leave the loop's thread. DuckDB.NET implements no async execution, so + /// every read completes synchronously on the calling thread and an AnalyzeAsync invoked + /// inline ran to completion BEFORE the timeout race was reached — the race could never fire for + /// the phase that would actually be slow. + /// + /// Something must raise the cancellation at the budget, or the token is decoration. + /// + /// And the skip branch must report. That bare continue is the whole defect: it is + /// the line that made a permanently wedged server invisible. + /// + [Fact] + public void ScheduledAnalysis_OffloadsThePass_ArmsTheBudget_AndReportsAStuckServer() + { + /* Line endings are normalised because the assertions below span lines and the working + copy is CRLF on Windows and LF elsewhere. */ + var source = File.ReadAllText( + Path.Combine(FindRepoDirectory(Path.Combine("Lite", "Services")), "CollectionBackgroundService.cs")) + .Replace("\r\n", "\n", StringComparison.Ordinal); + + /* The offload and the token have to be on the SAME call — a Task.Run somewhere in the file + and an AnalyzeAsync somewhere else would satisfy two independent Contains checks while + leaving the pass running inline on the loop's thread. */ + Assert.Matches( + new Regex(@"Task\.Run\(\s*\(\)\s*=>\s*analysisService\.AnalyzeAsync\(serverId, serverName, hoursBack: 4, cts\.Token\)"), + source); + + Assert.Contains( + "passCts.CancelAfter(timeout)", + source, + StringComparison.Ordinal); + + Assert.Contains( + "ReportStuckAnalysis(serverId, serverName, timeout);", + source, + StringComparison.Ordinal); + + /* The reporting has to back off. Scheduled analysis runs on a 30-minute default cadence, so + a fixed repeat would be either slower than the cadence or one line per cycle forever. */ + Assert.Contains( + "StuckAnalysisMaxBackoffDoublings", + source, + StringComparison.Ordinal); + + /* And the shutdown hold has to wait on an UNCANCELLED token. Offloading the pass onto the + pool is what makes the hold necessary at all — the store work can now still be in flight + when the loop is told to stop — and handing the already-fired stopping token to the wait + would collapse it instantly, which looks identical to having waited. */ + Assert.Contains( + "await analyzeTask.WaitAsync(AnalysisUnwindGrace, CancellationToken.None);", + source, + StringComparison.Ordinal); + } + + private static string FindRepoDirectory(string relativePath) + { + var dir = AppContext.BaseDirectory; + for (int i = 0; i < 8 && dir is not null; i++) + { + var candidate = Path.Combine(dir, relativePath); + if (Directory.Exists(candidate)) + { + return candidate; + } + dir = Path.GetDirectoryName(dir); + } + + throw new DirectoryNotFoundException($"Could not locate {relativePath} above {AppContext.BaseDirectory}"); + } +} diff --git a/Lite.Tests/AnalysisPassTokenThreadingTests.cs b/Lite.Tests/AnalysisPassTokenThreadingTests.cs new file mode 100644 index 000000000..574b94173 --- /dev/null +++ b/Lite.Tests/AnalysisPassTokenThreadingTests.cs @@ -0,0 +1,403 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Diagnostics; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Text.RegularExpressions; +using System.Threading; +using System.Threading.Tasks; +using PerformanceMonitorLite.Analysis; +using PerformanceMonitorLite.Database; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// The analysis pass's token has to reach the reads, not just the gaps between them (#2443). +/// +/// #2419 armed a per-pass budget and put a checkpoint ahead of every store-touching stage, which +/// closed the user-visible defect in #2412: a wedged server became loudly stuck instead of silently +/// skipped forever. What it did not do was hand the token to the store layer — its own diff touched no +/// store file — so abandonment happened BETWEEN stages. A collector that had already started ran to +/// completion, and the pass that had given up went on holding the read lock and the connection it was +/// no longer waiting for. +/// +/// The threading is mechanical; this file is the part that lasts. A sweep of 203 call sites only +/// stays swept if something counts it, and the failure mode it guards against is specific: adding +/// CancellationToken cancellationToken = default to signatures and stopping there, which leaves +/// every call site passing while LOOKING threaded. So the pin is on +/// the CALL, not the signature — a no-argument ExecuteReaderAsync() is what goes red, and no +/// default can hide it. The Darling twin is AnalysisPassTokenThreadingTests in Darling.Tests. +/// +public sealed class AnalysisPassTokenThreadingTests +{ + /// + /// The no-argument overloads, plus the read lock. Each has a token-taking sibling, so the empty + /// parentheses are the whole tell. AcquireReadLock() belongs on this list and not on a list + /// of its own: every store read on the pass takes that lock BEFORE it opens its connection, so a + /// pass that reached the reads with a token and the lock without one would still be uninterruptible + /// behind an archival — which is exactly the gap #2443 called out as Lite-only. + /// + private static readonly Regex s_untokenedStoreCall = new( + @"\.(?:ExecuteReaderAsync|ExecuteNonQueryAsync|ExecuteScalarAsync|ReadAsync|OpenAsync)\(\s*\)|AcquireReadLock\(\s*\)", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + /// + /// A member declaration, used only to attribute a call to the method that makes it. Deliberately + /// crude — it needs to name the enclosing method, not parse C#. The [^=;()]* is what keeps + /// field initializers and expression-bodied properties out. + /// + private static readonly Regex s_memberDeclaration = new( + @"^\s*(?:public|private|internal|protected)[^=;()]*\s(?[A-Za-z_][A-Za-z0-9_]*)\s*\(", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + private const string ExemptionMarker = "#2443 exempt"; + + /// + /// Every method allowed to make an untokened store call, and the reason it is allowed to — the same + /// five the Darling twin exempts, for the same two reasons. This is an enumeration and not a + /// wildcard on purpose: the exemption list this repo spent a day removing was a wildcard, and the + /// way that one grew was that nobody ever had to name a new entry. + /// + /// The four read-back surfaces are OFF the pass — the viewer, the MCP and the retention sweep + /// have no per-pass budget and no wedged analysis to abandon, so there is no token to thread. + /// InsertFindingAsync is the other kind: ON the pass, token deliberately withheld, because a + /// finding set cut in half is an outcome that reads like a healthy one (see + /// ). + /// + private static readonly Dictionary s_exempt = new(StringComparer.Ordinal) + { + ["GetRecentFindingsAsync"] = "read-back: the recommendations reader + MCP findings read", + ["GetLatestFindingsAsync"] = "read-back: the viewer's latest-run findings", + ["MuteStoryAsync"] = "off-pass write: the MCP/viewer mute verb", + ["CleanupOldFindingsAsync"] = "off-pass write: the retention sweep, its own lifetime", + ["InsertFindingAsync"] = "on-pass write, token withheld: a half-written finding set must not exist" + }; + + /// + /// The sweep, and the thing that keeps it swept. Scanned over the whole analysis directory rather + /// than a file list, because a file list is a thing a seventh DuckDbFactCollector partial can + /// be added outside of. + /// + [Fact] + public void NoStoreCallOnTheAnalysisPassRunsWithoutThePassToken() + { + var offenders = new List(); + + foreach (var (file, lines) in AnalysisSources()) + { + var enclosing = string.Empty; + + for (var i = 0; i < lines.Length; i++) + { + var declaration = s_memberDeclaration.Match(lines[i]); + if (declaration.Success) + { + enclosing = declaration.Groups["name"].Value; + } + + if (!s_untokenedStoreCall.IsMatch(lines[i])) + { + continue; + } + + if (s_exempt.ContainsKey(enclosing) + && DocBlockAbove(lines, IndexOfDeclaration(lines, i)).Contains(ExemptionMarker, StringComparison.Ordinal)) + { + continue; + } + + offenders.Add($"{file}:{i + 1} in {enclosing}(): {lines[i].Trim()}"); + } + } + + Assert.True(offenders.Count == 0, + "These store calls run without the pass's cancellation token, so the pass can be abandoned " + + "around them but never inside one. Pass context.CancellationToken (or the method's own token " + + $"parameter). If the call genuinely must complete, add its method to {nameof(s_exempt)} and put a " + + $"'{ExemptionMarker}' note in its doc comment saying why:\n" + + string.Join("\n", offenders)); + } + + /// + /// The other half of the agreement: an exemption that no longer corresponds to an untokened call is + /// a claim the code stopped making, and it must not be allowed to sit there authorising the next one. + /// + [Fact] + public void EveryExemptionIsStillEarningItsPlace() + { + var untokenedBy = new Dictionary(StringComparer.Ordinal); + var markedMethods = new HashSet(StringComparer.Ordinal); + + foreach (var (_, lines) in AnalysisSources()) + { + var enclosing = string.Empty; + var enclosingAt = -1; + + for (var i = 0; i < lines.Length; i++) + { + var declaration = s_memberDeclaration.Match(lines[i]); + if (declaration.Success) + { + enclosing = declaration.Groups["name"].Value; + enclosingAt = i; + if (DocBlockAbove(lines, enclosingAt).Contains(ExemptionMarker, StringComparison.Ordinal)) + { + markedMethods.Add(enclosing); + } + } + + if (s_untokenedStoreCall.IsMatch(lines[i]) && enclosingAt >= 0) + { + untokenedBy[enclosing] = untokenedBy.TryGetValue(enclosing, out var n) ? n + 1 : 1; + } + } + } + + Assert.Equal(s_exempt.Keys.OrderBy(k => k, StringComparer.Ordinal), markedMethods.OrderBy(k => k, StringComparer.Ordinal)); + Assert.Equal(s_exempt.Keys.OrderBy(k => k, StringComparer.Ordinal), untokenedBy.Keys.OrderBy(k => k, StringComparer.Ordinal)); + + /* Stated so a NEW untokened call inside an already-exempt method cannot ride in on the + exemption: 14 read-back calls across four methods, plus the one finding INSERT. */ + Assert.Equal(15, untokenedBy.Values.Sum()); + Assert.Equal(1, untokenedBy["InsertFindingAsync"]); + } + + /// + /// The fact collector is where the token would have been silently useless. Its per-query catches were + /// bare catch { } — deliberately, so a missing table degrades to "no facts" — which meant an + /// armed token produced 27 swallowed cancellations and a pass that carried on collecting under a + /// token that had already fired. Every catch there now lets an abandonment through, which is what + /// turns the threaded token into an exit rather than 27 wasted lock acquisitions. + /// + [Fact] + public void EveryCatchOnTheFactCollectorLetsAnAbandonmentThrough() + { + var bare = new List(); + var opensCatch = new Regex(@"^\s*catch\b", RegexOptions.Compiled | RegexOptions.CultureInvariant); + var classified = new Regex( + @"^\s*catch\s*\(\s*Exception\s+ex\s*\)\s*when\s*\(\s*!AnalysisAbandon\.IsExpected\(", + RegexOptions.Compiled | RegexOptions.CultureInvariant); + + foreach (var (file, lines) in AnalysisSources()) + { + if (!Path.GetFileName(file).StartsWith("DuckDbFactCollector.", StringComparison.Ordinal)) + { + continue; + } + + for (var i = 0; i < lines.Length; i++) + { + if (opensCatch.IsMatch(lines[i]) && !classified.IsMatch(lines[i])) + { + bare.Add($"{file}:{i + 1}: {lines[i].Trim()}"); + } + } + } + + Assert.True(bare.Count == 0, + "A fact-collector catch that does not classify swallows the abandonment the token was armed for, " + + "and the pass keeps collecting under a fired token:\n" + string.Join("\n", bare)); + + /* 27 collect methods guard their read this way; the number is stated so deleting one is a + decision. The other four have no catch at all and never had one — their reads propagate + straight to the pass, which is the same outcome by a shorter route. */ + Assert.Equal(27, AnalysisSources() + .Where(s => Path.GetFileName(s.File).StartsWith("DuckDbFactCollector.", StringComparison.Ordinal)) + .SelectMany(s => s.Lines) + .Count(line => classified.IsMatch(line))); + } + + /// + /// The decision #2443 asked to be made explicitly rather than assumed: what a cancelled PERSIST + /// means. Every row shares one analysis_time and the latest-findings read takes the newest + /// analysis_time, so a batch cut in half does not read as truncated; it reads as a complete + /// analysis that found fewer problems, and the server looks HEALTHIER for having been abandoned. + /// The lock and the connection open are therefore the last abandonment points: before the first + /// row, or not at all. + /// + /// #2448 closed the other half of that. #2443 could only reason about the CANCELLATION path, + /// and the same truncated set was still reachable from an ordinary store fault mid-batch, where no + /// amount of token discipline reaches. The batch is now one transaction, so it is all-or-nothing + /// against a fault as well. Deliberately the same answer as the Darling twin's, pinned the same + /// way, because a divergence here would be a parity bug rather than a local choice. + /// + [Fact] + public void TheFindingInsertIsAbandonedBeforeItStartsOrNotAtAll() + { + var store = File.ReadAllText(Path.Combine(AnalysisDirectory(), "FindingStore.cs")) + .Replace("\r\n", "\n", StringComparison.Ordinal); + + var insertBatch = Between(store, + "public async Task> InsertFindingsAsync(", + "public async Task> SaveFindingsAsync("); + + /* The abandonment points: both the lock wait and the connection open observe the pass token. */ + Assert.Contains("_duckDb.AcquireReadLock(context.CancellationToken)", insertBatch, StringComparison.Ordinal); + Assert.Contains("await connection.OpenAsync(context.CancellationToken)", insertBatch, StringComparison.Ordinal); + + /* And nothing after them does. A ThrowIfCancellationRequested between rows would MAKE the + partial set rather than prevent it, which is why there is none. */ + Assert.DoesNotContain("ThrowIfCancellationRequested", insertBatch, StringComparison.Ordinal); + + /* #2448: the batch is one transaction, and the rows are enlisted in it. Both halves are pinned + because either alone is silently useless — a transaction the rows do not join commits nothing + of theirs, and enlisting in a transaction nobody commits writes nothing at all. */ + Assert.Contains("connection.BeginTransaction();", insertBatch, StringComparison.Ordinal); + Assert.Contains("transaction.Commit();", insertBatch, StringComparison.Ordinal); + Assert.Contains("cmd.Transaction = transaction;", store, StringComparison.Ordinal); + + /* The row write states the decision where someone changing it will read it. */ + Assert.Contains(ExemptionMarker, + Between(store, "/// Inserts one finding on an already-open connection", "private static async Task InsertFindingAsync("), + StringComparison.Ordinal); + } + + /// + /// The classifier's truth table. Both halves of the predicate are load-bearing and the TYPE half + /// especially so: since #2419 this token fires on an ordinary timeout, so it is signalled during + /// perfectly normal running, and a filter that asked only "has the token fired?" would relabel any + /// genuine fault landing after the budget elapsed as an abandonment — swallowing the one line of + /// evidence it left, at forty-odd catch sites at once. + /// + [Fact] + public void AnAbandonmentIsAFiredTokenAndACancellationShape_NeverOneWithoutTheOther() + { + var fired = new CancellationToken(canceled: true); + + Assert.True(AnalysisAbandon.IsExpected(new OperationCanceledException(), fired)); + Assert.True(AnalysisAbandon.IsExpected(new TaskCanceledException(), fired)); + + /* A fault that merely coincides with the budget expiring is still a fault. */ + Assert.False(AnalysisAbandon.IsExpected(new InvalidOperationException("store"), fired)); + Assert.False(AnalysisAbandon.IsExpected(new TimeoutException("deadline"), fired)); + + /* And a cancellation shape with nothing cancelled means something threw it for another + reason, which must keep its error. */ + Assert.False(AnalysisAbandon.IsExpected(new OperationCanceledException(), CancellationToken.None)); + } + + /// + /// The Lite-only half of #2443, and the one place cancellation genuinely could not be replaced by + /// reporting. Every store read on the pass takes the read lock BEFORE opening its connection, and + /// AcquireReadLock() had no timeout and no token while its AcquireWriteLock(TimeSpan?) + /// sibling had one — so a pass queued behind a long archival sat in an uninterruptible + /// EnterReadLock() however carefully its reads were threaded. + /// + /// Held behind a real write lock on another thread, which is what an archival looks like from + /// here. The old overload would block for the full hold; the token'd one gives up when asked. + /// + [Fact] + public void TheReadLockWaitIsAbandonableWhileAWriterHoldsIt() + { + var initializer = new DuckDbInitializer(Path.Combine(Path.GetTempPath(), $"pm-lock-{Guid.NewGuid():N}.db")); + + var writerHasIt = new ManualResetEventSlim(false); + var releaseWriter = new ManualResetEventSlim(false); + + var writer = Task.Run(() => + { + using var write = initializer.AcquireWriteLock(); + writerHasIt.Set(); + releaseWriter.Wait(TimeSpan.FromSeconds(30)); + }); + + try + { + Assert.True(writerHasIt.Wait(TimeSpan.FromSeconds(5)), "the writer never took the lock"); + + using var cts = new CancellationTokenSource(TimeSpan.FromMilliseconds(250)); + var elapsed = Stopwatch.StartNew(); + + Assert.Throws(() => + { + using var read = initializer.AcquireReadLock(cts.Token); + }); + + /* Generously bounded — the assertion is "it gave up while the writer still held it", + not a latency measurement. The writer holds for up to 30s. */ + Assert.True(elapsed.Elapsed < TimeSpan.FromSeconds(10), + $"the read lock gave up only after {elapsed.Elapsed.TotalSeconds:F1}s"); + } + finally + { + releaseWriter.Set(); + writer.Wait(TimeSpan.FromSeconds(30)); + } + + /* And once the writer is gone it is an ordinary read lock again, token or no token. */ + using var after = initializer.AcquireReadLock(CancellationToken.None); + Assert.NotNull(after); + } + + private static string Between(string source, string start, string end) + { + var from = source.IndexOf(start, StringComparison.Ordinal); + Assert.True(from >= 0, $"anchor not found: {start}"); + var to = source.IndexOf(end, from, StringComparison.Ordinal); + Assert.True(to > from, $"anchor not found after {start}: {end}"); + return source[from..to]; + } + + /// The declaration line a call at belongs to (scanning up). + private static int IndexOfDeclaration(string[] lines, int callLine) + { + for (var i = callLine; i >= 0; i--) + { + if (s_memberDeclaration.IsMatch(lines[i])) + { + return i; + } + } + + return 0; + } + + /// The contiguous /// block immediately above a declaration, as one string. + private static string DocBlockAbove(string[] lines, int declarationLine) + { + var doc = new List(); + for (var i = declarationLine - 1; i >= 0 && lines[i].TrimStart().StartsWith("///", StringComparison.Ordinal); i--) + { + doc.Add(lines[i]); + } + + return string.Join("\n", doc); + } + + private static IEnumerable<(string File, string[] Lines)> AnalysisSources() + { + var files = Directory.GetFiles(AnalysisDirectory(), "*.cs", SearchOption.TopDirectoryOnly); + Assert.NotEmpty(files); + + foreach (var file in files.OrderBy(f => f, StringComparer.Ordinal)) + { + /* The working copy is CRLF; split on the LF so a line never carries a stray CR. */ + yield return (Path.GetFileName(file), + File.ReadAllText(file).Replace("\r\n", "\n", StringComparison.Ordinal).Split('\n')); + } + } + + private static string AnalysisDirectory() => Path.Combine(RepoRoot(), "Lite", "Analysis"); + + private static string RepoRoot([CallerFilePath] string thisFile = "") + { + var dir = Path.GetDirectoryName(thisFile)!; + while (dir is not null && !File.Exists(Path.Combine(dir, "PerformanceMonitor.sln")) && !Directory.Exists(Path.Combine(dir, ".git"))) + { + dir = Path.GetDirectoryName(dir); + } + + Assert.NotNull(dir); + return dir!; + } +} diff --git a/Lite.Tests/AsOfWindowAnchorTests.cs b/Lite.Tests/AsOfWindowAnchorTests.cs new file mode 100644 index 000000000..02b0c5124 --- /dev/null +++ b/Lite.Tests/AsOfWindowAnchorTests.cs @@ -0,0 +1,298 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitor.Common; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2495, Lite half: as_of moves the END of a windowed read off "now" so a caller can ask about a +/// past incident. Darling's twin (AsOfWindowAnchorLivePostgresTests) pins the same behaviour against +/// live Postgres; this pins it against a real DuckDB through the real tool methods. +/// +/// These run through the SERVICE, not just the helper, because the anchor has to survive the whole +/// path — tool parameter, shared validator, LocalDataService's window helper, and the SQL's two +/// bounds. A helper-only test would pass with the plumbing missing, which is exactly the shape of bug the +/// #2213 review found four of. +/// +/* The server-local half below reads ServerTimeHelper.UtcOffsetMinutes, a process-wide mutable static that a + sibling class SETS for its duration. xUnit runs test classes in parallel, so this joins the collection the + other offset-reading classes already use rather than racing them into a window hours away from its rows. */ +[Collection("server-time-helper")] +public sealed class AsOfWindowAnchorTests : IClassFixture, IDisposable +{ + /* Outside every default window on the surface (the widest default is 24 hours). The two windows are + told apart by CONTENT rather than by row count, so a read returning the wrong one cannot look right. */ + private const string IncidentWait = "PAGEIOLATCH_SH"; + private const string RecentWait = "CXPACKET"; + + private readonly string _tempDir; + private readonly DuckDbInitializer _duckDb; + private readonly LocalDataService _dataService; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private long _nextId = -1; + private DuckDBConnection? _seedConn; + + public AsOfWindowAnchorTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _tempDir = Path.Combine(Path.GetTempPath(), "AsOfAnchorTests_" + Guid.NewGuid().ToString("N")[..8]); + var configDir = Path.Combine(_tempDir, "config"); + Directory.CreateDirectory(configDir); + + _dataService = new LocalDataService(_duckDb); + _serverManager = new ServerManager(configDir); + + var server = new ServerConnection { ServerName = "TestServer", DisplayName = "TestServer" }; + _serverManager.AddServer(server); + + /* The derived id, not a literal: seeding under a hand-picked number makes the read return nothing + and the test pass for the wrong reason. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { if (Directory.Exists(_tempDir)) Directory.Delete(_tempDir, recursive: true); } + catch (IOException) { /* best-effort cleanup */ } + catch (UnauthorizedAccessException) { /* best-effort cleanup */ } + } + + /// + /// The whole point of the issue: the same read, anchored at a past incident, returns the incident — + /// and on the default anchor it does not. + /// + [Fact] + public async Task AnAnchoredRead_SeesThePastWindow_AndTheDefaultAnchorDoesNot() + { + var now = DateTime.UtcNow; + var incident = now.AddHours(-30); + + await SeedWaitAsync(incident, IncidentWait, 900_000); + await SeedWaitAsync(incident.AddMinutes(5), IncidentWait, 800_000); + await SeedWaitAsync(now.AddMinutes(-10), RecentWait, 1_000); + + var anchor = incident.AddMinutes(30).ToString("o"); + + var live = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer"); + Assert.Contains(RecentWait, WaitTypesIn(live)); + Assert.DoesNotContain(IncidentWait, WaitTypesIn(live)); + + var anchored = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4, 20, anchor); + Assert.Contains(IncidentWait, WaitTypesIn(anchored)); + Assert.DoesNotContain(RecentWait, WaitTypesIn(anchored)); + } + + /// + /// Widening hours_back is not the workaround it looks like. A window wide enough to reach the + /// incident reaches everything since it too, so the aggregate now describes both periods — a different + /// answer, not the same answer with more rows. + /// + [Fact] + public async Task WideningTheWindow_IsADifferentAnswer_NotTheSameOneWithMoreRows() + { + var now = DateTime.UtcNow; + var incident = now.AddHours(-30); + + await SeedWaitAsync(incident, IncidentWait, 900_000); + await SeedWaitAsync(now.AddMinutes(-10), RecentWait, 1_000); + + var anchored = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4, 20, incident.AddMinutes(30).ToString("o")); + var widened = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 48); + + Assert.Equal(new[] { IncidentWait }, WaitTypesIn(anchored)); + Assert.Equal(2, WaitTypesIn(widened).Length); + Assert.Contains(IncidentWait, WaitTypesIn(widened)); + Assert.Contains(RecentWait, WaitTypesIn(widened)); + } + + /// + /// Backward compatibility, stated as a test rather than as a claim: a caller that sends only + /// hours_back gets the window it always got. Seeded either side of the boundary so a window that + /// silently moved would show up as content, not as a timing wobble. + /// + [Fact] + public async Task WithNoAnchor_TheWindowIsStillTheLastHoursBackHours() + { + var now = DateTime.UtcNow; + await SeedWaitAsync(now.AddHours(-2), RecentWait, 5_000); + await SeedWaitAsync(now.AddHours(-10), IncidentWait, 5_000); + + Assert.Equal(new[] { RecentWait }, WaitTypesIn(await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4))); + + var both = WaitTypesIn(await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 24)); + Assert.Contains(RecentWait, both); + Assert.Contains(IncidentWait, both); + } + + /// + /// An anchor the read cannot use is refused with a message. It must NOT quietly answer as of now — that + /// answer is indistinguishable from a correct one, which is the failure this parameter exists to remove. + /// + [Fact] + public async Task AnUnusableAnchor_IsRefused_RatherThanAnsweredAsOfNow() + { + await SeedWaitAsync(DateTime.UtcNow.AddMinutes(-10), RecentWait, 1_000); + + var garbage = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4, 20, "last tuesday"); + Assert.StartsWith("Invalid as_of", garbage, StringComparison.Ordinal); + Assert.DoesNotContain(RecentWait, garbage, StringComparison.Ordinal); + + /* The one a general date parser would let through: "01/02/2026" is M/d/yyyy under the invariant + culture, so a caller who meant 1 February would silently get a window around 2 January. */ + var ambiguous = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4, 20, "01/02/2026"); + Assert.StartsWith("Invalid as_of", ambiguous, StringComparison.Ordinal); + + var future = await McpWaitTools.GetWaitStats( + _dataService, _serverManager, "TestServer", 4, 20, DateTime.UtcNow.AddDays(1).ToString("o")); + Assert.Contains("future", future, StringComparison.Ordinal); + Assert.DoesNotContain(RecentWait, future, StringComparison.Ordinal); + } + + /// + /// An anchor older than anything the store holds is NOT refused — the store's earliest row is + /// per-server and per-collector, so a floor would be a guess. It comes back as the read's own honest + /// miss, which is unambiguous because the caller chose the window. + /// + [Fact] + public async Task AnAnchorBeforeAnyData_ReturnsTheReadsOwnMiss_NotARefusal() + { + await SeedWaitAsync(DateTime.UtcNow.AddMinutes(-10), RecentWait, 1_000); + + var ancient = await McpWaitTools.GetWaitStats(_dataService, _serverManager, "TestServer", 4, 20, "1999-01-01T00:00:00Z"); + + using var doc = JsonDocument.Parse(ancient); + Assert.Equal("unavailable", doc.RootElement.GetProperty("status").GetString()); + } + + /// + /// The anchor reaches a read whose window is built by the SERVER-LOCAL helper, not just the UTC one. + /// Two families, two time bases, one instant — a bug here would shift the window by the monitored + /// server's offset and still look plausible. + /// + /// The offset is SET, not read (review catch). Reading whatever + /// ServerTimeHelper.UtcOffsetMinutes happens to be means running at 0 on any normal machine, and + /// at 0 a correct conversion and no conversion at all produce identical results — the test would pass + /// vacuously for exactly the bug its own summary names. UTC+5:30 is used because a half-hour offset also + /// catches an implementation that rounds to whole hours. Saved and restored, and this class shares the + /// server-time-helper collection so the write cannot land under a sibling class mid-read. + /// + [Fact] + public async Task TheAnchorAlsoMoves_AServerLocalWindow() + { + var savedOffset = ServerTimeHelper.UtcOffsetMinutes; + try + { + ServerTimeHelper.UtcOffsetMinutes = 330; // UTC+5:30 + + var now = DateTime.UtcNow; + var incident = now.AddHours(-30); + + /* sample_time is the monitored server's LOCAL wall clock, so the fixture is seeded in that base — + the read's window is server-local and the anchor arrives in UTC, and the point of the test is + that the conversion between them happens exactly once. */ + var offset = ServerTimeHelper.UtcOffsetMinutes; + await SeedCpuAsync(incident.AddMinutes(offset), 91); + await SeedCpuAsync(now.AddMinutes(-10).AddMinutes(offset), 12); + + var anchored = await McpCpuTools.GetCpuUtilization( + _dataService, _serverManager, "TestServer", 4, incident.AddMinutes(30).ToString("o")); + Assert.Equal(new[] { 91 }, SqlCpuIn(anchored)); + + var live = await McpCpuTools.GetCpuUtilization(_dataService, _serverManager, "TestServer"); + Assert.Equal(new[] { 12 }, SqlCpuIn(live)); + } + finally + { + ServerTimeHelper.UtcOffsetMinutes = savedOffset; + } + } + + // ── helpers ── + + private static string[] WaitTypesIn(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("waits").EnumerateArray() + .Select(w => w.GetProperty("wait_type").GetString()!) + .ToArray(); + } + + private static int[] SqlCpuIn(string json) + { + using var doc = JsonDocument.Parse(json); + return doc.RootElement.GetProperty("samples").EnumerateArray() + .Select(s => s.GetProperty("sql_server_cpu").GetInt32()) + .ToArray(); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedWaitAsync(DateTime at, string waitType, long deltaMs) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + + using var cmd = conn.CreateCommand(); + cmd.CommandText = @"INSERT INTO wait_stats + (collection_id, collection_time, server_id, server_name, wait_type, + waiting_tasks_count, wait_time_ms, signal_wait_time_ms, + delta_waiting_tasks, delta_wait_time_ms, delta_signal_wait_time_ms) + VALUES ($1, $2, $3, 'TestServer', $4, 0, 0, 0, 50, $5, 0)"; + void P(object v) => cmd.Parameters.Add(new DuckDBParameter { Value = v }); + P(_nextId--); + P(at); + P(_serverId); + P(waitType); + P(deltaMs); + await cmd.ExecuteNonQueryAsync(); + } + + private async Task SeedCpuAsync(DateTime at, int sqlCpu) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + + using var cmd = conn.CreateCommand(); + cmd.CommandText = @"INSERT INTO cpu_utilization_stats + (collection_id, collection_time, server_id, server_name, sample_time, + sqlserver_cpu_utilization, other_process_cpu_utilization) + VALUES ($1, $2, $3, 'TestServer', $4, $5, 0)"; + void P(object v) => cmd.Parameters.Add(new DuckDBParameter { Value = v }); + P(_nextId--); + P(at); + P(_serverId); + P(at); + P(sqlCpu); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/AuroraOnlySqlIsGatedTests.cs b/Lite.Tests/AuroraOnlySqlIsGatedTests.cs new file mode 100644 index 000000000..f24c7f60e --- /dev/null +++ b/Lite.Tests/AuroraOnlySqlIsGatedTests.cs @@ -0,0 +1,190 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Text.RegularExpressions; +using Xunit; + +namespace Lite.Tests; + +/// +/// Every reference to an Aurora-only SQL surface is accounted for — either the code path is gated to +/// Aurora, or it has a vanilla alternative chosen by flavor. +/// +/// +/// This class exists because the same defect shipped twice. #2625: pg_statement_stats read +/// aurora_stat_statements() behind an AppliesTo => IsAurora gate, so self-hosted +/// PostgreSQL had no answer at all to "which queries cost the most" — the question a database monitor +/// exists for. #2651: the statement TEXT store read the same function with no gate and no alternative, so +/// off Aurora it silently never populated, which made get_pg_top_queries return a null +/// query_text forever and made test_hypothetical_index unable to work at all. +/// +/// +/// +/// The second one survived the audit that found the first, and the reason is structural rather than +/// careless: pg_statement_text is not a collector. It is refreshed directly by the worker, so it +/// was in neither the catalog sweep nor the AppliesTo audit. A guard that only walks +/// ICollectorDefinition implementations cannot see it. This one walks the SOURCE. +/// +/// +/// +/// The allow-list is the point. Each entry is a claim that somebody looked at that file and decided +/// what its Aurora dependency means for a self-hosted target. A new file appearing here is not +/// necessarily a bug — but it IS a decision, and the failure message asks for it rather than letting the +/// file arrive unexamined. +/// +/// +public class AuroraOnlySqlIsGatedTests +{ + /// The Aurora-extended surfaces. Community PostgreSQL has none of them under any version. + private static readonly string[] AuroraSurfaces = + { + "aurora_stat_statements", + "aurora_stat_system_waits", + "aurora_stat_wait_type", + "aurora_stat_wait_event", + }; + + /// + /// Every file allowed to name one, and WHY it is allowed. Lowering an entry to "gated" or "paired" is + /// a decision someone made; adding a file is one someone must make. + /// + private static readonly Dictionary Accounted = new(StringComparer.OrdinalIgnoreCase) + { + ["PgWaitStatsCollector.cs"] = + "GATED. AppliesTo => IsAurora, and correctly: it reads Aurora's own wait instrumentation, which " + + "community PostgreSQL has no equivalent of. pg_wait_sampling answers the same question there, and " + + "since #2625 the permanent-gap message names it.", + + ["PgStatementStatsCollector.cs"] = + "PAIRED (#2625). Aurora reads aurora_stat_statements(); every other PostgreSQL reads the vanilla " + + "pg_stat_statements view with the Aurora-only columns as typed NULLs. Chosen in BuildQuery, not by " + + "AppliesTo, because the source differs and the capability does not.", + + ["PgStatementText.cs"] = + "PAIRED (#2651). Same split, one layer down: FetchSqlFor(isAurora, major) picks Aurora's function " + + "or the vanilla view. This is the one that shipped ungated and made test_hypothetical_index dead " + + "on self-hosted PostgreSQL.", + + ["CollectorEngineCapability.cs"] = + "PROSE. Names the surfaces in the sentence shown to an operator when a collector cannot run here. " + + "It describes the dependency rather than depending on it.", + }; + + [Fact] + public void EveryFileNamingAnAuroraOnlySurface_IsAccountedFor() + { + var found = FilesNamingAuroraSurfaces(); + + Assert.NotEmpty(found); + + var unaccounted = found.Where(f => !Accounted.ContainsKey(Path.GetFileName(f))) + .OrderBy(f => f, StringComparer.Ordinal) + .ToArray(); + + Assert.True( + unaccounted.Length == 0, + "These files reference an Aurora-only SQL surface and nothing records what that means for a " + + "self-hosted PostgreSQL target:\n " + string.Join("\n ", unaccounted) + + "\n\nAurora-only SQL is fine. Aurora-only SQL that NOBODY DECIDED ABOUT is how #2651 shipped: the " + + "statement-text store read aurora_stat_statements() with no gate and no alternative, so off Aurora " + + "it silently never populated and a feature that depends on it could not work at all.\n\n" + + "Decide which this is, then add it to Accounted:\n" + + " GATED - the collector's AppliesTo excludes non-Aurora targets, and the capability really is absent there.\n" + + " PAIRED - there is a vanilla path chosen by flavor, so the SOURCE differs and the capability does not.\n" + + " PROSE - it only names the surface in a message."); + } + + /// + /// The allow-list cannot outlive what it describes. An entry naming a file that no longer references + /// any Aurora surface is a note about a decision that has been undone — and it would silently permit + /// that file to reacquire one later. + /// + [Fact] + public void TheAllowListHasNoStaleEntries() + { + var found = FilesNamingAuroraSurfaces().Select(Path.GetFileName).ToHashSet(StringComparer.OrdinalIgnoreCase); + + var stale = Accounted.Keys.Where(k => !found.Contains(k)) + .OrderBy(k => k, StringComparer.Ordinal) + .ToArray(); + + Assert.True(stale.Length == 0, + "These allow-list entries name no Aurora surface any more — remove them, or the file could " + + "reacquire one unexamined: " + string.Join(", ", stale)); + } + + /// + /// The two PAIRED files must actually branch. An entry claiming a vanilla alternative while the code + /// has none would be worse than no entry: it records a decision that was never implemented. + /// + [Theory] + [InlineData("PgStatementStatsCollector.cs")] + [InlineData("PgStatementText.cs")] + public void APairedFileReallyReadsTheVanillaViewToo(string fileName) + { + var path = FilesNamingAuroraSurfaces().Single(f => string.Equals(Path.GetFileName(f), fileName, StringComparison.OrdinalIgnoreCase)); + var source = File.ReadAllText(path); + + Assert.Contains("pg_stat_statements", source, StringComparison.Ordinal); + + /* Case-INSENSITIVE: the collector branches on context.Target.IsAurora and the text store on a + parameter named isAurora, and a guard that only accepted one spelling would fail on correct + code. It did, on exactly this file, the first time this test ran in CI. */ + Assert.Contains("isaurora", source, StringComparison.OrdinalIgnoreCase); + } + + private static string[] FilesNamingAuroraSurfaces() + { + var root = RepoRoot(); + + return Directory.EnumerateFiles(root, "*.cs", SearchOption.AllDirectories) + .Where(f => !f.Contains($"{Path.DirectorySeparatorChar}obj{Path.DirectorySeparatorChar}", StringComparison.Ordinal)) + .Where(f => !f.Contains($"{Path.DirectorySeparatorChar}bin{Path.DirectorySeparatorChar}", StringComparison.Ordinal)) + .Where(f => !f.Contains($"{Path.DirectorySeparatorChar}.claude{Path.DirectorySeparatorChar}", StringComparison.Ordinal)) + /* Tests are excluded deliberately: a test naming the surface is asserting ABOUT it, which is + the opposite of depending on it. */ + .Where(f => !f.Contains(".Tests", StringComparison.OrdinalIgnoreCase)) + .Where(NamesAnAuroraSurfaceOutsideAComment) + .OrderBy(f => f, StringComparer.Ordinal) + .ToArray(); + } + + /// + /// Comment lines are skipped, or every explanatory paragraph about why Aurora is different would count + /// as a dependency on it — which would make the guard so noisy that the allow-list stopped being read. + /// + private static bool NamesAnAuroraSurfaceOutsideAComment(string path) + => File.ReadLines(path).Any(line => + { + var trimmed = line.TrimStart(); + + if (trimmed.StartsWith("//", StringComparison.Ordinal) + || trimmed.StartsWith("*", StringComparison.Ordinal) + || trimmed.StartsWith("/*", StringComparison.Ordinal)) + { + return false; + } + + return AuroraSurfaces.Any(s => line.Contains(s, StringComparison.Ordinal)); + }); + + private static string RepoRoot() + { + var directory = new DirectoryInfo(AppContext.BaseDirectory); + while (directory is not null && !File.Exists(Path.Combine(directory.FullName, "PerformanceMonitor.sln"))) + { + directory = directory.Parent; + } + + return directory?.FullName ?? throw new InvalidOperationException("PerformanceMonitor.sln not found above the test output directory."); + } +} diff --git a/Lite.Tests/AzureDatabaseSizeSiblingTests.cs b/Lite.Tests/AzureDatabaseSizeSiblingTests.cs new file mode 100644 index 000000000..eae5815fb --- /dev/null +++ b/Lite.Tests/AzureDatabaseSizeSiblingTests.cs @@ -0,0 +1,135 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Reflection; +using System.Text.RegularExpressions; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2643: on Azure SQL DB this collector reported the connected database and nothing else. +/// +/// +/// Correct — sys.database_files is database-scoped — and indistinguishable from a collector that +/// managed to find only master. A reporter with fifty databases pointed the Viewer at master, +/// saw master's two files on a grid headed "All Servers", and filed it. I told them the platform +/// made anything else impossible. It does not: sys.resource_stats is a master-only view carrying +/// storage_in_megabytes per database, verified against a live Azure SQL Database. +/// +/// +public class AzureDatabaseSizeSiblingTests +{ + private static string AzureSql => + (string)typeof(DatabaseSizeStatsCollector) + .GetField("AzureSqlDbQueryText", BindingFlags.NonPublic | BindingFlags.Static)! + .GetValue(null)!; + + [Fact] + public void TheAzureQueryReadsBothTheConnectedDatabaseAndItsSiblings() + { + Assert.Contains("sys.database_files", AzureSql, StringComparison.Ordinal); + Assert.Contains("sys.resource_stats", AzureSql, StringComparison.Ordinal); + } + + /// + /// The sibling read goes through sp_executesql, and that is the load-bearing detail. + /// + /// sys.resource_stats does not EXIST in a user database, and SQL Server resolves names at + /// PARSE time — so a plain UNION guarded by WHERE DB_NAME() = N'master' still fails with + /// error 208 on every user database, which is the common case. That is not a theory: the first + /// version of this shipped exactly that shape, and running it from a user database returned 208 + /// immediately. Deferring the reference until the branch runs is the only thing that fixes it. + /// + [Fact] + public void TheSiblingReadIsDeferred_BecauseTheViewDoesNotExistInAUserDatabase() + { + Assert.Contains("IF DB_NAME() = N'master'", AzureSql, StringComparison.Ordinal); + Assert.Contains("EXEC sys.sp_executesql", AzureSql, StringComparison.Ordinal); + + /* The reference must be INSIDE the deferred string, not in the outer batch where parsing reaches + it regardless of the branch. */ + var execIndex = AzureSql.IndexOf("EXEC sys.sp_executesql", StringComparison.Ordinal); + var viewIndex = AzureSql.IndexOf("sys.resource_stats", StringComparison.Ordinal); + + Assert.True(viewIndex > execIndex, + "sys.resource_stats is referenced in the outer batch — parsing reaches it on a user database and fails 208 before any guard runs."); + } + + /// + /// A sibling row is honest about being a database rather than a file. sys.resource_stats has no + /// per-file breakdown, so the row says so: a NULL file_id and a name that reads as a database. + /// A fabricated file name would make the grid look complete and be wrong. + /// + [Fact] + public void ASiblingRowIsLabelledAsAWholeDatabase_NotAFabricatedFile() + { + Assert.Contains("file_name = N''(whole database)''", AzureSql, StringComparison.Ordinal); + } + + /// + /// used_size_mb is not projected for a sibling, and the omission is the point: the table + /// variable defaults it to NULL. Zero would say the database is empty, which is a measurement nobody + /// took. + /// + [Fact] + public void TheSiblingInsertOmitsWhatItCannotMeasure() + { + /* Sliced from the INSERT's own column list, not from the first parenthesis after the IF — that one + belongs to DB_NAME(), and the first version of this assertion happily tested the string "(". */ + var insert = AzureSql[AzureSql.IndexOf("IF DB_NAME() = N'master'", StringComparison.Ordinal)..]; + var listStart = insert.IndexOf("@database_sizes", StringComparison.Ordinal); + var open = insert.IndexOf('(', listStart); + var columnList = insert[open..insert.IndexOf(')', open)]; + + Assert.DoesNotContain("used_size_mb", columnList, StringComparison.Ordinal); + Assert.DoesNotContain("auto_growth_mb", columnList, StringComparison.Ordinal); + Assert.Contains("total_size_mb", columnList, StringComparison.Ordinal); + } + + /// + /// The connected database is excluded from the sibling arm — the file arm already reported it, with + /// real files. Without this every Azure entry reports its own database twice, once properly and once + /// as a sizeless "(whole database)" row. + /// + [Fact] + public void TheConnectedDatabaseIsNotReportedTwice() + => Assert.Contains("r.database_name <> DB_NAME()", AzureSql, StringComparison.Ordinal); + + /// + /// Newest sample per database. sys.resource_stats keeps roughly fourteen days at five-minute + /// grain, so without this every database arrives a few thousand times. + /// + [Fact] + public void OnlyTheNewestSamplePerDatabaseIsTaken() + { + Assert.Contains("ROW_NUMBER() OVER (PARTITION BY r.database_name ORDER BY r.end_time DESC)", AzureSql, StringComparison.Ordinal); + Assert.Contains("WHERE rs.rn = 1", AzureSql, StringComparison.Ordinal); + } + + /// + /// The final projection must still match PayloadColumns exactly — the collector writes by + /// position, and a table variable makes it easy to reorder one and not the other. + /// + [Fact] + public void TheFinalProjectionMatchesThePayloadColumnsInOrder() + { + var final = AzureSql[AzureSql.LastIndexOf("FROM @database_sizes", StringComparison.Ordinal)..]; + var select = AzureSql[..AzureSql.LastIndexOf("FROM @database_sizes", StringComparison.Ordinal)]; + select = select[select.LastIndexOf("SELECT", StringComparison.Ordinal)..]; + + var projected = Regex.Matches(select, @"ds\.(\w+)").Select(m => m.Groups[1].Value).ToArray(); + var declared = DatabaseSizeStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Equal(declared, projected); + Assert.NotEmpty(final); + } +} diff --git a/Lite.Tests/AzureDeadlockTelemetryTests.cs b/Lite.Tests/AzureDeadlockTelemetryTests.cs new file mode 100644 index 000000000..257a14ff2 --- /dev/null +++ b/Lite.Tests/AzureDeadlockTelemetryTests.cs @@ -0,0 +1,211 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Reflection; +using System.Threading; +using System.Threading.Tasks; +using System.Text.RegularExpressions; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2641: Azure SQL DB deadlock capture had one source, and it was the wrong shape twice over. +/// +/// +/// The session we create is database-scoped, so a connection to master captures only +/// master's deadlocks — a reporter with fifty user databases got essentially none of theirs — and +/// its ring buffer is memory-resident, so an Azure failover empties it without notice. +/// +/// +/// +/// sys.fn_xe_telemetry_blob_target_read_file is Azure's own file-backed telemetry, and it is +/// master-scoped. That last fact is the whole trick, and it cost two wrong conclusions before it was +/// found: called from a user database the identical statement returns zero rows, silently, with no error. +/// I reported the source unusable on that basis before testing it from master, where it returned the +/// deadlock immediately — with the USER database's name attached. +/// +/// +/// +/// Verified against a live Azure SQL Database, running the shipped query text itself rather than a +/// paraphrase: from master it returns source_database_name = 'Erik' for a deadlock that +/// happened in the user database; from that user database the same text returns the ring-buffer row with a +/// NULL source database, unchanged. +/// +/// +public class AzureDeadlockTelemetryTests +{ + private static string AzureSql => + (string)typeof(DeadlocksCollector) + .GetField("AzureQueryText", BindingFlags.NonPublic | BindingFlags.Static)! + .GetValue(null)!; + + private static string ServerScopedSql => + (string)typeof(DeadlocksCollector) + .GetField("ServerScopedQueryText", BindingFlags.NonPublic | BindingFlags.Static)! + .GetValue(null)!; + + [Fact] + public void TheAzureQueryReadsTheDurableTelemetryBlob_NotOnlyTheRingBuffer() + { + Assert.Contains("sys.dm_xe_database_session_targets", AzureSql, StringComparison.Ordinal); + Assert.Contains("sys.fn_xe_telemetry_blob_target_read_file('dl', NULL, NULL, NULL)", AzureSql, StringComparison.Ordinal); + Assert.Contains("UNION ALL", AzureSql, StringComparison.Ordinal); + } + + /// + /// The guard, and it is not an optimization. From a user database the telemetry call returns zero rows + /// with no error — indistinguishable from "this server had no deadlocks". Restricting the read to the + /// only connection that can answer keeps a silent zero from ever being produced. + /// + [Fact] + public void TheTelemetryArmIsGuardedOnBeingConnectedToMaster() + { + Assert.Contains("DB_NAME() = N'master'", AzureSql, StringComparison.Ordinal); + + /* Inside the telemetry subquery, not applied to the whole union — the ring-buffer arm must keep + working from a user database, which is where most entries point. */ + var telemetryIndex = AzureSql.IndexOf("fn_xe_telemetry_blob_target_read_file", StringComparison.Ordinal); + var guardIndex = AzureSql.IndexOf("DB_NAME() = N'master'", StringComparison.Ordinal); + + Assert.True(guardIndex > telemetryIndex, + "The master guard must sit inside the telemetry subquery; ahead of it, it would gate the ring buffer too."); + } + + /// + /// The telemetry event carries the USER database's name, and it has to win over the connection's. + /// + /// That arm is read while connected to master, so CurrentDatabaseName is + /// "master" for every row of it. Taking the connection's database would stamp every deadlock on + /// the server as master's — turning the one source that spans all fifty databases into fifty rows + /// about the wrong one. + /// + [Fact] + public void TheTelemetryArmProjectsTheEventsOwnDatabaseName() + { + /* Asserted as a substring, not a regex: the value under test contains double quotes, and the + escaping needed to match it in a pattern is exactly the kind of thing that makes a guard pass + for the wrong reason. The first version of this assertion did — it matched the doubled quotes + of the C# verbatim literal rather than the single quotes of the runtime string. */ + Assert.Contains( + "source_database_name = tel.evt.value('(/event/data[@name=\"database_name\"]/value)[1]', 'nvarchar(128)')", + AzureSql, + StringComparison.Ordinal); + } + + /// + /// Every arm of every engine projects the column, so one reader serves them all. A union whose arms + /// disagree on width does not compile; one whose arms disagree on ORDER compiles and silently swaps + /// two values, which is the failure this asserts against. + /// + [Theory] + [InlineData(true)] + [InlineData(false)] + public void EveryProjectionCarriesTheColumn_SoOneReaderServesBoth(bool azure) + { + var sql = azure ? AzureSql : ServerScopedSql; + + var occurrences = Regex.Matches(sql, @"source_database_name\s*=").Count; + + Assert.True(occurrences >= 1, "The projection lost source_database_name — ReadAsync reads it by name and would go silently null."); + } + + /// + /// The ring-buffer arms project it as a TYPED null. An untyped NULL in the first arm of a union + /// takes its type from that arm, which on some paths is int — and the telemetry arm's + /// nvarchar then fails to convert at runtime, on Azure only, where nothing we own would notice. + /// + [Fact] + public void TheNullArmsAreTypedNulls() + { + Assert.Contains("source_database_name = CONVERT(nvarchar(128), NULL)", AzureSql, StringComparison.Ordinal); + Assert.Contains("source_database_name = CONVERT(nvarchar(128), NULL)", ServerScopedSql, StringComparison.Ordinal); + } + + /// + /// The column is read by ORDINAL, and it is the last one. + /// + /// The reader contract in this codebase is positional — the test fake throws from + /// GetName deliberately, so that a collector cannot quietly depend on a capability the + /// production readers have and the fixtures do not. My first version read the column by name and + /// passed every SQL assertion here while failing the two existing ReadAsync fixtures. + /// + /// Which ordinal moves with CapturePlanXml, because the victim plan is spliced in at 3 + /// only when it is on — so this is the same conditional the plan column uses, one place along. + /// + [Theory] + [InlineData(false, 3)] + [InlineData(true, 4)] + public void TheColumnIsReadByOrdinal_AndItMovesWithThePlanCapture(bool capturePlan, int expected) + { + var ordinal = typeof(DeadlocksCollector) + .GetMethod("SourceDatabaseNameOrdinal", BindingFlags.NonPublic | BindingFlags.Static)! + .Invoke(null, new object[] { new CollectorContext + { + ServerId = 1, + ServerName = "s", + CollectionTime = new DateTime(2026, 8, 26, 0, 0, 0, DateTimeKind.Utc), + Deltas = new Helpers.RecordingCollectorDeltaCalculator(), + Target = new CollectorTargetInfo { IsAzureSqlDb = true }, + CapturePlanXml = capturePlan, + } }); + + Assert.Equal(expected, ordinal); + } + + /// + /// The behaviour, not only the SQL: a telemetry row carries its own database and that database WINS + /// over the connection's, while a ring-buffer row (NULL in that column) still falls back to it. + /// + /// This is the assertion that matters. The telemetry arm is read while connected to + /// master, so CurrentDatabaseName is "master" for every row of it — and taking + /// the connection's database would report every deadlock on a fifty-database server as + /// master's, which is a more convincing wrong answer than the empty grid it replaced. + /// + [Fact] + public async Task ATelemetryRowsOwnDatabaseWins_AndARingBufferRowStillFallsBack() + { + var context = new CollectorContext + { + ServerId = 1, + ServerName = "azure-server", + CollectionTime = new DateTime(2026, 8, 26, 12, 0, 0, DateTimeKind.Utc), + Deltas = new Helpers.RecordingCollectorDeltaCalculator(), + Target = new CollectorTargetInfo { IsAzureSqlDb = true }, + /* What the per-database loop stamps while reading the telemetry arm from master. */ + CurrentDatabaseName = "master", + }; + + var deadlockTime = new DateTime(2026, 8, 26, 12, 13, 15, DateTimeKind.Utc); + + using var reader = new Helpers.FakeCollectorDataReader( + /* telemetry arm: the event named the user database */ + new object[] { deadlockTime, "process20a9deb0478", "", "AppDb" }, + /* ring-buffer arm: no source database of its own */ + new object[] { deadlockTime, "process20a9deb0479", "", DBNull.Value }); + + var rows = await DeadlocksCollector.Instance.ReadAsync(reader, context, CancellationToken.None); + + Assert.Equal(2, rows.Count); + Assert.Equal("AppDb", rows[0].DatabaseName); + Assert.Equal("master", rows[1].DatabaseName); + } + + /// + /// Both arms honour the watermark. An unfiltered telemetry arm would re-read the whole blob every cycle + /// and re-insert every deadlock it has ever held — the collector would look like it was working + /// perfectly while multiplying its own history. + /// + [Fact] + public void BothArmsFilterOnTheCutoff() + => Assert.True(Regex.Matches(AzureSql, @"> @cutoff_time").Count >= 2, + "One arm of the Azure union does not filter on @cutoff_time and would re-read its whole source every cycle."); +} diff --git a/Lite.Tests/BlockingTrendEmptyToolTests.cs b/Lite.Tests/BlockingTrendEmptyToolTests.cs new file mode 100644 index 000000000..bc950ab67 --- /dev/null +++ b/Lite.Tests/BlockingTrendEmptyToolTests.cs @@ -0,0 +1,282 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_blocking_trend / get_deadlock_trend (#2485), twins of Darling's. Both used to serialize +/// { server, hours_back, trend } unconditionally, so a server that never blocked and a server that never +/// collected produced the SAME bytes — and an MCP client, which has only the JSON, can reasonably read the +/// first as "there is no data" when the true answer is "no, and that is good news". +/// +/// The denominator is what makes the empty answer honest, and it cannot come from the read itself: these +/// are EDGE tables, where an absent capture and a capture that found nothing are both an absence of rows. It +/// comes from collection_log, which records a SUCCESS with zero rows for a collector that ran and stored +/// nothing. +/// +/// The messages are Darling's word for word. A user moving between the SKUs must not be told a different +/// story about the same state — the parity half of #2485. +/// +public sealed class BlockingTrendEmptyToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "BlockingTrendSrv"; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private DuckDBConnection? _seedConn; + private long _nextId = 810000; + + public BlockingTrendEmptyToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-blocktrend-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + + /* Derived, not stored -- seeding under a hardcoded id would write rows the tool looks past. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task NeverCollected_IsNotReportedAsAnAllClear() + { + var service = new LocalDataService(_duckDb); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("NOT an all-clear", text, StringComparison.Ordinal); + Assert.Contains("EVER", text, StringComparison.Ordinal); + Assert.Equal(0, root.GetProperty("hints").GetProperty("capture_count").GetInt64()); + } + + [Fact] + public async Task CollectedBeforeButNotInTheWindow_IsAGap_NotANeverCollected() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("blocked_process_report", DateTime.UtcNow.AddHours(-48)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 1)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("NOT an all-clear", text, StringComparison.Ordinal); + Assert.Contains("widen hours_back", text, StringComparison.Ordinal); + + /* Same status as the never-collected case and it must not reach for the same word. */ + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + } + + [Fact] + public async Task CapturesRanAndSawNothing_IsAGenuineAllClear_AndSaysSoDifferently() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("blocked_process_report", DateTime.UtcNow.AddMinutes(-10)); + await SeedRunAsync("dmv_blocking_snapshot", DateTime.UtcNow.AddMinutes(-10)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("genuine all-clear", text, StringComparison.Ordinal); + + /* Same zero rows as both other cases, and it must share wording with neither. */ + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + Assert.DoesNotContain("NOT an all-clear", text, StringComparison.Ordinal); + + /* "No blocking" means something different across two captures than across two hundred, and the + caller cannot supply that number itself. */ + var hints = root.GetProperty("hints"); + Assert.Equal(2, hints.GetProperty("capture_count").GetInt64()); + Assert.Equal(2, hints.GetProperty("captures").GetArrayLength()); + } + + /// + /// The two-capture-path guard. An RDS instance cannot run the blocked-process-report XE session, so its + /// blocking arrives entirely through the DMV snapshot the trend falls back to — and a probe counting only + /// blocked_process_report runs would tell that server its blocking has never been captured. + /// + [Fact] + public async Task DmvOnlyCapture_StillCountsAsHavingLooked() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("dmv_blocking_snapshot", DateTime.UtcNow.AddMinutes(-10)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + Assert.Contains("genuine all-clear", root.GetProperty("message").GetString()!, StringComparison.Ordinal); + Assert.Equal(1, root.GetProperty("hints").GetProperty("capture_count").GetInt64()); + } + + /// + /// A blocking capture is NOT a deadlock capture. Without this the deadlock probe would be satisfied by the + /// neighbouring collector's runs, and a server with the deadlocks collector switched off would be told its + /// empty deadlock trend is an all-clear. + /// + [Fact] + public async Task DeadlockTrend_IsNotSatisfiedByABlockingCapture() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("blocked_process_report", DateTime.UtcNow.AddMinutes(-10)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetDeadlockTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + Assert.Contains("EVER", root.GetProperty("message").GetString()!, StringComparison.Ordinal); + } + + [Fact] + public async Task DeadlockTrend_WithItsOwnCapture_IsAGenuineAllClear() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("deadlocks", DateTime.UtcNow.AddMinutes(-10)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetDeadlockTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("genuine all-clear", text, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + Assert.Equal(1, root.GetProperty("hints").GetProperty("capture_count").GetInt64()); + } + + /// + /// A run that FAILED is not a run that looked. Counting a PERMISSIONS row would manufacture an all-clear + /// out of a collector that never saw the window — the failure this whole change exists to prevent, arrived + /// at from the other direction. + /// + [Fact] + public async Task AFailedRunDoesNotCountAsHavingLooked() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("blocked_process_report", DateTime.UtcNow.AddMinutes(-10), status: "PERMISSIONS"); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + Assert.Contains("EVER", root.GetProperty("message").GetString()!, StringComparison.Ordinal); + } + + [Fact] + public async Task ARealEventStillReturnsTheTrendPayload() + { + var service = new LocalDataService(_duckDb); + await SeedRunAsync("blocked_process_report", DateTime.UtcNow.AddMinutes(-10)); + await SeedBlockedProcessReportAsync(DateTime.UtcNow.AddMinutes(-5)); + + var root = JsonDocument.Parse( + await McpBlockingTools.GetBlockingTrend(service, _serverManager, ServerName, 4)).RootElement; + + Assert.False(root.TryGetProperty("status", out _)); + Assert.Equal(1, root.GetProperty("trend").GetArrayLength()); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + /// A collector run that SUCCEEDED and stored nothing — the row that makes an empty trend + /// interpretable, and the one the edge tables can never produce. + private async Task SeedRunAsync(string collector, DateTime collectionTimeUtc, string status = "SUCCESS") + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO collection_log + (log_id, server_id, server_name, collector_name, collection_time, + duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = collector }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = 100 }); + cmd.Parameters.Add(new DuckDBParameter { Value = status }); + cmd.Parameters.Add(new DuckDBParameter { Value = DBNull.Value }); + cmd.Parameters.Add(new DuckDBParameter { Value = 0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 80 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 20 }); + await cmd.ExecuteNonQueryAsync(); + } + + private async Task SeedBlockedProcessReportAsync(DateTime eventTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO blocked_process_reports + (blocked_report_id, collection_time, server_id, server_name, event_time, database_name, + blocked_spid, blocking_spid, wait_time_ms, lock_mode, blocked_sql_text, blocking_sql_text, + blocked_process_report_xml, contentious_object) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13, $14)"; + var naive = DateTime.SpecifyKind(eventTimeUtc, DateTimeKind.Unspecified); + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = "AppDb" }); + cmd.Parameters.Add(new DuckDBParameter { Value = 55 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 60 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 8000L }); + cmd.Parameters.Add(new DuckDBParameter { Value = "X" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "SELECT 1" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "UPDATE Orders SET Total = 1" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "dbo.Orders" }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/CollectionHealthLatestNoteTests.cs b/Lite.Tests/CollectionHealthLatestNoteTests.cs index 0da5ed983..419d5abaa 100644 --- a/Lite.Tests/CollectionHealthLatestNoteTests.cs +++ b/Lite.Tests/CollectionHealthLatestNoteTests.cs @@ -12,6 +12,7 @@ using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; using PerformanceMonitorLite.Database; using PerformanceMonitorLite.Services; using Xunit; @@ -191,6 +192,56 @@ Seeded so the OTHER collector's note is both newer and lexicographically greater Assert.Equal("zzz a different collector's note", health.Single(h => h.CollectorName == "wait_stats").LastNote); } + /// + /// #2460, Lite half: the two duration statistics, against a REAL DuckDB, on the population that + /// motivated them. A source pin cannot tell PERCENTILE_DISC from AVG — both are valid SQL returning + /// one number — and the whole finding is that one of those numbers describes no run that ever ran. + /// + /// The fixture is prod-sql-use2-multi-49's query_store at 1/11.55 scale, same 83/17 shape: 83 + /// runs carrying the empty-enumeration note at the 36 ms prod-sql-use2-alpha-01 measurably pays for it, + /// and 17 productive runs at the ~80,933 ms the store's own numbers force. The assertion that matters + /// is the last pair: the MEAN sits comfortably inside a 60,000 ms sweep budget while a heavy run costs + /// more than the whole budget by itself, which is the sentence the mean alone could never produce. + /// Darling.Tests pins the identical fixture against live Postgres — two engines, one answer. + /// + [Fact] + public async Task TheDurationStatistics_SplitABimodalCollectorTheMeanBlendsAway() + { + var service = new LocalDataService(_duckDb); + + for (var i = 0; i < 83; i++) + { + await SeedDurationAsync("query_store", MinutesAgo(600 - i), 36, + EnumeratedCollectorDriver.EmptyEnumerationMessage); + } + + for (var i = 0; i < 17; i++) + { + await SeedDurationAsync("query_store", MinutesAgo(400 - i), 80_933, null); + } + + /* A failed run with no duration recorded: the new aggregates must ignore it exactly as AVG + already does, or a collector that errors occasionally reports a NULL p95 and silently falls + back to its mean. Verified here rather than assumed, because the two engines had to agree. */ + await SeedDurationAsync("query_store", MinutesAgo(300), null, null, "ERROR"); + + var row = await ReadAsync(service, "query_store"); + + Assert.Equal(101, row.TotalRuns); + Assert.Equal(83, row.NoteCount); + + Assert.Equal(13_788.49, row.AvgDurationMs, 2); + Assert.Equal(80_933, row.MaxDurationMs, 3); + Assert.Equal(80_933, row.P95DurationMs, 3); + + /* The finding, as an assertion: one number says this collector fits a body four times over, the + other says one run of it does not fit at all, and both are honest about the same 101 runs. */ + Assert.True(row.AvgDurationMs < SweepPressureClassifier.SweepBudgetMs, + $"the mean was {row.AvgDurationMs} ms"); + Assert.True(row.P95DurationMs > SweepPressureClassifier.SweepBudgetMs, + $"a heavy run was {row.P95DurationMs} ms"); + } + /* ── helpers ── */ private async Task ReadAsync(LocalDataService service, string collector) => @@ -213,7 +264,12 @@ private async Task SeedConnectionAsync() return _seedConn; } - private async Task SeedAsync(string collector, DateTime collectionTimeUtc, string status, string? message) + private Task SeedAsync(string collector, DateTime collectionTimeUtc, string status, string? message) => + SeedDurationAsync(collector, collectionTimeUtc, 100, message, status); + + /// #2460: the same seed with the run's own duration_ms spelled out, null included. + private async Task SeedDurationAsync( + string collector, DateTime collectionTimeUtc, int? durationMs, string? message, string status = "SUCCESS") { using var readLock = _duckDb.AcquireReadLock(); var connection = await SeedConnectionAsync(); @@ -228,7 +284,7 @@ INSERT INTO collection_log cmd.Parameters.Add(new DuckDBParameter { Value = "TestSrv" }); cmd.Parameters.Add(new DuckDBParameter { Value = collector }); cmd.Parameters.Add(new DuckDBParameter { Value = collectionTimeUtc }); - cmd.Parameters.Add(new DuckDBParameter { Value = 100 }); + cmd.Parameters.Add(new DuckDBParameter { Value = (object?)durationMs ?? DBNull.Value }); cmd.Parameters.Add(new DuckDBParameter { Value = status }); cmd.Parameters.Add(new DuckDBParameter { Value = (object?)message ?? DBNull.Value }); cmd.Parameters.Add(new DuckDBParameter { Value = 10 }); diff --git a/Lite.Tests/CollectionLogToolTests.cs b/Lite.Tests/CollectionLogToolTests.cs new file mode 100644 index 000000000..30d79b527 --- /dev/null +++ b/Lite.Tests/CollectionLogToolTests.cs @@ -0,0 +1,181 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_collection_log (#2484), the twin of Darling's. Lite gained the tool in the same change that +/// added it to Darling rather than being parked on the divergence ratchet, so its behaviour is pinned here +/// for the same reasons -- and because the SKUs make PROMISES to each other that only a test can hold: +/// the same two words for the two kinds of empty, and store_duration_ms as one field name over two different +/// storage engines. +/// +/// Written at the TOOL level, not the reader level. The reader is the easy half; the cap contract and +/// the truncation signal live in the tool, and a reader-level test would have missed both. +/// +public sealed class CollectionLogToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "CollectionLogSrv"; + + /* + Lite does not store a server id -- ServerResolver DERIVES it from the storage name, so the seeded + rows have to be written under the same derived value the tool will resolve to. Hardcoding an id + here would seed rows the tool then looks straight past, and the test would pass its + never-collected assertion for entirely the wrong reason. + */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public CollectionLogToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-collog-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task EmptyWindow_AndNeverCollected_AreDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + /* Registered, never collected: a fault, and it must not be described as an empty window. */ + var never = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 24, 200); + var neverRoot = JsonDocument.Parse(never).RootElement; + Assert.Equal("unavailable", neverRoot.GetProperty("status").GetString()); + var neverText = neverRoot.GetProperty("message").GetString()!; + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + Assert.DoesNotContain("widen", neverText, StringComparison.OrdinalIgnoreCase); + + /* Collected, but outside the asked-for window: a true negative, and widening IS the move. */ + await SeedLogAsync("query_store", DateTime.UtcNow.AddHours(-48)); + + var quiet = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 1, 200); + var quietRoot = JsonDocument.Parse(quiet).RootElement; + Assert.Equal("empty", quietRoot.GetProperty("status").GetString()); + var quietText = quietRoot.GetProperty("message").GetString()!; + Assert.DoesNotContain("EVER", quietText, StringComparison.Ordinal); + Assert.Contains("widen", quietText, StringComparison.OrdinalIgnoreCase); + } + + [Fact] + public async Task TruncationIsObserved_NotInferredFromTheRowCount() + { + var service = new LocalDataService(_duckDb); + await SeedLogAsync("query_store", DateTime.UtcNow.AddMinutes(-10)); + await SeedLogAsync("deadlocks", DateTime.UtcNow.AddMinutes(-5)); + + /* Exactly the cap, nothing beyond it. Comparing count to limit reports truncated here, wrongly. */ + var exact = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 24, 2); + var exactRoot = JsonDocument.Parse(exact).RootElement; + Assert.Equal(2, exactRoot.GetProperty("run_count").GetInt32()); + Assert.False(exactRoot.GetProperty("truncated").GetBoolean()); + + /* Under the cap: there really is more, and it says so. */ + var capped = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 24, 1); + var cappedRoot = JsonDocument.Parse(capped).RootElement; + Assert.Equal(1, cappedRoot.GetProperty("run_count").GetInt32()); + Assert.True(cappedRoot.GetProperty("truncated").GetBoolean()); + } + + [Fact] + public async Task AnOutOfRangeCap_IsRefused_NotSilentlyClamped() + { + var service = new LocalDataService(_duckDb); + await SeedLogAsync("query_store", DateTime.UtcNow.AddMinutes(-10)); + + var tooBig = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 24, 5000); + Assert.Contains("exceeds maximum of", tooBig, StringComparison.Ordinal); + Assert.Contains("1000", tooBig, StringComparison.Ordinal); + } + + [Fact] + public async Task ThePayloadSplitsDuration_AndNamesTheStoreFieldTheWayDarlingDoes() + { + var service = new LocalDataService(_duckDb); + await SeedLogAsync("query_store", DateTime.UtcNow.AddMinutes(-10)); + + var hit = await McpHealthTools.GetCollectionLog(service, _serverManager, ServerName, 24, 200); + var run = JsonDocument.Parse(hit).RootElement.GetProperty("runs")[0]; + + Assert.Equal(80, run.GetProperty("sql_duration_ms").GetInt32()); + + /* Lite's column is duckdb_duration_ms. The FIELD is store_duration_ms on both SKUs, because the + storage engine differs and the question the caller is asking does not. */ + Assert.Equal(20, run.GetProperty("store_duration_ms").GetInt32()); + Assert.False(run.TryGetProperty("duckdb_duration_ms", out _)); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedLogAsync(string collector, DateTime collectionTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO collection_log + (log_id, server_id, server_name, collector_name, collection_time, + duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = collector }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = 100 }); + cmd.Parameters.Add(new DuckDBParameter { Value = "SUCCESS" }); + cmd.Parameters.Add(new DuckDBParameter { Value = DBNull.Value }); + cmd.Parameters.Add(new DuckDBParameter { Value = 10 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 80 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 20 }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/CollectionResetGateTests.cs b/Lite.Tests/CollectionResetGateTests.cs new file mode 100644 index 000000000..509ea8849 --- /dev/null +++ b/Lite.Tests/CollectionResetGateTests.cs @@ -0,0 +1,168 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Diagnostics; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using PerformanceMonitorLite.Database; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// The gate that stops a size-triggered reset deleting the database out from under a running collection +/// (#2594). +/// +/// The field failure this encodes. Opening a server tab calls +/// RunAllCollectorsForServerAsync on a bare Task.Run, sequenced against nothing. When the store +/// crossed 512 MB while such a collection was running, the reset deleted and recreated monitor.duckdb +/// underneath it — index_object_stats was five seconds into a fifty-five second run — and the +/// collection's final collection_log insert failed with Table with name collection_log does not +/// exist. Nothing was lost on disk and a restart cleared it, which is the signature of stale in-process +/// state rather than a damaged store. +/// +/// These run serially: the gate is process-wide static state, so concurrent tests would see each +/// other's registrations. +/// +[Collection("CollectionResetGate")] +public class CollectionResetGateTests : IDisposable +{ + public CollectionResetGateTests() => CollectionResetGate.ResetForTests(); + + public void Dispose() => CollectionResetGate.ResetForTests(); + + [Fact] + public async Task AReset_DoesNotStartWhileACollectionIsRunning() + { + using var collection = await CollectionResetGate.BeginCollectionAsync(); + + Assert.Equal(1, CollectionResetGate.CollectionsInFlight); + + /* The whole point: the reset must not get the gate here. It waits for the drain timeout and then + defers, which the caller reports and retries on the next archival check. */ + var reset = await WithShortDrain(() => CollectionResetGate.TryBeginResetAsync()); + + Assert.Null(reset); + } + + [Fact] + public async Task AReset_ProceedsOnceTheCollectionFinishes() + { + var collection = await CollectionResetGate.BeginCollectionAsync(); + collection.Dispose(); + + Assert.Equal(0, CollectionResetGate.CollectionsInFlight); + + using var reset = await CollectionResetGate.TryBeginResetAsync(); + + Assert.NotNull(reset); + } + + [Fact] + public async Task AResetWaits_RatherThanFailingImmediately_WhenACollectionEndsShortly() + { + var collection = await CollectionResetGate.BeginCollectionAsync(); + + /* Released while the reset is already waiting, which is the ordinary case: a collection finishes + and the reset proceeds without anyone retrying. */ + _ = Task.Run(async () => + { + await Task.Delay(TimeSpan.FromMilliseconds(300)); + collection.Dispose(); + }); + + var stopwatch = Stopwatch.StartNew(); + using var reset = await CollectionResetGate.TryBeginResetAsync(); + stopwatch.Stop(); + + Assert.NotNull(reset); + Assert.True( + stopwatch.Elapsed >= TimeSpan.FromMilliseconds(200), + $"the reset returned in {stopwatch.ElapsedMilliseconds} ms, so it did not actually wait for the " + + "collection to finish."); + } + + /// + /// A collection arriving mid-reset must wait, not start against a database that is about to be deleted. + /// This is the other half of the race and the half that is easy to forget: gating only the reset leaves + /// the tab-open path free to start one millisecond before the file disappears. + /// + [Fact] + public async Task ACollection_DoesNotStartWhileAResetIsRunning() + { + var reset = await CollectionResetGate.TryBeginResetAsync(); + Assert.NotNull(reset); + + var started = CollectionResetGate.BeginCollectionAsync(); + + var raced = await Task.WhenAny(started, Task.Delay(TimeSpan.FromMilliseconds(400))); + Assert.NotSame(started, raced); + + reset!.Dispose(); + + using var collection = await started.WaitAsync(TimeSpan.FromSeconds(5)); + Assert.Equal(1, CollectionResetGate.CollectionsInFlight); + } + + /// + /// Collections must NOT serialise against each other. The background sweep and a tab-open sweep run + /// concurrently by design, and turning that into a queue would be a behaviour change wearing a bug fix. + /// + [Fact] + public async Task Collections_RunConcurrentlyWithEachOther() + { + using var first = await CollectionResetGate.BeginCollectionAsync(); + + var second = CollectionResetGate.BeginCollectionAsync().WaitAsync(TimeSpan.FromSeconds(2)); + + using var acquired = await second; + + Assert.Equal(2, CollectionResetGate.CollectionsInFlight); + } + + [Fact] + public async Task DisposingACollectionTwice_DoesNotDriveTheCounterNegative() + { + var collection = await CollectionResetGate.BeginCollectionAsync(); + collection.Dispose(); + collection.Dispose(); + + Assert.Equal(0, CollectionResetGate.CollectionsInFlight); + + /* A negative counter would let a reset run while a real collection was still writing. */ + using var other = await CollectionResetGate.BeginCollectionAsync(); + Assert.Equal(1, CollectionResetGate.CollectionsInFlight); + } + + /// + /// The production drain timeout is three minutes, which no test should sit through. This shortens the + /// wait by racing the call rather than by making the timeout configurable — a knob that exists only for + /// tests is a knob somebody eventually sets in production. + /// + private static async Task WithShortDrain(Func> attempt) + { + var call = attempt(); + var finished = await Task.WhenAny(call, Task.Delay(TimeSpan.FromMilliseconds(500))); + + if (finished == call) + { + return await call; + } + + /* Still draining after the grace period, which is the assertion the caller wants: it did not take + the gate. The underlying call abandons itself when the collection scope is disposed in teardown. */ + return null; + } +} + +[CollectionDefinition("CollectionResetGate", DisableParallelization = true)] +public sealed class CollectionResetGateCollection +{ +} diff --git a/Lite.Tests/CollectorGateSurfacePinTests.cs b/Lite.Tests/CollectorGateSurfacePinTests.cs index 6e8c34b9a..9da05e62e 100644 --- a/Lite.Tests/CollectorGateSurfacePinTests.cs +++ b/Lite.Tests/CollectorGateSurfacePinTests.cs @@ -23,8 +23,10 @@ namespace Lite.Tests; /// the running_jobs/job_history/agent_status msdb gate) was silently ignored by Darling. That layer is gone. /// These pins assert: /// -/// each moved gate's AppliesTo truth table (Azure SQL DB / AWS RDS / no-msdb / pre-2016), so a -/// gate can't silently regress; +/// each moved gate's AppliesTo truth table (Azure SQL DB / AWS RDS / pre-2016), so a +/// gate can't silently regress. msdb access is deliberately NOT among them since #2559 — it is a grant +/// rather than an engine capability, so the Agent collectors attempt and fail into PERMISSIONS instead of +/// gating off on a cached probe; /// that the by-name surface Lite dispatches through () /// agrees with the definition's own AppliesTo that Darling's runner calls — the proof both SKUs now /// gate identically off ONE surface; @@ -62,23 +64,44 @@ public void ServerConfig_AppliesTo_SkipsOnlyAzureSqlDb() } /// - /// #2150 field report: this fired 11x consecutive on an Azure SQL DB elastic pool with error 262, - /// "VIEW DATABASE PERFORMANCE STATE permission denied in database 'tempdb'". The query reads - /// tempdb.sys.dm_db_file_space_usage three-part, which a non-administrative login on Azure - /// SQL DB cannot be granted, so the collector could only ever fail there. - /// Managed Instance must KEEP collecting — it has a real tempdb — which is why this asserts - /// both directions rather than just the skip. + /// tempdb_stats applies EVERYWHERE, Azure SQL Database included (#2512). + /// + /// This assertion was flipped, and it is worth being precise about what it used to pin. + /// It was written for the #2150 field report — 11x consecutive on an Azure SQL DB elastic pool with + /// error 262, "VIEW DATABASE PERFORMANCE STATE permission denied in database 'tempdb'" — and it pinned + /// TWO different claims in one Assert.False. The first, that the three-part + /// tempdb.sys.dm_db_file_space_usage reference cannot be served on Azure SQL Database, was + /// checkable and is false: the collector's SQL runs verbatim on GP_S_Gen5_2 and HS_S_Gen5_2 and returns + /// real, moving numbers (see for the measurement). The + /// second, that a login might not be able to READ it, is true — but it is a property of the login, not + /// of the tier, so it belongs to the fault classifier + /// (, which now covers 262) and not to a gate + /// that denies every properly-permissioned Azure target to spare the one that is not. + /// + /// The other half of this pin never changed and is the reason it survives rather than being + /// deleted. Managed Instance was never gated — it has a real tempdb and full DMV access — and the + /// original comment says outright that asserting BOTH directions is what stops an "anything Azure" + /// gate creeping in. That risk runs the other way now: this must not be re-narrowed to Azure SQL DB + /// later on the strength of the stale doc comment, so both directions still get asserted. /// [Fact] - public void TempDbStats_AppliesTo_SkipsOnlyAzureSqlDb() + public void TempDbStats_AppliesTo_EveryTarget_IncludingAzureSqlDb() { - Assert.False(TempDbStatsCollector.Instance.AppliesTo(AzureSqlDb)); /* error 262 in tempdb */ + /* #2512: the gate is gone. The DMVs bind and return real data on both Azure SQL DB tiers. */ + Assert.True(TempDbStatsCollector.Instance.AppliesTo(AzureSqlDb)); + /* Never gated, and must stay that way — MI has a real tempdb. */ Assert.True(TempDbStatsCollector.Instance.AppliesTo(AzureMi)); Assert.True(TempDbStatsCollector.Instance.AppliesTo(AwsRds)); Assert.True(TempDbStatsCollector.Instance.AppliesTo(OnPrem2016)); Assert.True(TempDbStatsCollector.Instance.AppliesTo(OnPrem2014)); Assert.True(TempDbStatsCollector.Instance.AppliesTo(NoMsdb)); Assert.True(TempDbStatsCollector.Instance.AppliesTo(Unknown)); + + /* The surface the runners actually call — the composed engine gate — must agree, or Darling + would still skip what Lite now runs. AppliesTo alone cannot see that half. */ + Assert.True(CollectorCatalog.AppliesTo(TempDbStatsCollector.Instance, AzureSqlDb)); + Assert.True(CollectorCatalog.AppliesTo(TempDbStatsCollector.Instance, AzureMi)); + Assert.True(CollectorCatalog.AppliesTo(TempDbStatsCollector.Instance.Name, AzureSqlDb)); } [Fact] @@ -112,21 +135,21 @@ public void QueryStore_AppliesTo_SkipsOnlyPreSql2016OnPrem() } [Fact] - public void RunningJobs_AppliesTo_SkipsAzureSqlDbRdsAndNoMsdb() + public void RunningJobs_AppliesTo_SkipsAzureSqlDbAndRds_ButAttemptsWithoutMsdb() { Assert.False(RunningJobsCollector.Instance.AppliesTo(AzureSqlDb)); Assert.False(RunningJobsCollector.Instance.AppliesTo(AwsRds)); /* joins msdb.dbo.syssessions */ - Assert.False(RunningJobsCollector.Instance.AppliesTo(NoMsdb)); + Assert.True(RunningJobsCollector.Instance.AppliesTo(NoMsdb)); // #2559: reported, not dispatched on Assert.True(RunningJobsCollector.Instance.AppliesTo(AzureMi)); Assert.True(RunningJobsCollector.Instance.AppliesTo(OnPrem2016)); Assert.True(RunningJobsCollector.Instance.AppliesTo(Unknown)); } [Fact] - public void JobHistory_AppliesTo_SkipsAzureSqlDbAndNoMsdb_ButNotRds() + public void JobHistory_AppliesTo_SkipsAzureSqlDb_ButNotRdsAndNotNoMsdb() { Assert.False(JobHistoryCollector.Instance.AppliesTo(AzureSqlDb)); - Assert.False(JobHistoryCollector.Instance.AppliesTo(NoMsdb)); + Assert.True(JobHistoryCollector.Instance.AppliesTo(NoMsdb)); // #2559: reported, not dispatched on Assert.True(JobHistoryCollector.Instance.AppliesTo(AwsRds)); /* never touches syssessions */ Assert.True(JobHistoryCollector.Instance.AppliesTo(AzureMi)); Assert.True(JobHistoryCollector.Instance.AppliesTo(OnPrem2016)); @@ -134,11 +157,11 @@ public void JobHistory_AppliesTo_SkipsAzureSqlDbAndNoMsdb_ButNotRds() } [Fact] - public void AgentStatus_AppliesTo_SkipsAzureSqlDbRdsAndNoMsdb() + public void AgentStatus_AppliesTo_SkipsAzureSqlDbAndRds_ButAttemptsWithoutMsdb() { Assert.False(AgentStatusCollector.Instance.AppliesTo(AzureSqlDb)); Assert.False(AgentStatusCollector.Instance.AppliesTo(AwsRds)); /* no sys.dm_server_services */ - Assert.False(AgentStatusCollector.Instance.AppliesTo(NoMsdb)); + Assert.True(AgentStatusCollector.Instance.AppliesTo(NoMsdb)); // #2559: reported, not dispatched on Assert.True(AgentStatusCollector.Instance.AppliesTo(AzureMi)); Assert.True(AgentStatusCollector.Instance.AppliesTo(OnPrem2016)); Assert.True(AgentStatusCollector.Instance.AppliesTo(Unknown)); @@ -185,10 +208,12 @@ public void CatalogGate_UnknownCollectorName_IsNotGated() } [Fact] - public void HasMsdbAccess_DefaultsToTrue_SoUnknownTargetsStillAttemptAgentCollectors() + public void HasMsdbAccess_DefaultsToTrue_AndTheAgentCollectorsAttemptRegardless() { - /* The probe returns NULL ⇒ assume access; every bare CollectorTargetInfo must mirror that, or the - three Agent collectors would silently gate off on the unknown path. */ + /* The probe returns NULL ⇒ assume access, and every bare CollectorTargetInfo mirrors that. Since + #2559 the three Agent collectors attempt whatever this says, so the default no longer decides + dispatch — but it is still the value reported on a connection surface, and it staying true is + what keeps an unclassified target from being described as having no msdb access. */ Assert.True(new CollectorTargetInfo().HasMsdbAccess); Assert.True(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo())); Assert.True(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo())); diff --git a/Lite.Tests/CollectorPayloadArityTests.cs b/Lite.Tests/CollectorPayloadArityTests.cs new file mode 100644 index 000000000..928a831cd --- /dev/null +++ b/Lite.Tests/CollectorPayloadArityTests.cs @@ -0,0 +1,145 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Linq; +using System.Reflection; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Every collector writes exactly as many payload values as it declares payload columns. +/// +/// The runtime already checks this — and that was not enough. EndPayload throws when the +/// counts disagree, but only when the collector actually RUNS against a real target. #2599 added +/// database_name to pg_extension_availability's columns, its SQL, its row and its reader, and +/// missed WritePayload. The collector is DAILY, so it did not re-run after the upgrade that shipped +/// the change, and the store simply had no rows from it — which looks exactly like a quiet server. It was +/// found by pointing the service at a live self-hosted PostgreSQL, three schema versions later. +/// +/// A build-time check turns that into a compile-cycle failure. The columns and the writes are two +/// halves of one statement about the table; nothing but a test holds them together. +/// +public class CollectorPayloadArityTests +{ + /// Counts writes and accepts anything, so a collector's own value types never matter here. + private sealed class CountingWriter : ICollectorRowWriter + { + public int Count { get; private set; } + + private ICollectorRowWriter Bump() + { + Count++; + return this; + } + + public ICollectorRowWriter Value(string? value) => Bump(); + public ICollectorRowWriter Value(long value) => Bump(); + public ICollectorRowWriter Value(long? value) => Bump(); + public ICollectorRowWriter Value(int value) => Bump(); + public ICollectorRowWriter Value(int? value) => Bump(); + public ICollectorRowWriter Value(short value) => Bump(); + public ICollectorRowWriter Value(short? value) => Bump(); + public ICollectorRowWriter Value(double value) => Bump(); + public ICollectorRowWriter Value(double? value) => Bump(); + public ICollectorRowWriter Value(decimal value) => Bump(); + public ICollectorRowWriter Value(decimal? value) => Bump(); + public ICollectorRowWriter Value(bool value) => Bump(); + public ICollectorRowWriter Value(bool? value) => Bump(); + public ICollectorRowWriter Value(DateTime value) => Bump(); + public ICollectorRowWriter Value(DateTime? value) => Bump(); + public ICollectorRowWriter Value(byte[]? value) => Bump(); + public ICollectorRowWriter NullValue() => Bump(); + + public void BeginPayload() => Count = 0; + public void EndPayload(int expectedPayloadColumns) { } + } + + [Fact] + public void EveryCollectorWritesAsManyValuesAsItDeclaresColumns() + { + var mismatches = new List(); + var skipped = new List(); + + foreach (var schema in CollectorCatalog.All) + { + var write = schema.GetType().GetMethods(BindingFlags.Public | BindingFlags.Instance) + .FirstOrDefault(m => m.Name == "WritePayload" && m.GetParameters().Length == 3); + + if (write is null) + { + skipped.Add(schema.Name); + continue; + } + + var rowType = write.GetParameters()[0].ParameterType; + + /* A default row is enough: this counts WRITES, not values. A collector that branched on row + CONTENT to decide how many values to write would already be broken against the fixed column + list, so the default is not a weaker test — it is the same test. */ + var row = rowType.IsValueType ? Activator.CreateInstance(rowType) : null; + + if (row is null) + { + skipped.Add(schema.Name); + continue; + } + + var writer = new CountingWriter(); + writer.BeginPayload(); + + try + { + write.Invoke(schema, new[] { row, writer, MakeContext() }); + } + catch (TargetInvocationException) + { + /* A default row is not valid for every collector - the delta-based SQL Server ones + dereference fields the default leaves null. That is the PROBE being too blunt for + those, not a defect in them, so they count as uncovered rather than as failures. The + coverage assertion below is what keeps that from quietly swallowing the catalog. */ + skipped.Add(schema.Name); + continue; + } + + if (writer.Count != schema.PayloadColumns.Count) + { + mismatches.Add( + $"{schema.Name}: writes {writer.Count} value(s) but declares " + + $"{schema.PayloadColumns.Count} column(s)"); + } + } + + Assert.True( + mismatches.Count == 0, + "Collector(s) write a different number of payload values than they declare columns. Every row " + + "they produce will be rejected at EndPayload, and a collector that runs rarely can hide that " + + "for a long time (#2599 hid for three schema versions): " + + string.Join(" | ", mismatches)); + + /* The guard is only worth having if it actually covered the catalog. */ + Assert.True( + skipped.Count < CollectorCatalog.All.Count / 4, + $"{skipped.Count} of {CollectorCatalog.All.Count} collectors could not be exercised, so this " + + "guard is covering far less than it appears to: " + string.Join(", ", skipped)); + } + + private static CollectorContext MakeContext() + => new() + { + ServerId = 1, + ServerName = "arity-probe", + CollectionTime = new DateTime(2026, 8, 26, 0, 0, 0, DateTimeKind.Utc), + Deltas = new RecordingCollectorDeltaCalculator(), + Target = new CollectorTargetInfo(), + }; +} diff --git a/Lite.Tests/CrossAppMcpToolInventoryPinTests.cs b/Lite.Tests/CrossAppMcpToolInventoryPinTests.cs index 983236a30..040cafa7e 100644 --- a/Lite.Tests/CrossAppMcpToolInventoryPinTests.cs +++ b/Lite.Tests/CrossAppMcpToolInventoryPinTests.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor Lite. @@ -55,14 +55,35 @@ private static readonly (string Lite, string Darling)[] KnownNamingDrift = // system_health parser tools). A NEW Darling-only tool must be either ported to Lite or added here. private static readonly HashSet KnownLiteMissingMcpTools = new(StringComparer.Ordinal) { - /* The eight PostgreSQL reads. Darling-ONLY by architecture, not "not ported yet", so these are the + /* The PostgreSQL reads. Darling-ONLY by architecture, not "not ported yet", so these are the same kind of entry as get_store_metrics rather than a to-do: Lite has no PostgreSQL target and cannot acquire one (the engine gate never dispatches a PostgreSQL definition there), and Lite does not even create the tables — DuckDbSchemaGenerator.StoredCollectors filters them out, so there is nothing for a Lite twin to read. If Lite ever gains a PostgreSQL target, port these and delete them from here; the ratchet only shrinks. */ "get_pg_wait_stats", + /* #2629: the stock-PostgreSQL counterparts. Same entry, same reason — Lite has no PostgreSQL + target at all, so these are a SKU boundary rather than a porting to-do. */ + "get_pg_wait_sampling", + "get_pg_kernel_stats", + "get_pg_predicate_stats", + "get_pg_index_bloat", + "get_pg_column_stats", + "get_pg_buffer_usage", + "get_pg_extensions", + "get_pg_lock_stats", + "get_pg_write_stats", + "get_pg_server_config", + "get_pg_deadlocks", + "get_pg_wait_trend", + "get_pg_query_duration_trend", + "get_pg_io_trend", + "get_pg_database_trend", + "get_pg_deadlock_detail", + "get_pg_server_config_changes", + "get_pg_replication_stats", "get_pg_top_queries", + "get_pg_plans", "get_pg_wraparound_risk", "get_pg_xmin_horizon", "get_pg_replication_slots", @@ -74,6 +95,38 @@ cannot acquire one (the engine gate never dispatches a PostgreSQL definition the this to Lite would require a PostgreSQL target Lite cannot have, so it belongs here with the rest. */ "get_pg_blocking", + /* get_pg_database_stats (#2539) — the pg_stat_database counters. Same architectural reason as the + eight above: Lite has no PostgreSQL target and cannot acquire one, and DuckDbSchemaGenerator + filters the table out, so there is nothing for a Lite twin to read. */ + "get_pg_database_stats", + + /* get_pg_index_usage (#2541) and get_pg_table_bloat (#2542) - per-index usage and the per-table + bloat estimate. Same architectural reason as the nine above rather than a porting backlog: Lite + has no PostgreSQL target and cannot acquire one, DuckDbSchemaGenerator.StoredCollectors filters + both tables out of every generation loop, and Lite passes engineKind: null explicitly - so there + is no Lite twin for these to be missing FROM. + + Worth being explicit about get_pg_index_usage in particular, because Lite DOES ship + get_index_usage over index_object_stats and the two look like twins. They are not: that one reads + SQL Server DMVs and reports seeks/scans/lookups with lock and latch waits, this one reads + pg_stat_user_indexes and reports the constraint, replica-identity and validity facts that decide + whether a PostgreSQL index can be dropped at all. Conflating them would put T-SQL on a + PostgreSQL path, which is the #2213 class of defect. */ + "get_pg_index_usage", + "get_pg_table_bloat", + + /* get_pg_session_states (#2540) - which sessions are holding a transaction open and which of them + actually pins the xmin horizon. Same architectural reason as every entry above: Lite has no + PostgreSQL target, so there is no Lite twin for this to be missing from. + + Worth naming the near-twin explicitly, because Lite ships get_active_queries and the two sound + alike. They are not the same read. That one is a SQL Server DMV snapshot of what is EXECUTING; + this one is a stored history of what is NOT executing but has a transaction open, which is the + condition SQL Server has no equivalent of - no SQL Server session pins a cluster-wide cleanup + horizon by sitting idle inside a transaction. Porting the name across would put T-SQL on a + PostgreSQL path, which is the #2213 class of defect. */ + "get_pg_session_states", + /* #2068: the store self-metrics read (get_store_metrics) over collect.store_metrics — the central Postgres store measuring ITSELF (hypertable sizes/compression, payload dims, whole-store growth) for capacity forecasting. Darling-ONLY by architecture, not a "not ported yet" item: Lite is a diff --git a/Lite.Tests/CurrentWaitsTrendToolTests.cs b/Lite.Tests/CurrentWaitsTrendToolTests.cs new file mode 100644 index 000000000..845f80055 --- /dev/null +++ b/Lite.Tests/CurrentWaitsTrendToolTests.cs @@ -0,0 +1,155 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_current_waits_trend (#2484), twin of Darling's. Ported in the same change rather than parked +/// on the divergence ratchet. +/// +/// The empty case is the one that matters, and it is sharper here than for most reads: the WRONG +/// answer is the reassuring one. "Nothing was waiting" reads as an all-clear, and a caller who believes it +/// stops looking — so a server the collector has never sampled must not be described that way. +/// +public sealed class CurrentWaitsTrendToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "CurrentWaitsSrv"; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private DuckDBConnection? _seedConn; + private long _nextId = 700000; + + public CurrentWaitsTrendToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-waits-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + + /* Derived, not stored -- seeding under a hardcoded id would write rows the tool looks past. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task NeverSampled_IsNotReportedAsAnAllClear() + { + var service = new LocalDataService(_duckDb); + + var never = await McpHealthTools.GetCurrentWaitsTrend(service, _serverManager, ServerName, 4, null); + var root = JsonDocument.Parse(never).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("NOT an all-clear", text, StringComparison.Ordinal); + Assert.Contains("EVER", text, StringComparison.Ordinal); + } + + [Fact] + public async Task SampledButQuiet_IsAGenuineAllClear_AndSaysSoDifferently() + { + var service = new LocalDataService(_duckDb); + await SeedWaitAsync(DateTime.UtcNow.AddHours(-48), "LCK_M_X", 500, blockingSessionId: 99, database: "AppDb"); + + var clear = await McpHealthTools.GetCurrentWaitsTrend(service, _serverManager, ServerName, 1, null); + var root = JsonDocument.Parse(clear).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("genuine all-clear", text, StringComparison.Ordinal); + + /* Same zero rows as the never-sampled case, and it must not reach for the same word. */ + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + } + + [Fact] + public async Task BothSeriesComeBackTogether_AndOnlyBlockedRowsCountAsBlocked() + { + var service = new LocalDataService(_duckDb); + await SeedWaitAsync(DateTime.UtcNow.AddMinutes(-10), "LCK_M_X", 500, blockingSessionId: 99, database: "AppDb"); + await SeedWaitAsync(DateTime.UtcNow.AddMinutes(-9), "PAGEIOLATCH_SH", 250, blockingSessionId: 0, database: "AppDb"); + + var hit = await McpHealthTools.GetCurrentWaitsTrend(service, _serverManager, ServerName, 4, null); + var root = JsonDocument.Parse(hit).RootElement; + + Assert.Equal(2, root.GetProperty("waiting_tasks").GetArrayLength()); + + /* + The PAGEIOLATCH row waits on IO and blocks on nothing, so it is in the wait series and NOT the + blocked series. That difference is exactly why the two are returned together: the same wait + spike means a resource problem without blocked sessions and contention with them. + */ + var blocked = root.GetProperty("blocked_sessions"); + Assert.Equal(1, blocked.GetArrayLength()); + Assert.Equal("AppDb", blocked[0].GetProperty("database_name").GetString()); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedWaitAsync( + DateTime collectionTimeUtc, string waitType, long waitMs, int blockingSessionId, string database) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO waiting_tasks + (collection_id, collection_time, server_id, server_name, session_id, wait_type, + wait_duration_ms, blocking_session_id, resource_description, database_name) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = 55 }); + cmd.Parameters.Add(new DuckDBParameter { Value = waitType }); + cmd.Parameters.Add(new DuckDBParameter { Value = waitMs }); + cmd.Parameters.Add(new DuckDBParameter { Value = blockingSessionId }); + cmd.Parameters.Add(new DuckDBParameter { Value = DBNull.Value }); + cmd.Parameters.Add(new DuckDBParameter { Value = database }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/DailySummaryRangeToolTests.cs b/Lite.Tests/DailySummaryRangeToolTests.cs new file mode 100644 index 000000000..82c2c2813 --- /dev/null +++ b/Lite.Tests/DailySummaryRangeToolTests.cs @@ -0,0 +1,225 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_daily_summary_range (#2484), the twin of Darling's — the Performance Calendar's month grid, +/// which had no read on either SKU. GetDailySummaryRangeAsync already existed and nothing called it, +/// so get_daily_summary could only ever answer one day. +/// +/// What is worth pinning is not "the aggregate runs" — get_daily_summary proves that off the same SQL. +/// It is the three things only the RANGE form can get wrong: that a missing day means missing COLLECTION +/// rather than a quiet day, that the band is computed per day so two days in one result can disagree, and +/// that the anchor moves the range rather than widening it. +/// +public sealed class DailySummaryRangeToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "CalendarSrv"; + + /* Lite DERIVES its server id from the storage name; a hardcoded one would seed rows the tool looks + straight past and pass the never-collected assertion for the wrong reason. */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public DailySummaryRangeToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-calendar-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task TheCalendar_BandsEachDaySeparately_AndShowsCollectionGapsAsMissingDays() + { + var service = new LocalDataService(_duckDb); + var today = DateTime.UtcNow.Date; + var twoDaysAgo = today.AddDays(-2); + + /* 1. nothing collected at all: an empty calendar is a collection fault, not a quiet fortnight. */ + var never = Root(await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 3)); + Assert.Equal("unavailable", never.GetProperty("status").GetString()); + var neverText = never.GetProperty("message").GetString()!; + Assert.Contains("nothing has been collected", neverText, StringComparison.Ordinal); + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + + /* + 2. two collected days with a hole between them. Today collected cleanly; two days ago collected + AND recorded an ERROR run, which bands that day Critical. Yesterday is deliberately left alone: + it is the gap, and the point of the read is that a gap is an ABSENT day rather than a quiet one. + */ + await SeedRunAsync(Truncate(DateTime.UtcNow), "wait_stats", "SUCCESS", error: null); + await SeedRunAsync(twoDaysAgo.AddHours(12), "wait_stats", "SUCCESS", error: null); + await SeedRunAsync(twoDaysAgo.AddHours(13), "query_store", "ERROR", error: "seeded failure"); + + var range = Root(await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 3)); + + Assert.Equal(ServerName, range.GetProperty("server").GetString()); + Assert.Equal(3, range.GetProperty("days_back").GetInt32()); + + /* The bounds the read used, echoed back so a caller can tell which days they asked for from which + days they got — the whole distinction this payload exists to make. */ + Assert.Equal(twoDaysAgo.ToString("yyyy-MM-dd"), range.GetProperty("from_date").GetString()); + Assert.Equal(today.ToString("yyyy-MM-dd"), range.GetProperty("to_date").GetString()); + + var days = range.GetProperty("days").EnumerateArray().ToArray(); + + /* Two collected days out of three asked for; day_count counts days WITH data, and the gap between + it and days_back is the useful number. */ + Assert.Equal(2, days.Length); + Assert.Equal(2, range.GetProperty("day_count").GetInt32()); + Assert.DoesNotContain(days, d => d.GetProperty("summary_date").GetString() == today.AddDays(-1).ToString("yyyy-MM-dd")); + + /* Oldest first, as the aggregate returns them — a calendar read backwards is a bug nobody would + spot in a table. */ + Assert.Equal(twoDaysAgo.ToString("yyyy-MM-dd"), days[0].GetProperty("summary_date").GetString()); + Assert.Equal(today.ToString("yyyy-MM-dd"), days[1].GetProperty("summary_date").GetString()); + + /* The two days disagree, which is what makes this a calendar rather than one verdict. */ + Assert.Equal("Critical", days[0].GetProperty("overall_health").GetString()); + Assert.Equal(1, days[0].GetProperty("collection_errors").GetInt64()); + Assert.Equal("Healthy", days[1].GetProperty("overall_health").GetString()); + Assert.Equal(0, days[1].GetProperty("collection_errors").GetInt64()); + } + + /// + /// The anchor moves the range, and a range outside this server's history is a different answer from a + /// server that has never collected. + /// + [Fact] + public async Task TheAnchor_MovesTheRange_AndAnEmptyRangeIsNotAnUncollectedServer() + { + var service = new LocalDataService(_duckDb); + var today = DateTime.UtcNow.Date; + var tenDaysAgo = today.AddDays(-10); + + await SeedRunAsync(Truncate(DateTime.UtcNow), "wait_stats", "SUCCESS", error: null); + + /* A range this server has no history for, on a server that HAS collected. Same zero rows as the + never-collected branch, and it must not reach for the same word. */ + var beforeHistory = Root(await McpHealthTools.GetDailySummaryRange( + service, _serverManager, ServerName, 1, tenDaysAgo.ToString("yyyy-MM-dd"))); + + Assert.Equal("empty", beforeHistory.GetProperty("status").GetString()); + var beforeText = beforeHistory.GetProperty("message").GetString()!; + Assert.Contains("A day with ANY collection appears here even when every signal was quiet", beforeText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", beforeText, StringComparison.Ordinal); + + /* Now seed that day, and the SAME anchored call sees it — proof by content that the anchor reached + the query rather than being validated and thrown away. */ + await SeedRunAsync(tenDaysAgo.AddHours(12), "wait_stats", "SUCCESS", error: null); + + var anchored = Root(await McpHealthTools.GetDailySummaryRange( + service, _serverManager, ServerName, 1, tenDaysAgo.ToString("yyyy-MM-dd"))); + var anchoredDay = Assert.Single(anchored.GetProperty("days").EnumerateArray().ToArray()); + Assert.Equal(tenDaysAgo.ToString("yyyy-MM-dd"), anchoredDay.GetProperty("summary_date").GetString()); + + /* And the same one-day span unanchored answers about TODAY instead, so the anchor moved the range + rather than widening it. */ + var unanchored = Root(await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 1)); + var unanchoredDay = Assert.Single(unanchored.GetProperty("days").EnumerateArray().ToArray()); + Assert.Equal(today.ToString("yyyy-MM-dd"), unanchoredDay.GetProperty("summary_date").GetString()); + } + + [Fact] + public async Task OutOfRangeKnobs_AreRefused_NotSilentlyClamped() + { + var service = new LocalDataService(_duckDb); + + Assert.StartsWith( + "Invalid days_back value '0'", + await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 0), + StringComparison.Ordinal); + + Assert.StartsWith( + "Invalid days_back value '367'", + await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 367), + StringComparison.Ordinal); + + /* An anchor we cannot use is refused, never silently treated as today. */ + Assert.StartsWith( + "Invalid as_of", + await McpHealthTools.GetDailySummaryRange(service, _serverManager, ServerName, 30, "last tuesday"), + StringComparison.Ordinal); + } + + private static JsonElement Root(string json) => JsonDocument.Parse(json).RootElement; + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedRunAsync(DateTime collectionTimeUtc, string collector, string status, string? error) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO collection_log + (log_id, server_id, server_name, collector_name, collection_time, + duration_ms, status, error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = collector }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = 100 }); + cmd.Parameters.Add(new DuckDBParameter { Value = status }); + cmd.Parameters.Add(new DuckDBParameter { Value = (object?)error ?? DBNull.Value }); + cmd.Parameters.Add(new DuckDBParameter { Value = 10 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 80 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 20 }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/DataGridRowMarkTests.cs b/Lite.Tests/DataGridRowMarkTests.cs new file mode 100644 index 000000000..d78b8ffcd --- /dev/null +++ b/Lite.Tests/DataGridRowMarkTests.cs @@ -0,0 +1,149 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.RegularExpressions; +using PerformanceMonitor.Ui; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2645: session-scoped row marks, so an operator can record "it is done" / "to do" / "do not do" +/// against rows they are working through — asked for on the Index Analysis grid, where you decide index +/// by index and need to remember which ones you have dealt with. +/// +/// +/// The marks are held against the row OBJECT. That is a deliberate limit: UpdateData replaces the +/// row objects on every refresh, so on a live grid a mark lasts until the next one, and on a +/// run-on-demand grid — which is where this was asked for — it lasts as long as the result set. Keying on +/// row CONTENT instead would survive refreshes and needs a key function per grid, fifty-two of them, where +/// a wrong key silently moves somebody's "do not do" onto a different index. For a note whose entire value +/// is being right about which row it is on, that trade is the wrong way round. +/// +/// +public class DataGridRowMarkTests +{ + private sealed class Row + { + public string Name { get; init; } = ""; + } + + [Fact] + public void AMarkIsRememberedAgainstTheRowItWasSetOn() + { + var a = new Row { Name = "ix_a" }; + var b = new Row { Name = "ix_b" }; + + DataGridRowMarks.Set(a, DataGridRowMark.Done); + + Assert.Equal(DataGridRowMark.Done, DataGridRowMarks.Get(a)); + Assert.Equal(DataGridRowMark.None, DataGridRowMarks.Get(b)); + } + + [Fact] + public void SettingAMarkAgainReplacesIt_RatherThanStacking() + { + var row = new Row(); + + DataGridRowMarks.Set(row, DataGridRowMark.ToDo); + DataGridRowMarks.Set(row, DataGridRowMark.DoNot); + + Assert.Equal(DataGridRowMark.DoNot, DataGridRowMarks.Get(row)); + } + + [Fact] + public void NoneClearsTheMark() + { + var row = new Row(); + + DataGridRowMarks.Set(row, DataGridRowMark.Done); + DataGridRowMarks.Set(row, DataGridRowMark.None); + + Assert.Equal(DataGridRowMark.None, DataGridRowMarks.Get(row)); + } + + /// + /// Identity, not equality. Two rows that happen to carry the same values are different rows, and a + /// refresh that produces an equal-looking object must NOT inherit the old one's mark — that is the + /// content-keyed behaviour this deliberately does not implement, and inheriting it by accident would + /// be the worst of both. + /// + [Fact] + public void MarksAreByIdentity_NotByValue() + { + var original = new Row { Name = "ix_a" }; + var afterRefresh = new Row { Name = "ix_a" }; + + DataGridRowMarks.Set(original, DataGridRowMark.Done); + + Assert.Equal(DataGridRowMark.None, DataGridRowMarks.Get(afterRefresh)); + } + + [Fact] + public void ANullRowIsNeitherMarkedNorThrows() + { + DataGridRowMarks.Set(null, DataGridRowMark.Done); + + Assert.Equal(DataGridRowMark.None, DataGridRowMarks.Get(null)); + } + + /// + /// Every grid that shows the mark items must also paint them. + /// + /// The menu is one shared resource, so adding the items put them on all twenty FinOps grids at + /// once — while the paint hangs off each grid's own LoadingRow. A grid with the menu and no + /// hook offers an action that appears to do nothing, which is worse than not offering it. I shipped + /// exactly that on the first cut: menu on twenty, hook on two. + /// + [Fact] + public void EveryGridOfferingTheMarkMenu_AlsoPaintsIt() + { + var xaml = File.ReadAllText(Path.Combine(RepoRoot(), "Lite", "Controls", "FinOpsTab.xaml")); + + Assert.Contains("Click=\"MarkRow_Click\"", xaml, StringComparison.Ordinal); + + var unwired = Regex.Matches(xaml, @"\w+)""(?.*?)>", RegexOptions.Singleline) + .Where(m => !m.Groups["body"].Value.Contains("MarkedGrid_LoadingRow", StringComparison.Ordinal)) + .Select(m => m.Groups["name"].Value) + .ToArray(); + + Assert.True(unwired.Length == 0, + "These grids carry the shared context menu (and so the mark items) but have no LoadingRow hook, " + + "so marking them would appear to do nothing: " + string.Join(", ", unwired)); + } + + /// + /// All four items share one handler and differ only by Tag, so a fifth mark is a XAML line + /// rather than a fifth handler — and a typo'd Tag falls to None, which clears rather than + /// mismarks. + /// + [Fact] + public void TheFourMarkItemsShareOneHandler() + { + var xaml = File.ReadAllText(Path.Combine(RepoRoot(), "Lite", "Controls", "FinOpsTab.xaml")); + + foreach (var tag in new[] { "Done", "ToDo", "DoNot", "None" }) + { + Assert.Contains($"Tag=\"{tag}\" Click=\"MarkRow_Click\"", xaml, StringComparison.Ordinal); + } + } + + private static string RepoRoot() + { + var directory = new DirectoryInfo(AppContext.BaseDirectory); + while (directory is not null && !File.Exists(Path.Combine(directory.FullName, "PerformanceMonitor.sln"))) + { + directory = directory.Parent; + } + + return directory?.FullName ?? throw new InvalidOperationException("PerformanceMonitor.sln not found above the test output directory."); + } +} diff --git a/Lite.Tests/DuckDbLockModelTests.cs b/Lite.Tests/DuckDbLockModelTests.cs new file mode 100644 index 000000000..ef8ac4c29 --- /dev/null +++ b/Lite.Tests/DuckDbLockModelTests.cs @@ -0,0 +1,217 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Threading; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2463: the lock MODEL, rather than any one caller's use of it. +/// +/// Lite had two conventions for locking a DuckDB write and only one of them was written down. +/// The resolution is that they are answers to two different questions — the read lock excludes +/// MAINTENANCE, the write lock additionally excludes OTHER WRITERS OF THE SAME ROWS — and the axis +/// belongs to DuckDB's optimistic concurrency control rather than to anything in +/// DuckDbInitializer. These tests pin the two facts that rule rests on, and pin that it is +/// stated where a caller will actually find it. +/// +/// Only one of the three can be red on dev, and that is honest rather than a gap: the +/// other two pin properties that are ALREADY TRUE and whose value is that they announce when they stop +/// being true. A test for a fact is a tripwire, not a regression test, and writing one that fails today +/// would mean breaking the thing first. +/// +public sealed class DuckDbLockModelTests +{ + /// + /// The premise the whole store layer rests on: DuckDB.NET's async completes SYNCHRONOUSLY, so a lock + /// entered before an await is still owned by the entering thread when it is released after one. + /// + /// This is a tripwire on the driver, not a test of our code. + /// is thread-affine — ExitReadLock throws from a thread that + /// did not enter — and Lite holds that lock across await at every store call site (66 in + /// Lite/Analysis alone per #2443), on a pass that runs under Task.Run with no + /// SynchronizationContext, where a continuation is free to resume anywhere. Nothing in our + /// code makes that safe. What makes it safe is that no await here ever yields, so no continuation is + /// ever scheduled. + /// + /// If a DuckDB.NET bump makes any of these genuinely asynchronous this goes red, and the + /// paragraph on DuckDbInitializer.LockReleaser says what to do about it — including why + /// guarding Dispose with IsReadLockHeld is the wrong answer, which the test below + /// measures. + /// + /// The insert loop is long on purpose. A single moved continuation could land back on the same + /// pool thread by chance; two hundred of them landing there is not chance. + /// + [Fact] + public async Task DuckDbAsyncStillCompletesOnTheCallingThread() + { + var dir = Path.Combine(Path.GetTempPath(), $"pm-2463-{Guid.NewGuid():N}"); + Directory.CreateDirectory(dir); + try + { + /* Task.Run, so there is no SynchronizationContext to pin the continuations for us -- + the same shape the analysis pass runs in. */ + await Task.Run(async () => + { + var entered = Environment.CurrentManagedThreadId; + + using var connection = new DuckDBConnection($"Data Source={Path.Combine(dir, "lockmodel.duckdb")}"); + await connection.OpenAsync(); + Assert.Equal(entered, Environment.CurrentManagedThreadId); + + using (var ddl = connection.CreateCommand()) + { + ddl.CommandText = "CREATE TABLE t (v INTEGER)"; + await ddl.ExecuteNonQueryAsync(); + } + Assert.Equal(entered, Environment.CurrentManagedThreadId); + + for (var i = 0; i < 200; i++) + { + using var insert = connection.CreateCommand(); + insert.CommandText = $"INSERT INTO t VALUES ({i})"; + await insert.ExecuteNonQueryAsync(); + } + Assert.Equal(entered, Environment.CurrentManagedThreadId); + + using (var read = connection.CreateCommand()) + { + read.CommandText = "SELECT v FROM t"; + using var reader = await read.ExecuteReaderAsync(); + while (await reader.ReadAsync()) + { + } + } + + Assert.Equal(entered, Environment.CurrentManagedThreadId); + }); + } + finally + { + Directory.Delete(dir, recursive: true); + } + } + + /// + /// Why LockReleaser.Dispose is NOT guarded with IsReadLockHeld, kept as an executable + /// measurement because "the obvious fix is worse" is exactly the claim a later reader will doubt. + /// + /// The guard looks free: skip the exit when this thread did not enter, and the + /// goes away. What actually goes away is the diagnosis. + /// The entry the ORIGINAL thread took is still held, nothing will ever release it, and every + /// maintenance operation in the process — CHECKPOINT, archival, compaction — blocks on it for the + /// life of the app. An exception is loud, attributable and survivable; a leaked reader is none of + /// those. + /// + [Fact] + public void GuardingTheReleaserWouldTradeAnExceptionForAPermanentlyWedgedLock() + { + using var rwLock = new ReaderWriterLockSlim(LockRecursionPolicy.NoRecursion); + rwLock.EnterReadLock(); + try + { + /* What the guard would see on a continuation that resumed elsewhere. */ + var guardWouldSkip = RunOnAnotherThread(() => !rwLock.IsReadLockHeld); + Assert.True(guardWouldSkip, "a foreign thread does not observe the entry, so the guard skips the exit"); + + /* What it is being offered instead of: loud and attributable. */ + var thrown = RunOnAnotherThread(() => + { + try + { + rwLock.ExitReadLock(); + return null; + } + catch (Exception ex) + { + return ex.GetType(); + } + }); + Assert.Equal(typeof(SynchronizationLockException), thrown); + + /* And what skipping costs: the entry is still held and maintenance cannot get in. */ + Assert.Equal(1, rwLock.CurrentReadCount); + var maintenanceGotIn = RunOnAnotherThread(() => + { + if (!rwLock.TryEnterWriteLock(TimeSpan.FromMilliseconds(250))) + return false; + rwLock.ExitWriteLock(); + return true; + }); + Assert.False(maintenanceGotIn, "the leaked read entry blocks every maintenance writer, permanently"); + } + finally + { + rwLock.ExitReadLock(); + } + } + + /// + /// The rule is stated where a caller will find it, and the doc comment that used to be read as the + /// house rule no longer is. This is the one here that goes red on dev. + /// + /// A source pin because what #2463 produced is a WRITTEN RULE — no call site changed, so there + /// is no behavior to assert. The thing that can regress is the sentence, and it regresses by being + /// deleted or by drifting back to "use the write lock for INSERT", which is the phrasing that made + /// eleven callers look like a house rule and FindingStore look like a bug. + /// + [Fact] + public void TheLockRuleIsWrittenWhereEveryCallerAlreadyLooks() + { + var initializer = ParitySource.ReadFile("Lite/Database/DuckDbInitializer.cs"); + + /* The rule itself, on the lock it governs. */ + Assert.Contains("WHAT IT MUST EXCLUDE, NOT BY WHETHER IT READS OR", initializer, StringComparison.Ordinal); + Assert.Contains("#2463", initializer, StringComparison.Ordinal); + + /* The measurement that makes it a rule rather than an opinion: DuckDB fails, rather than + queues, the loser of a write-write collision -- and an append cannot collide. */ + Assert.Contains("Conflict on update!", initializer, StringComparison.Ordinal); + Assert.Contains("Conflict on tuple deletion!", initializer, StringComparison.Ordinal); + + /* And the latent thread-affinity hazard, with the reason its obvious mitigation is refused. */ + Assert.Contains("IsReadLockHeld", initializer, StringComparison.Ordinal); + Assert.Contains("bug amplifier", initializer, StringComparison.Ordinal); + + /* LocalDataService must no longer read as the house rule for the whole app. */ + var localData = ParitySource.ReadFile("Lite/Services/LocalDataService.cs"); + Assert.DoesNotContain( + "Use for UPDATE/DELETE/INSERT operations that must not race with archival or compaction.", + localData, + StringComparison.Ordinal); + Assert.Contains("#2463", localData, StringComparison.Ordinal); + + /* Every store on either side of the split points at the one rule, so whichever a reader lands + on first is where they find it. */ + foreach (var file in new[] + { + "Lite/Analysis/FindingStore.cs", + "Lite/Services/DuckDbAlertHistoryStore.cs", + "Lite/Services/DuckDbMuteRuleStore.cs", + }) + { + var source = ParitySource.ReadFile(file); + Assert.Contains("#2463", source, StringComparison.Ordinal); + Assert.Contains("DuckDbInitializer.s_dbLock", source, StringComparison.Ordinal); + } + } + + private static T RunOnAnotherThread(Func body) + { + var result = default(T)!; + var thread = new Thread(() => result = body()); + thread.Start(); + Assert.True(thread.Join(TimeSpan.FromSeconds(10)), "the probe thread did not finish"); + return result; + } +} diff --git a/Lite.Tests/EngineCapabilityMissTests.cs b/Lite.Tests/EngineCapabilityMissTests.cs new file mode 100644 index 000000000..82d3b2166 --- /dev/null +++ b/Lite.Tests/EngineCapabilityMissTests.cs @@ -0,0 +1,182 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitor.Collectors; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2511 on Lite: the SAME read, on two servers that differ only in servers.sql_engine_edition, must +/// answer differently — not_collected on Azure SQL Database, where the collector serving the read +/// cannot run at all, and the read's own empty/unavailable on an engine that does collect it. +/// +/// Both directions, always. A pin that only asserted the Azure branch would pass equally well +/// if the read had stopped distinguishing anything and started answering not_collected to everyone, +/// which is a worse defect than the one being fixed: it would hide real collection outages behind a +/// confident "this engine cannot do that". +/// +/// Lite derives its server id from the storage name rather than storing one, so the seeded registry +/// row has to be written under the same derived value the tool resolves to — a hardcoded id would seed a row +/// the tool looks straight past, and the Azure assertion would pass for the wrong reason. +/// +public sealed class EngineCapabilityMissTests : IClassFixture, IDisposable +{ + private const string AzureServerName = "LiteEngineCapAzure"; + private const string BoxServerName = "LiteEngineCapBox"; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private readonly int _azureServerId; + private readonly int _boxServerId; + private DuckDBConnection? _seedConn; + + public EngineCapabilityMissTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-enginecap-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + _azureServerId = Register(AzureServerName); + _boxServerId = Register(BoxServerName); + } + + private int Register(string name) + { + var server = new ServerConnection { Id = Guid.NewGuid().ToString(), ServerName = name, IsEnabled = true }; + _serverManager.AddServer(server); + return RemoteCollectorService.GetDeterministicHashCode(RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task AnEmptyRead_AnswersNotCollectedOnAzureSqlDb_AndKeepsItsOwnMissOnABox() + { + await SeedServerRowAsync(_azureServerId, AzureServerName, CollectorEngineCapability.AzureSqlDatabaseEngineEdition); + await SeedServerRowAsync(_boxServerId, BoxServerName, engineEdition: 3); + + var service = new LocalDataService(_duckDb); + + /* Neither server has a single collected row: the ONLY difference is the engine edition. */ + + var azureHealth = await McpHealthParserTools.GetSystemHealth(service, _serverManager, AzureServerName); + Assert.Equal("not_collected", StatusOf(azureHealth)); + Assert.Contains("Azure SQL Database", azureHealth, StringComparison.Ordinal); + Assert.Contains("EngineEdition 5", azureHealth, StringComparison.Ordinal); + Assert.Contains("system_health_events", azureHealth, StringComparison.Ordinal); + Assert.Contains(AzureServerName, azureHealth, StringComparison.Ordinal); + + /* The never-captured branch of get_health_parser_significant_waits carried the message the issue + quoted. It must no longer tell an Azure caller to start a session that cannot exist there. */ + var azureWaits = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, AzureServerName); + Assert.Equal("not_collected", StatusOf(azureWaits)); + Assert.DoesNotContain("system_health session is started", azureWaits, StringComparison.Ordinal); + + /* A read from a different family, sharing nothing with the above but the helper. */ + var azureFlags = await McpConfigTools.GetTraceFlags(service, _serverManager, AzureServerName); + Assert.Equal("not_collected", StatusOf(azureFlags)); + Assert.Contains("trace_flags", azureFlags, StringComparison.Ordinal); + + /* A third family, and deliberately NOT get_tempdb_trend: #2512 measured the tempdb DMVs returning + real data on Azure SQL Database, so #2516 opens that gate and tempdb_stats stops being a permanent + gap. Picking it as the example here would tie this test to a gate that is moving; the default + trace is absent from the engine itself, so its gate is a durable one to demonstrate with. */ + var azureTrace = await McpDefaultTraceTools.GetDefaultTraceEvents(service, _serverManager, AzureServerName); + Assert.Equal("not_collected", StatusOf(azureTrace)); + Assert.Contains("default_trace_events", azureTrace, StringComparison.Ordinal); + + /* ── The box, same empty store: every one of them keeps the answer it gave before. ── */ + Assert.Equal("empty", StatusOf(await McpHealthParserTools.GetSystemHealth(service, _serverManager, BoxServerName))); + + var boxWaits = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, BoxServerName); + Assert.Equal("unavailable", StatusOf(boxWaits)); + Assert.Contains("system_health session is started", boxWaits, StringComparison.Ordinal); + + Assert.Equal("empty", StatusOf(await McpConfigTools.GetTraceFlags(service, _serverManager, BoxServerName))); + Assert.Equal("empty", StatusOf(await McpDefaultTraceTools.GetDefaultTraceEvents(service, _serverManager, BoxServerName))); + + /* A read whose collector runs on every engine is untouched on BOTH servers — the helper must not + have become a blanket "Azure gets not_collected" rule. */ + Assert.Equal("unavailable", StatusOf(await McpConfigTools.GetDatabaseConfig(service, _serverManager, AzureServerName))); + Assert.Equal("unavailable", StatusOf(await McpConfigTools.GetDatabaseConfig(service, _serverManager, BoxServerName))); + } + + /// + /// A registry row with no probed edition — a server that has never completed a connect — keeps its old + /// miss. "We do not know" rendering as "this will never work" would be the same defect wearing the fix's + /// clothes, and it is the state every server passes through on its first cycle. + /// + [Fact] + public async Task AServerWithNoProbedEdition_KeepsItsOldMiss() + { + await SeedServerRowAsync(_boxServerId, BoxServerName, engineEdition: CollectorEngineCapability.UnknownEngineEdition); + + var service = new LocalDataService(_duckDb); + + Assert.Equal("empty", StatusOf(await McpHealthParserTools.GetSystemHealth(service, _serverManager, BoxServerName))); + Assert.Equal("empty", StatusOf(await McpDefaultTraceTools.GetDefaultTraceEvents(service, _serverManager, BoxServerName))); + Assert.Equal("empty", StatusOf(await McpConfigTools.GetTraceFlags(service, _serverManager, BoxServerName))); + } + + /// + /// A server the registry has no row for at all reads as unknown, not as a capability gap. The MCP + /// surface resolves against the ServerManager, so a freshly added server can be asked about before the + /// collector has ever written its servers row. + /// + [Fact] + public async Task AServerWithNoRegistryRow_KeepsItsOldMiss() + { + var service = new LocalDataService(_duckDb); + + Assert.Equal(CollectorEngineCapability.UnknownEngineEdition, await service.GetSqlEngineEditionAsync(_azureServerId)); + Assert.Equal("empty", StatusOf(await McpHealthParserTools.GetSystemHealth(service, _serverManager, AzureServerName))); + } + + private static string StatusOf(string json) => + JsonDocument.Parse(json).RootElement.GetProperty("status").GetString()!; + + /// Column list copied from TestDataSeeder's servers seed, plus the one column these + /// tests exist to vary. + private async Task SeedServerRowAsync(int serverId, string serverName, int engineEdition) + { + using var readLock = _duckDb.AcquireReadLock(); + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + + using var cmd = _seedConn.CreateCommand(); + cmd.CommandText = @" +INSERT INTO servers (server_id, server_name, display_name, use_windows_auth, is_enabled, sql_engine_edition) +VALUES ($1, $2, $3, true, true, $4)"; + cmd.Parameters.Add(new DuckDBParameter { Value = serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = serverName }); + cmd.Parameters.Add(new DuckDBParameter { Value = serverName }); + cmd.Parameters.Add(new DuckDBParameter { Value = engineEdition }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/FindingStoreTests.cs b/Lite.Tests/FindingStoreTests.cs index 71a22a7eb..cf20cb836 100644 --- a/Lite.Tests/FindingStoreTests.cs +++ b/Lite.Tests/FindingStoreTests.cs @@ -565,6 +565,233 @@ INSERT INTO analysis_muted (mute_id, server_id, story_path_hash, story_path, mut await cmd.ExecuteNonQueryAsync(); } + /// + /// #2448: a finding batch that faults partway through must persist NOTHING, not the rows that + /// happened to land first. + /// + /// The damage this prevents is invisible by construction, which is why it is worth a + /// behavioural test against a real DuckDB rather than a source pin. Every row in a batch shares + /// one analysis_time and reads the + /// newest analysis_time, so two committed rows of an intended five do not read as a + /// truncated set — they read as a complete analysis that found two problems. The server looks + /// HEALTHIER for the store having failed, and nothing anywhere says otherwise. + /// + /// The fault is a duplicate finding_id, which is the reachable per-row fault on + /// this table (finding_id BIGINT PRIMARY KEY) and not a contrivance: the ids come from + /// _nextId++ seeded off DateTime.UtcNow.Ticks, so two stores in one process can + /// genuinely produce one — see #2455. + /// + /// Reverted, rows 1 and 2 commit before row 3 throws and this fails on + /// Assert.Single with three rows present — the truncated set stated exactly. + /// + [Fact] + public async Task AFaultedBatch_PersistsNothing_RatherThanATruncatedSetThatReadsAsComplete() + { + const int serverId = -646464; + const long earlierPassId = 5_646_464L; + + var store = new FindingStore(_duckDb); + var context = new AnalysisContext + { + ServerId = serverId, + ServerName = "partial-batch", + TimeRangeStart = DateTime.UtcNow.AddHours(-4), + TimeRangeEnd = DateTime.UtcNow + }; + + /* An earlier pass's row, and the id the doomed batch collides with. */ + await store.InsertFindingsAsync( + [PartialBatchFinding(earlierPassId, serverId, "an earlier pass")], context); + + /* Row 3 of 5 collides, so rows 1-2 have already run by the time it fails. */ + var doomed = new List + { + PartialBatchFinding(earlierPassId + 1, serverId, "row 1"), + PartialBatchFinding(earlierPassId + 2, serverId, "row 2"), + PartialBatchFinding(earlierPassId, serverId, "row 3 - collides"), + PartialBatchFinding(earlierPassId + 3, serverId, "row 4"), + PartialBatchFinding(earlierPassId + 4, serverId, "row 5") + }; + + await Assert.ThrowsAsync(() => store.InsertFindingsAsync(doomed, context)); + + /* The assertion #2448 exists for: not "fewer rows", NO rows from the doomed batch. Only the + earlier pass survives, and it is still stamped with its own analysis_time — stale and + saying so, rather than fresh and understating the server. */ + var persisted = await store.GetRecentFindingsAsync(serverId); + Assert.Equal(earlierPassId, Assert.Single(persisted).FindingId); + + /* And the rollback is not the store simply refusing to write: the same five rows without + the collision commit in full, through the same code path. */ + var clean = new List + { + PartialBatchFinding(earlierPassId + 10, serverId, "row 1"), + PartialBatchFinding(earlierPassId + 11, serverId, "row 2"), + PartialBatchFinding(earlierPassId + 12, serverId, "row 3"), + PartialBatchFinding(earlierPassId + 13, serverId, "row 4"), + PartialBatchFinding(earlierPassId + 14, serverId, "row 5") + }; + + await store.InsertFindingsAsync(clean, context); + Assert.Equal(6, (await store.GetRecentFindingsAsync(serverId)).Count); + } + + private static AnalysisFinding PartialBatchFinding(long findingId, int serverId, string storyText) => + new() + { + FindingId = findingId, + AnalysisTime = DateTime.UtcNow, + ServerId = serverId, + ServerName = "partial-batch", + Severity = 1.0, + Confidence = 0.9, + Category = "waits", + StoryPath = "WRITELOG", + StoryPathHash = "an1-partial-batch", + StoryText = storyText, + RootFactKey = "WRITELOG", + FactCount = 1 + }; + + /// #2455: two FindingStore instances must never issue the same id. + /// + /// The filed defect was that _nextId++ ran under a lock that admits concurrent + /// holders, and it was real. The bigger half is that Interlocked.Increment on that field + /// would not have fixed anything, because there is no shared field to make atomic: Lite builds + /// TWO stores — AnalysisService and RecommendationsTab — each seeding its own + /// counter from DateTime.UtcNow.Ticks at construction. Two built in the same timer tick + /// start from the same value and then walk the same range independently. + /// + /// finding_id and mute_id are both PRIMARY KEY in DuckDB, so a collision is a + /// hard INSERT failure, and since #2448 made the batch atomic it costs the entire analysis rather + /// than one row. + /// + /// The two stores are constructed on adjacent lines and each issues 1,000 ids, so the old + /// per-instance seeds would have to differ by more than 1,000 ticks (100 microseconds) to avoid + /// overlapping — orders of magnitude more clock than two adjacent constructor calls consume. On + /// Windows, where this suite runs, the seeds are simply identical: the interrupt-timer granularity + /// behind DateTime.UtcNow is ~15.6 ms, which is ~156,000 ticks. + /// + [Fact] + public async Task TwoFindingStoresNeverIssueTheSameId() + { + /* Adjacent on purpose: this is the case the old seeding could not survive. */ + var analysisPass = new FindingStore(_duckDb); + var recommendationsTab = new FindingStore(_duckDb); + + var context = TestDataSeeder.CreateTestContext(); + var stories = ManyStories(1000); + + var fromPass = await analysisPass.FilterMutedFindingsAsync(stories, context); + var fromTab = await recommendationsTab.FilterMutedFindingsAsync(stories, context); + + Assert.Equal(1000, fromPass.Count); + Assert.Equal(1000, fromTab.Count); + + var ids = new HashSet(); + foreach (var finding in fromPass) + Assert.True(ids.Add(finding.FindingId), $"the analysis pass reissued {finding.FindingId}"); + foreach (var finding in fromTab) + Assert.True(ids.Add(finding.FindingId), $"the second store reissued {finding.FindingId}"); + + Assert.Equal(2000, ids.Count); + } + + /// + /// #2455: the read lock around the three WRITE paths is deliberate, and the reason has to stay + /// written down. + /// + /// An unexplained read-lock-to-write reads as a bug on every inspection, and the obvious + /// "fix" — swapping in AcquireWriteLock — is the one change that would actually cost + /// something: it serializes every finding batch against every UI read for the length of the batch, + /// and #2443 had just made the read-lock WAIT abandonable precisely so the analysis pass could + /// yield to a long archival rather than become the thing archival waits on. + /// + /// The lock coordinates everyone against MAINTENANCE (CHECKPOINT, archive DELETEs, + /// compaction), which takes the exclusive write lock; a held read lock blocks + /// EnterWriteLock, so holding one is how a write says "not while I am in flight". That is + /// the only exclusion these paths need — concurrency between writers is DuckDB's own job. So this + /// pins the choice AND its explanation together: changing the lock should be a decision someone + /// makes against the stated reason, not a tidy-up. + /// + [Fact] + public void TheWritePathsTakeAReadLockOnPurpose_AndSayWhy() + { + var source = File.ReadAllText(FindingStoreSourcePath()).Replace("\r\n", "\n", StringComparison.Ordinal); + + /* The choice. The write lock is not banned outright — it is banned from these three without a + reason, which is what a failing test forces someone to supply. */ + Assert.DoesNotContain("AcquireWriteLock", source, StringComparison.Ordinal); + Assert.Equal(6, CountOf(source, "_duckDb.AcquireReadLock(")); + + /* The explanation, at the class and at each write site — whichever one a reader lands on. */ + Assert.Contains("The read lock around the WRITES is deliberate (#2455)", source, StringComparison.Ordinal); + Assert.Contains("A READ lock", Between(source, "public async Task> InsertFindingsAsync(", "public async Task> SaveFindingsAsync("), StringComparison.Ordinal); + Assert.Contains("see the class note (#2455)", Between(source, "public async Task MuteStoryAsync(", "await cmd.ExecuteNonQueryAsync();"), StringComparison.Ordinal); + Assert.Contains("see the class note (#2455)", Between(source, "public async Task CleanupOldFindingsAsync(", "await cmd.ExecuteNonQueryAsync();"), StringComparison.Ordinal); + + /* And the reason the shared lock is SUFFICIENT: no read-modify-write is left under it. */ + Assert.DoesNotContain("private long _nextId", source, StringComparison.Ordinal); + Assert.Contains("FindingId = NextId(),", source, StringComparison.Ordinal); + Assert.Contains("Value = NextId() }", source, StringComparison.Ordinal); + Assert.Contains("CollectionIdGenerator.Next()", source, StringComparison.Ordinal); + } + + private static string FindingStoreSourcePath() + { + var dir = new DirectoryInfo(AppContext.BaseDirectory); + while (dir is not null && !File.Exists(Path.Combine(dir.FullName, "Lite", "Analysis", "FindingStore.cs"))) + dir = dir.Parent; + + Assert.NotNull(dir); + return Path.Combine(dir!.FullName, "Lite", "Analysis", "FindingStore.cs"); + } + + private static string Between(string source, string start, string end) + { + var from = source.IndexOf(start, StringComparison.Ordinal); + Assert.True(from >= 0, $"anchor not found: {start}"); + var to = source.IndexOf(end, from, StringComparison.Ordinal); + Assert.True(to > from, $"anchor not found after {start}: {end}"); + return source[from..to]; + } + + private static int CountOf(string source, string needle) + { + var count = 0; + var at = 0; + while ((at = source.IndexOf(needle, at, StringComparison.Ordinal)) >= 0) + { + count++; + at += needle.Length; + } + + return count; + } + + private static System.Collections.Generic.List ManyStories(int count) + { + var stories = new System.Collections.Generic.List(count); + + for (var i = 0; i < count; i++) + { + stories.Add(new AnalysisStory + { + RootFactKey = "WRITELOG", + RootFactValue = i, + Severity = 1.0, + Confidence = 1.0, + Category = "waits", + StoryPath = $"WRITELOG_{i}", + StoryPathHash = $"id-collision-{i}", + StoryText = $"story {i}", + FactCount = 1 + }); + } + + return stories; + } + private static System.Collections.Generic.List CreateTestStories() { return diff --git a/Lite.Tests/GoldenCollectorSchema.cs b/Lite.Tests/GoldenCollectorSchema.cs index 30c43ead0..4b472daeb 100644 --- a/Lite.Tests/GoldenCollectorSchema.cs +++ b/Lite.Tests/GoldenCollectorSchema.cs @@ -14,7 +14,8 @@ // ever land at the END of its table; reflecting it here keeps the oracle describing the schema a fresh // store is actually built with, while every pre-existing column stays frozen. Precedent: ag_replica_role // (v36), replica_role (v47), runtime_stats_interval_id + interval_start_time_utc (v49, #1841 tier 2). A -// NOT NULL relaxation is NOT an append — that goes in IntentionalStorageDivergences instead. +// NOT NULL relaxation is NOT an append — that goes in IntentionalStorageDivergences instead. Also +// max_size_mb on tempdb_stats (v56, #2515). // using System.Collections.Generic; @@ -132,7 +133,8 @@ total_reserved_mb DECIMAL(18,2), unallocated_mb DECIMAL(18,2), total_sessions_using_tempdb BIGINT, top_session_id INTEGER, - top_session_tempdb_mb DECIMAL(18,2) + top_session_tempdb_mb DECIMAL(18,2), + max_size_mb DECIMAL(18,2) )", ["memory_grant_stats"] = @"CREATE TABLE IF NOT EXISTS memory_grant_stats ( collection_id BIGINT PRIMARY KEY, diff --git a/Lite.Tests/JobHistoryCollectorDefinitionTests.cs b/Lite.Tests/JobHistoryCollectorDefinitionTests.cs index 6e58a26b2..0927987c2 100644 --- a/Lite.Tests/JobHistoryCollectorDefinitionTests.cs +++ b/Lite.Tests/JobHistoryCollectorDefinitionTests.cs @@ -122,9 +122,12 @@ public void AppliesTo_CollectsEverywhereExceptAzureSqlDbAndNoMsdb() { /* No SQL Agent / sysjobhistory on Azure SQL DB; Managed Instance and on-prem / RDS have it. */ Assert.False(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAzureSqlDb = true })); - /* No msdb access → every table this reads lives in msdb; skip (gate collapsed from - IsCollectorSupported into the shared AppliesTo). */ - Assert.False(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); + /* NOT gated on msdb access (#2559). It is a GRANT rather than an engine capability, and it was + probed once and cached for the connection's life - so running the GRANT we advise did nothing + until a restart. This now attempts and fails into PERMISSIONS, which error 916 already maps to, + and CollectorHealthClassifier bands a never-permitted collector as NO_PERMISSIONS ahead of + FAILING, so it does not read as broken. */ + Assert.True(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); /* NOT gated on AWS RDS — unlike running_jobs it never touches syssessions, so history reads fine. */ Assert.True(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAwsRds = true })); Assert.True(JobHistoryCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAzureManagedInstance = true })); diff --git a/Lite.Tests/LaneAxisAlignerWiringTests.cs b/Lite.Tests/LaneAxisAlignerWiringTests.cs new file mode 100644 index 000000000..f6d8774eb --- /dev/null +++ b/Lite.Tests/LaneAxisAlignerWiringTests.cs @@ -0,0 +1,100 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Runtime.CompilerServices; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2535: Lite's Overview control is a third copy of the same five-separate-WpfPlot surface that +/// carried #2533 on the Darling viewer - identical X-axis LIMITS across the lanes, but nothing gave them +/// identical pixel geometry, so a lane whose Y ticks read six digits started its data area further right +/// than one reading a fraction. Values stayed correct throughout on Darling (tooltips read true, only the +/// data->pixel mapping skewed), and the same is true here: this is a layout defect, not a data one. +/// +/// The fix already exists in the shared PerformanceMonitor.Ui.LaneAxisAligner that Lite +/// already references - see Darling.Tests/LaneAxisAlignerTests.cs for the behavioral coverage +/// (headless ScottPlot renders proving the gutter is independent of plot width and non-decreasing in plot +/// height). Those are properties of the shared helper, not of either SKU, so they are not duplicated here. +/// What IS SKU-specific is whether Lite's control actually calls it - the defect on Darling was an absent +/// call site, not a wrong calculation, so only a wiring pin against Lite's own control proves anything. +/// +/// This lives in Lite.Tests rather than beside the behavioral tests in Darling.Tests: +/// the CI path filters in .github/workflows/build.yml gate on directories, and a Lite-only edit to +/// this control fires the lite filter, which this suite is already triggered by. Parsed from +/// Darling.Tests instead, this would be a guard that silently stops guarding on exactly the change +/// it exists to catch, unless the darling filter also named this file - the same trap the parity +/// pins already documented there (Lite/Controls/ServerTab.xaml et al.) exist to avoid. +/// +public sealed class LaneAxisAlignerWiringTests +{ + /// + /// The helper working proves nothing if the control never calls it. Parse Lite's real control and + /// assert the aligner runs inside SyncXAxes and BEFORE the loop that refreshes the lanes - a + /// floor applied after the render would not take effect until something else redrew. + /// + [Fact] + public void TheLiteOverviewControl_AlignsItsLanes_BeforeItRefreshesThem() + { + var source = File.ReadAllText(ControlSourcePath()); + const string What = "Lite's Overview control"; + + string body = SyncXAxesBody(source, What); + + int align = body.IndexOf("LaneAxisAligner.AlignLeftGutters(", StringComparison.Ordinal); + Assert.True(align >= 0, + $"{What} sets identical X LIMITS but never gives the lanes one left gutter (#2535) - they will skew again the moment one lane's Y labels get wide"); + + int refresh = body.IndexOf(".Refresh();", StringComparison.Ordinal); + Assert.True(refresh > align, + $"{What} aligns the lanes after refreshing them - the shared gutter would not appear until something else redrew"); + } + + /// + /// The body of SyncXAxes, bounded by brace matching rather than by the next member's name. + /// Lite's control has an AddGhostLine helper a few members after SyncXAxes that also + /// calls Refresh(), so anchoring on "the next Refresh() call in the file" instead of on the + /// method's own closing brace would let the wrong method's code satisfy the assertions above. + /// + private static string SyncXAxesBody(string source, string what) + { + int start = source.IndexOf("private void SyncXAxes(", StringComparison.Ordinal); + Assert.True(start >= 0, $"{what} no longer has a SyncXAxes - re-anchor this pin rather than deleting it"); + + int open = source.IndexOf('{', start); + Assert.True(open > start, $"{what}: SyncXAxes has no body"); + + int depth = 0; + int end = -1; + for (int i = open; i < source.Length && end < 0; i++) + { + if (source[i] == '{') + { + depth++; + } + else if (source[i] == '}' && --depth == 0) + { + end = i; + } + } + + Assert.True(end > open, $"{what}: SyncXAxes body is unbalanced"); + return source[open..(end + 1)]; + } + + /// Lite's control, resolved from this test file's compile-time path (Lite.Tests is a sibling + /// of the Lite project). + private static string ControlSourcePath([CallerFilePath] string thisFile = "") + { + var testDir = Path.GetDirectoryName(thisFile)!; + return Path.GetFullPath(Path.Combine(testDir, "..", "Lite", "Controls", "CorrelatedTimelineLanesControl.xaml.cs")); + } +} diff --git a/Lite.Tests/Lite.Tests.csproj b/Lite.Tests/Lite.Tests.csproj index 458d72963..b3fbaa555 100644 --- a/Lite.Tests/Lite.Tests.csproj +++ b/Lite.Tests/Lite.Tests.csproj @@ -37,5 +37,13 @@ + + + + + diff --git a/Lite.Tests/LiteAlertForwardingTests.cs b/Lite.Tests/LiteAlertForwardingTests.cs index 9809a7300..85969814a 100644 --- a/Lite.Tests/LiteAlertForwardingTests.cs +++ b/Lite.Tests/LiteAlertForwardingTests.cs @@ -710,6 +710,35 @@ public async Task TempDb_ToastBody_MatchesTheOldLoop() Assert.Equal("tempdb 91% used", fired.ShortMessage); /* :435 toast body */ } + /// + /// #2515, the SKU-parity half: Lite and Darling share the engine and the type, so the ceiling has to + /// change the decision identically on both. Darling's twin is + /// AlertEngineTests.TempDb_TheAzureCeiling_SuppressesTheAlert_ThatTheAllocationWouldFire, with the + /// same numbers — an operator must not get a page on one product and silence on the other for the same + /// tempdb. + /// + [Fact] + public async Task TempDb_TheAzureCeiling_SuppressesTheAlert_ThatTheAllocationWouldFire() + { + DisableAllChecks(); + App.AlertTempDbSpaceEnabled = true; + Assert.Equal(80, App.AlertTempDbSpaceThresholdPercent); + + var h = new Harness(); + /* GP_S_Gen5_2 with one ~57 MB #temp table: 62.44 MB allocated, 65,536 MB of headroom behind it. */ + h.Adapter.TempDb = new TempDbSpaceInfo { TotalReservedMb = 59.75, UnallocatedMb = 2.69, MaxSizeMb = 65_536 }; + await h.Build().EvaluateServerAsync(Harness.Snapshot()); + Assert.Empty(h.Deliverer.Outcomes); + + /* The identical snapshot with the ceiling unmeasured is the pre-#2515 reading, and it pages. */ + var withoutCeiling = new Harness(); + withoutCeiling.Adapter.TempDb = new TempDbSpaceInfo { TotalReservedMb = 59.75, UnallocatedMb = 2.69 }; + await withoutCeiling.Build().EvaluateServerAsync(Harness.Snapshot()); + var fired = Assert.Single(withoutCeiling.Deliverer.Outcomes); + Assert.Equal("96% used (60 MB)", fired.CurrentValue); + Assert.Equal("tempdb 96% used", fired.ShortMessage); + } + [Fact] public async Task LongRunningJob_ToastBody_CarriesJobNamePercentAndMinutes() { diff --git a/Lite.Tests/LiteOverviewCardExplainsItselfTests.cs b/Lite.Tests/LiteOverviewCardExplainsItselfTests.cs new file mode 100644 index 000000000..512514ae1 --- /dev/null +++ b/Lite.Tests/LiteOverviewCardExplainsItselfTests.cs @@ -0,0 +1,386 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Windows.Media; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2437 / #2422: Lite's Overview card stops keeping the answer to itself. +/// +/// What was reported. Against the Darling viewer, but the defect is verbatim here: a card said a +/// word in a colour and would not say what about. @ehaar's question was "what is it that this text warns me +/// about?", and the reader's only recourse was to scan the metric rows guessing which one the card meant — +/// once per card, on every card. #2429 answered it on the viewer; Lite/MainWindow.xaml:442 was the same +/// bare Text="{Binding StatusDisplay}" with no ToolTip. +/// +/// What is pinned here. The property that makes the viewer's version work, and the one thing a +/// port could quietly lose: the sentence is assembled from the card's OWN metric displays, gated on the SAME +/// predicates the row brushes are painted from, so it can never name a metric whose row is green or stay +/// silent about one that is not. asserts that +/// against the shipped brushes rather than against a copy of the thresholds — a re-derivation of severity +/// would pass every text assertion below and still be worse than no tooltip. +/// +/// What Lite does NOT have. There is no AwaitingFirstCollection flag on Lite's card, so the +/// status/tooltip desync family #2429 spent four review rounds on cannot arise. What Lite has instead is the +/// inverse conflation, pinned by : its status word is +/// a CONNECTION word, so a card in real metric trouble reads a green "Online" while its border is red, and its +/// amber "Warning" is about failing collectors and says nothing about the metrics at all. +/// +public sealed class LiteOverviewCardExplainsItselfTests +{ + /* Every fixture leaves OtherProcessCpuPercent null on purpose. CpuPercentForAlert reads the process-wide + App.AlertCpuMode, which another suite legitimately flips; with no other-process figure both modes read + the same number, so nothing here depends on which one happens to be set. */ + + private static ServerSummaryItem Healthy() => + new() { DisplayName = "calm", ServerId = 1, IsOnline = true, CpuPercent = 5 }; + + private static ServerSummaryItem Busy() => + new() { DisplayName = "busy", ServerId = 2, IsOnline = true, CpuPercent = 96, DeadlockCount = 2 }; + + private static ServerSummaryItem CollectorsFailing() => + new() { DisplayName = "erroring", ServerId = 3, IsOnline = true, CpuPercent = 5, HasCollectorErrors = true }; + + private static ServerSummaryItem Offline() => + new() { DisplayName = "dark", ServerId = 4, IsOnline = false, CpuPercent = 96, DeadlockCount = 2 }; + + private static ServerSummaryItem NotYetChecked() => + new() { DisplayName = "queued", ServerId = 5, IsOnline = null, CpuPercent = 62 }; + + private static IEnumerable EveryState() => + new[] { Healthy(), Busy(), CollectorsFailing(), Offline(), NotYetChecked() }; + + // ── the card explains itself ─────────────────────────────────────────────────────────────────── + + /// + /// The invariant the whole change rests on, asserted against the SHIPPED brushes rather than a copy of the + /// thresholds: the tooltip names a metric exactly when that metric's row is not green. A Lite tooltip that + /// recomputed severity independently would satisfy every wording assertion below and still drift from the + /// rows sitting directly underneath it, which is the one failure mode a tooltip must not have. + /// + [Fact] + public void TheTooltip_NamesExactlyTheRowsTheCardHasColoured() + { + var green = Healthy().CpuBrush.Color; + + foreach (var card in EveryState()) + { + /* Offline is the documented exception and is checked on its own below: the card draws a dimming + overlay over those rows because the numbers under it are pre-blackout. */ + if (card.IsOffline) + { + continue; + } + + var tooltip = card.StatusTooltip; + + Assert.Equal(card.CpuBrush.Color != green, tooltip.Contains("CPU ", StringComparison.Ordinal)); + Assert.Equal(card.BlockingBrush.Color != green, tooltip.Contains("Blocking ", StringComparison.Ordinal)); + Assert.Equal(card.DeadlockBrush.Color != green, tooltip.Contains("Deadlocks ", StringComparison.Ordinal)); + } + } + + /// The numbers in the sentence are the card's own display strings, character for character — + /// not a second formatting of the same values, which is how a tooltip ends up rounding differently from + /// the row it is explaining. + [Fact] + public void TheTooltip_QuotesTheCardsOwnDisplays() + { + var card = Busy(); + + Assert.Equal("96%", card.CpuDisplay); + Assert.Equal("2", card.DeadlockDisplay); + Assert.Equal("CPU 96%, Deadlocks 2", card.StatusReason); + Assert.Contains("Needs attention: CPU 96%, Deadlocks 2", card.StatusTooltip, StringComparison.Ordinal); + } + + /// + /// Lite's conflation, which is not the viewer's. There the amber "Warning" means the collection has gone + /// stale; here it means collectors are erroring, and the metric rows have no say in the status word at all — + /// so a card in real trouble reads a green "Online" over a red border. Both halves have to be named or the + /// tooltip just repeats the word. + /// + [Fact] + public void TheTooltip_SaysWhichAxisTheStatusWordIsAbout() + { + var erroring = CollectorsFailing(); + + Assert.Equal("Warning", erroring.StatusDisplay); + Assert.Contains("collectors are failing", erroring.StatusTooltip, StringComparison.Ordinal); + /* ...and says so without claiming a metric problem the rows do not show. */ + Assert.Contains("Every metric on this card is inside its threshold", erroring.StatusTooltip, StringComparison.Ordinal); + + var busy = Busy(); + + Assert.Equal("Online", busy.StatusDisplay); + Assert.Contains("Needs attention: CPU 96%, Deadlocks 2", busy.StatusTooltip, StringComparison.Ordinal); + /* The border is already red for this card. The word never was, which is the whole point. */ + Assert.NotEqual(Healthy().CardBorderBrush.Color, busy.CardBorderBrush.Color); + } + + /// A calm card gets an all-clear, not a demand. The viewer had to guard this because its ranking's + /// "Needs attention" fallback would otherwise reach every healthy card in the fleet; the same sentence is + /// available to get wrong here, and the card grid shows EVERY server. + [Fact] + public void TheTooltip_OnACalmCard_DoesNotClaimItNeedsAttention() + { + var tooltip = Healthy().StatusTooltip; + + Assert.Contains("Online — the last connection check succeeded", tooltip, StringComparison.Ordinal); + Assert.Contains("Every metric on this card is inside its threshold", tooltip, StringComparison.Ordinal); + Assert.DoesNotContain("Needs attention", tooltip, StringComparison.Ordinal); + } + + /// + /// An offline card is not told to act on the numbers behind its own blackout overlay. Those are the last + /// values collected before the server went dark; the card dims them for that reason, and a tooltip + /// demanding attention for them would contradict the card while the reader is looking at it. + /// + [Fact] + public void TheTooltip_OnAnOfflineCard_DoesNotDemandActionOnPreBlackoutNumbers() + { + var card = Offline(); + var tooltip = card.StatusTooltip; + + Assert.True(card.IsOffline); + Assert.Contains("Offline — the last connection check failed", tooltip, StringComparison.Ordinal); + Assert.DoesNotContain("Needs attention", tooltip, StringComparison.Ordinal); + Assert.DoesNotContain("CPU", tooltip, StringComparison.Ordinal); + + /* The reason itself is still computed — the omission is a tooltip decision, made in one place, not a + hole in what the card knows. */ + Assert.Equal("CPU 96%, Deadlocks 2", card.StatusReason); + } + + /// "Unknown" is the one word with nothing behind it: no connection check has run. Not knowing + /// whether a server is reachable is no reason to withhold the CPU number that WAS collected. + [Fact] + public void TheTooltip_OnAnUncheckedCard_StillReportsWhatWasCollected() + { + var tooltip = NotYetChecked().StatusTooltip; + + Assert.Contains("Unknown — this server has not been connection-checked yet", tooltip, StringComparison.Ordinal); + Assert.Contains("Needs attention: CPU 62%", tooltip, StringComparison.Ordinal); + } + + /// Every card gets a tooltip, every tooltip ends on the gesture that acts on it, and no tooltip is + /// ever the bare word the reader already read. + [Fact] + public void EveryCard_GetsATooltipThatSaysSomethingTheWordDidNot() + { + foreach (var card in EveryState()) + { + var tooltip = card.StatusTooltip; + + Assert.NotEqual(card.StatusDisplay, tooltip); + Assert.Contains(card.StatusDisplay + " — ", tooltip, StringComparison.Ordinal); + Assert.EndsWith("Double-click the card to open this server's tab", tooltip, StringComparison.Ordinal); + } + } + + /// + /// The word, the colour and the tooltip's first line are three renderings of ONE discriminant, which is what + /// makes them incapable of disagreeing rather than merely observed not to. The viewer arrived at the same + /// collapse over two review rounds on #2429, each of which found a different flag pair where two independent + /// readings contradicted each other. + /// + /// Since #2458 the ladder is ServerCardStatusRules.Classify rather than a member of this class, + /// and the count below spans the sidebar's file too. The reason is the shape of what this pin missed: it + /// counted one literal in ONE file, which was true and stayed true while a fourth copy of the same ladder sat + /// in Lite/Models/ServerConnection.cs driving the sidebar dot. A pin scoped to one file cannot see the + /// duplicate it exists to forbid. + /// + [Fact] + public void TheCardStatus_IsTheOnlyPlaceTheStatusFlagsAreRead() + { + Assert.Equal(ServerCardStatus.Online, Healthy().CardStatus); + Assert.Equal(ServerCardStatus.CollectorErrors, CollectorsFailing().CardStatus); + Assert.Equal(ServerCardStatus.Offline, Offline().CardStatus); + Assert.Equal(ServerCardStatus.Unknown, NotYetChecked().CardStatus); + + /* An offline card's collector-error marker must not turn it amber — the flags are read once, in order. */ + var offlineAndErroring = Offline(); + offlineAndErroring.HasCollectorErrors = true; + Assert.Equal(ServerCardStatus.Offline, offlineAndErroring.CardStatus); + Assert.Equal("Offline", offlineAndErroring.StatusDisplay); + Assert.Contains("Offline", offlineAndErroring.StatusTooltip, StringComparison.Ordinal); + + var source = ReadRepoFile(Path.Combine("Lite", "Services", "LocalDataService.Overview.cs")); + Assert.Contains( + "public static ServerCardStatus Classify(bool? isOnline, bool hasCollectorErrors) => isOnline switch", + source, StringComparison.Ordinal); + Assert.Contains( + "public ServerCardStatus CardStatus => ServerCardStatusRules.Classify(IsOnline, HasCollectorErrors);", + source, StringComparison.Ordinal); + Assert.Contains("public string StatusDisplay => CardStatus.Word();", source, StringComparison.Ordinal); + Assert.Contains("public SolidColorBrush StatusBrush => MakeBrush(CardStatus switch", source, StringComparison.Ordinal); + Assert.Contains("private string StatusHeadline => CardStatus.Headline();", source, StringComparison.Ordinal); + + /* And exactly one switch on the flag itself, counted across BOTH files that render it. This is the + assertion behind the claim in ServerCardStatus's doc comment; without it the claim is prose that + review has to re-check by hand, which is how it came to overstate what was true in the first place + (raised on #2451). The literal moved with the ladder — it is now the classifier's own parameter + list — and the card may not switch on the property again. The sidebar's file is scanned for either + casing, because the copy #2458 removed was the property-cased one. */ + var sidebar = ReadRepoFile(Path.Combine("Lite", "Models", "ServerConnection.cs")); + Assert.Equal(1, CountOccurrences(source, "isOnline switch")); + Assert.Equal(0, CountOccurrences(source, "IsOnline switch")); + Assert.Equal(0, CountOccurrences(sidebar, "sOnline switch")); + } + + /// + /// The card's BORDER renders the same discriminant its word does. Review on #2451 found the last place it + /// did not: CardBorderBrush read IsOffline and HasCollectorErrors raw, so it agreed + /// with the status word by coincidence rather than by construction — and on one pair it already did not + /// agree. An unchecked card carrying a collector-error marker drew the amber "collectors failing" border + /// while its word read "Unknown". The loader only sets that marker when the connection check succeeded, so + /// the pair is unreached in practice, which is the argument for making it unrepresentable rather than + /// leaving a caller to keep avoiding it. + /// + [Fact] + public void TheCardBorder_RendersTheSameDiscriminantTheWordDoes() + { + var neutral = Healthy().CardBorderBrush.Color; + + var notChecked = NotYetChecked(); + notChecked.HasCollectorErrors = true; + + Assert.Equal(ServerCardStatus.Unknown, notChecked.CardStatus); + Assert.Equal("Unknown", notChecked.StatusDisplay); + Assert.DoesNotContain("collectors", notChecked.StatusTooltip, StringComparison.Ordinal); + Assert.Equal(neutral, notChecked.CardBorderBrush.Color); + + /* A card that IS erroring still gets its amber border — and it is now literally the same amber the + status word is painted, because both render the same state. */ + var erroring = CollectorsFailing(); + Assert.NotEqual(neutral, erroring.CardBorderBrush.Color); + Assert.Equal(erroring.StatusBrush.Color, erroring.CardBorderBrush.Color); + + /* Precedence is unchanged: a dark server outranks its own metrics. */ + Assert.Equal(Offline().CardBorderBrush.Color, Busy().CardBorderBrush.Color); + + var source = ReadRepoFile(Path.Combine("Lite", "Services", "LocalDataService.Overview.cs")); + Assert.Contains("CardStatus == ServerCardStatus.Offline ? \"#E57373\"", source, StringComparison.Ordinal); + Assert.Contains("CardStatus == ServerCardStatus.CollectorErrors ? \"#FFD54F\"", source, StringComparison.Ordinal); + Assert.Contains("public bool IsOffline => CardStatus == ServerCardStatus.Offline;", source, StringComparison.Ordinal); + Assert.DoesNotContain("IsOffline ? \"#E57373\"", source, StringComparison.Ordinal); + } + + /// + /// The row brushes and the reason read the SAME gates. This is the source-level half of + /// : that test proves they agree today, this + /// one forbids the second copy of a threshold that would let them stop agreeing later. Memory is absent from + /// both because the Memory row carries no severity brush — there is no band on this card to report. + /// + [Fact] + public void TheRowBrushes_AndTheReason_ReadOneGateEach() + { + var source = ReadRepoFile(Path.Combine("Lite", "Services", "LocalDataService.Overview.cs")); + + Assert.Contains("private bool CpuIsElevated => CpuPercentForAlert >= 50;", source, StringComparison.Ordinal); + Assert.Contains("private bool CpuIsCritical => CpuPercentForAlert >= 80;", source, StringComparison.Ordinal); + Assert.Contains("private bool BlockingIsElevated => BlockingCount > 0;", source, StringComparison.Ordinal); + Assert.Contains("private bool DeadlocksAreElevated => DeadlockCount > 0;", source, StringComparison.Ordinal); + + /* Both readers of each gate, named: the row's brush, and the reason. */ + Assert.Contains("MakeBrush(CpuIsCritical ? \"#E57373\" : CpuIsElevated ? \"#FFB74D\"", source, StringComparison.Ordinal); + Assert.Contains("MakeBrush(BlockingIsElevated ? \"#FFB74D\"", source, StringComparison.Ordinal); + Assert.Contains("MakeBrush(DeadlocksAreElevated ? \"#E57373\"", source, StringComparison.Ordinal); + Assert.Contains("if (CpuIsElevated)", source, StringComparison.Ordinal); + Assert.Contains("if (BlockingIsElevated)", source, StringComparison.Ordinal); + Assert.Contains("if (DeadlocksAreElevated)", source, StringComparison.Ordinal); + + /* And no second copy of any threshold. These are the literals the brushes carried before the gates + existed; a tooltip written against its own copy of them is the drift this whole class is about. */ + Assert.DoesNotContain("CpuPercentForAlert >= 80 ? \"#FFB74D\"", source, StringComparison.Ordinal); + Assert.DoesNotContain("BlockingCount > 0 ? \"#FFB74D\"", source, StringComparison.Ordinal); + Assert.DoesNotContain("DeadlockCount > 0 ? \"#E57373\"", source, StringComparison.Ordinal); + } + + // ── the wiring, which only source can show ───────────────────────────────────────────────────── + + /// + /// The card's status line actually carries the tooltip. Background="Transparent" is load-bearing + /// rather than decorative: a TextBlock with a null Background hit-tests on its rendered glyphs alone, so + /// the tooltip would appear over the letters of "Warning" and nowhere in the space around them. Removing + /// either attribute compiles perfectly clean, which is why this reads XAML. + /// + [Fact] + public void TheCardStatusLine_IsBoundToTheTooltip_AndIsHoverable() + { + var xaml = ReadRepoFile(Path.Combine("Lite", "MainWindow.xaml")); + + var at = xaml.IndexOf("Text=\"{Binding StatusDisplay}\"", StringComparison.Ordinal); + Assert.True(at > 0, "the Overview card's status TextBlock is gone — find where it moved before editing this test"); + + var element = xaml[at..(xaml.IndexOf("/>", at, StringComparison.Ordinal) + 2)]; + + Assert.Contains("ToolTip=\"{Binding StatusTooltip}\"", element, StringComparison.Ordinal); + Assert.Contains("Background=\"Transparent\"", element, StringComparison.Ordinal); + + /* The status dot carries it too — it is the same signal, and it is the thing a reader points at first. */ + var dot = xaml.IndexOf("Fill=\"{Binding StatusBrush}\"", StringComparison.Ordinal); + Assert.True(dot > 0, "the Overview card's status dot is gone"); + Assert.Contains("ToolTip=\"{Binding StatusTooltip}\"", xaml[dot..(xaml.IndexOf("/>", dot, StringComparison.Ordinal) + 2)], StringComparison.Ordinal); + } + + /// + /// The closing line is the viewer's, verbatim. The reason all three surfaces are being fixed together is + /// that a reader moving between Lite, the viewer and the web dashboard should meet one vocabulary rather + /// than three levels of helpfulness — so this is a pin, not a coincidence. + /// + [Fact] + public void TheTooltipsClosingLine_MatchesTheDarlingViewers() + { + const string action = "Double-click the card to open this server's tab"; + + Assert.EndsWith(action, Healthy().StatusTooltip, StringComparison.Ordinal); + + var viewer = ReadRepoFile(Path.Combine( + "Darling", "PerformanceMonitor.Darling.Viewer", "ViewerDataService.Fleet.cs")); + Assert.Contains("\"" + action + "\"", viewer, StringComparison.Ordinal); + } + + // ── helpers ──────────────────────────────────────────────────────────────────────────────────── + + private static int CountOccurrences(string haystack, string needle) + { + var count = 0; + var index = 0; + while ((index = haystack.IndexOf(needle, index, StringComparison.Ordinal)) >= 0) + { + count++; + index += needle.Length; + } + return count; + } + + /// Locates a repo file by walking up from this file's compile-time path — the + /// AlertFiringLogTests / ThemeCompletenessTests idiom, no build-output copying. + private static string ReadRepoFile(string relative, [CallerFilePath] string thisFile = "") + { + for (var dir = new DirectoryInfo(Path.GetDirectoryName(thisFile)!); dir is not null; dir = dir.Parent) + { + var candidate = Path.Combine(dir.FullName, relative); + if (File.Exists(candidate)) + { + return File.ReadAllText(candidate).Replace("\r\n", "\n", StringComparison.Ordinal); + } + } + + throw new FileNotFoundException($"Could not locate {relative} walking up from {thisFile}"); + } +} diff --git a/Lite.Tests/LiteRuntimePrerequisiteDocsTests.cs b/Lite.Tests/LiteRuntimePrerequisiteDocsTests.cs new file mode 100644 index 000000000..1a22ea81a --- /dev/null +++ b/Lite.Tests/LiteRuntimePrerequisiteDocsTests.cs @@ -0,0 +1,630 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.Globalization; +using System.IO; +using System.Linq; +using System.Text.Json; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2489/#2501: Lite ships TWO artifacts, and what a reader must install before either one starts is +/// decided nowhere near the sentence that states it. #2489 was the docs having the two backwards — +/// PerformanceMonitorLite-win-Setup.exe is packed from a --self-contained publish and +/// needs no runtime at all, yet it was the one artifact carrying a runtime requirement in the README, +/// while the portable ZIP was framework-dependent and had no prerequisites documented anywhere. +/// #2501 then made the ZIP self-contained too, so the correct answer for BOTH artifacts is now +/// "nothing" — and every runtime sentence #2499 added had to come back out. +/// +/// The reason this is worth a guard rather than a one-time correction is that nothing in the product +/// forces the docs and the artifacts to agree. Both facts are decided in files nobody edits while +/// writing docs — the publish shape lives in the workflows, and the framework list is a build OUTPUT +/// that changes when a package reference changes (ASP.NET Core is on that list only because +/// ModelContextProtocol.AspNetCore drags the framework reference in transitively). A prose +/// sentence cannot notice either one moving. +/// +/// So every assertion here is DERIVED from the shipped artifact, never from a list kept beside it: +/// the runtimes a framework-dependent build would demand come from the built +/// PerformanceMonitorLite.runtimeconfig.json, and which artifacts are self-contained comes from +/// parsing the dotnet publish lines in the two workflows that build them. Every claim is stated +/// BOTH ways round, because #2501 proved a one-way assertion is not enough: under #2499's version of +/// this file, flipping the ZIP to self-contained left two facts still green — the docs kept demanding +/// two runtimes and nothing noticed. Now, revert either publish line to framework-dependent and the +/// prose has to come back; leave them self-contained and the prose must not be there. +/// +/// A framework this file has no mapping for is still a hard failure, because an undocumentable +/// prerequisite is exactly the state that produced #2489. +/// +/// Note for whoever touches the CI path filters: README.md and Lite/README.md are named +/// explicitly in build.yml's lite filter. They have to be. Every area filter carves markdown +/// out (dir/**/!(*.md)), and a docs-only PR additionally engages the fast path that skips .NET +/// setup entirely — so without those entries this guard would be unrunnable on precisely the change it +/// exists to catch. +/// +public sealed class LiteRuntimePrerequisiteDocsTests +{ + private const string RootReadmePath = "README.md"; + private const string LiteReadmePath = "Lite/README.md"; + private const string ShippedNoticePath = "Lite/READ-ME-FIRST.txt"; + private const string LiteProjectPath = "Lite/PerformanceMonitorLite.csproj"; + private const string LiteLockFilePath = "Lite/packages.lock.json"; + private const string BuildWorkflowPath = ".github/workflows/build.yml"; + private const string NightlyWorkflowPath = ".github/workflows/nightly.yml"; + + private const string PublishCommand = "dotnet publish Lite/PerformanceMonitorLite.csproj"; + private const string SetupExeName = "PerformanceMonitorLite-win-Setup.exe"; + private const string PortableZipName = "PerformanceMonitorLite-.zip"; + + /// + /// The root README's Lite section heading. The runtime-download rule is scoped to this section + /// rather than applied to the whole file on purpose: Darling IS framework-dependent (#2481), and + /// the day its prerequisites get written down in the root README they will name the same two + /// downloads. A whole-file rule would then go red for a reason that has nothing to do with Lite. + /// + private const string RootReadmeLiteHeading = "## Quick Start — Lite"; + + /// + /// Shared-framework name to the name of the runtime a human downloads to satisfy it, as a format + /// string over the major version. This is a TRANSLATION table, not an inventory: the set of + /// frameworks is read out of the built runtimeconfig, and a name missing from here fails the suite + /// rather than being skipped. + /// + /// Microsoft.NETCore.App maps to nothing on purpose. Both installers a reader could be sent + /// to contain it, so naming it in the docs would send an operator to a third download they do not + /// need. + /// + private static readonly Dictionary InstallerNameFormats = new(StringComparer.Ordinal) + { + ["Microsoft.NETCore.App"] = string.Empty, + ["Microsoft.WindowsDesktop.App"] = ".NET Desktop Runtime {0}", + ["Microsoft.AspNetCore.App"] = "ASP.NET Core Runtime {0}", + }; + + /// + /// Phrases that can only be on a doc line to tell a reader they must go and install something + /// before the artifact that line names will start. Checked against lines that NAME an artifact, so + /// a sentence about the build system, or about Darling, is not caught by them. + /// + private static readonly string[] RuntimeDemandPhrases = + [ + "requires", + "Desktop Runtime", + "ASP.NET Core Runtime", + "framework-dependent", + ]; + + /// One shared framework a framework-dependent build names, and the major version it wants. + private sealed record FrameworkRequirement(string Name, int Major, string SourcePath); + + /// + /// The frameworks a FRAMEWORK-DEPENDENT Lite build asks the host for, read from its own + /// runtimeconfig.json under Lite/bin. + /// + /// Reading the build output rather than the csproj is deliberate. The csproj never mentions + /// Microsoft.AspNetCore.App; that framework arrives transitively and only the SDK's own + /// output says so, which is the whole reason the requirement was a surprise. A self-contained + /// runtimeconfig carries includedFrameworks instead of frameworks and is skipped — + /// it states what is bundled, which by definition is not a prerequisite. + /// + /// This still resolves after #2501 because the ordinary RID-less dotnet build that CI runs + /// to compile this suite is framework-dependent; only the PUBLISH is pinned to win-x64. That is + /// what keeps the "if it were framework-dependent, here is what it would demand" half of this file + /// derivable at all. + /// + private static IReadOnlyList FrameworkDependentRequirements() + { + var binRoot = Path.Combine(ParitySource.RepoRoot(), "Lite", "bin"); + Assert.True( + Directory.Exists(binRoot), + $"{binRoot} does not exist, so this guard cannot read what Lite actually asks the .NET host for. " + + "Lite.Tests references Lite, so building the suite builds it - if this fires, the build layout moved."); + + var configs = Directory.GetFiles(binRoot, "PerformanceMonitorLite.runtimeconfig.json", SearchOption.AllDirectories); + var requirements = new List(); + + foreach (var configPath in configs) + { + using var document = JsonDocument.Parse(File.ReadAllText(configPath)); + + if (!document.RootElement.TryGetProperty("runtimeOptions", out var runtimeOptions) || + !runtimeOptions.TryGetProperty("frameworks", out var frameworks)) + { + /* Self-contained output (includedFrameworks), or a shape this does not understand. */ + continue; + } + + foreach (var framework in frameworks.EnumerateArray()) + { + var name = framework.GetProperty("name").GetString(); + var version = framework.GetProperty("version").GetString(); + Assert.False(string.IsNullOrWhiteSpace(name), $"Nameless framework entry in {configPath}"); + Assert.False(string.IsNullOrWhiteSpace(version), $"Versionless framework entry in {configPath}"); + + var major = int.Parse(version!.Split('.')[0], CultureInfo.InvariantCulture); + requirements.Add(new FrameworkRequirement(name!, major, configPath)); + } + } + + Assert.True( + requirements.Count > 0, + $"No framework-dependent PerformanceMonitorLite.runtimeconfig.json found under {binRoot}. " + + $"Searched {configs.Length} runtimeconfig file(s). Without one, nothing here is derived and the " + + "docs would be guarded by a list instead of by the build."); + + return requirements; + } + + /// The runtime names a reader would be told to download if a Lite artifact were framework-dependent. + private static string[] DerivedInstallerNames() => + FrameworkDependentRequirements() + .Where(r => InstallerNameFormats[r.Name].Length > 0) + .Select(r => string.Format(CultureInfo.InvariantCulture, InstallerNameFormats[r.Name], r.Major)) + .Distinct(StringComparer.Ordinal) + .ToArray(); + + /// + /// Every dotnet publish of Lite in one workflow, as output directory to "was it published + /// --self-contained". This is what makes "neither artifact needs a runtime" a derived claim + /// rather than a remembered one. + /// + private static Dictionary LitePublishShapes(string workflowPath) + { + var shapes = new Dictionary(StringComparer.OrdinalIgnoreCase); + + foreach (var rawLine in ParitySource.ReadFile(workflowPath).Split('\n')) + { + var line = rawLine.Trim(); + if (!line.Contains(PublishCommand, StringComparison.Ordinal)) + { + continue; + } + + var tokens = line.Split(' ', StringSplitOptions.RemoveEmptyEntries); + var outputFlag = Array.IndexOf(tokens, "-o"); + Assert.True( + outputFlag >= 0 && outputFlag + 1 < tokens.Length, + $"A Lite publish in {workflowPath} has no '-o ' this guard can read: {line}"); + + shapes[tokens[outputFlag + 1]] = line.Contains("--self-contained", StringComparison.Ordinal); + } + + Assert.True(shapes.Count > 0, $"No '{PublishCommand}' line found in {workflowPath}."); + return shapes; + } + + /// Both workflows' Lite publish shapes in one dictionary, keyed by workflow path plus output dir. + private static Dictionary AllLitePublishShapes() + { + var all = new Dictionary(StringComparer.OrdinalIgnoreCase); + + foreach (var workflow in new[] { BuildWorkflowPath, NightlyWorkflowPath }) + { + foreach (var (dir, selfContained) in LitePublishShapes(workflow)) + { + all[$"{workflow} -> {dir}"] = selfContained; + } + } + + return all; + } + + /// + /// Every runtime identifier a Lite publish pins with -r, across both workflows. Empty when + /// every publish is RID-agnostic, which is the state the committed lock file was written for + /// before #2501. + /// + private static string[] PinnedRuntimeIdentifiers() => + new[] { BuildWorkflowPath, NightlyWorkflowPath } + .SelectMany(path => ParitySource.ReadFile(path).Split('\n')) + .Select(l => l.Trim()) + .Where(l => l.Contains(PublishCommand, StringComparison.Ordinal)) + .Select(l => l.Split(' ', StringSplitOptions.RemoveEmptyEntries)) + .Select(tokens => (Tokens: tokens, Flag: Array.IndexOf(tokens, "-r"))) + .Where(t => t.Flag >= 0 && t.Flag + 1 < t.Tokens.Length) + .Select(t => t.Tokens[t.Flag + 1]) + .Distinct(StringComparer.Ordinal) + .ToArray(); + + /// + /// The three places a reader can land on Lite's "what do I install first" answer, each already + /// narrowed to the text that is about LITE. Lite's own README and the notice inside the ZIP are + /// wholly Lite's; the root README is shared with Darling, so only its Lite section counts. + /// + private static IEnumerable<(string Doc, string Text)> LiteRuntimeDocScopes() + { + yield return (RootReadmePath, SectionOf(RootReadmePath, RootReadmeLiteHeading)); + yield return (LiteReadmePath, ParitySource.ReadFile(LiteReadmePath)); + yield return (ShippedNoticePath, ParitySource.ReadFile(ShippedNoticePath)); + } + + /// + /// One markdown section: from to the next same-level heading. Taken + /// from the document's own structure rather than by line number, so re-ordering the README cannot + /// silently move the window somewhere else. + /// + private static string SectionOf(string docPath, string heading) + { + var lines = ParitySource.ReadFile(docPath).Split('\n').Select(l => l.TrimEnd('\r')).ToArray(); + + var start = Array.FindIndex(lines, l => l.Trim().Equals(heading, StringComparison.Ordinal)); + Assert.True( + start >= 0, + $"{docPath} has no '{heading}' heading. This guard scopes Lite's runtime claims to that " + + "section; renaming it needs this constant updated in the same change, or the rule quietly " + + "starts checking nothing."); + + var end = Array.FindIndex(lines, start + 1, l => l.StartsWith("## ", StringComparison.Ordinal)); + + return string.Join('\n', lines[start..(end < 0 ? lines.Length : end)]); + } + + /// The last path segment of a workflow path token, quotes and a trailing /* removed. + private static string DirectoryLeaf(string token) => + Path.GetFileName(token.Trim('\'', '"').TrimEnd('*').TrimEnd('/', '\\')); + + /// Every line of a doc that mentions . + private static string[] LinesMentioning(string docPath, string needle) => + ParitySource.ReadFile(docPath) + .Split('\n') + .Select(l => l.TrimEnd('\r')) + .Where(l => l.Contains(needle, StringComparison.OrdinalIgnoreCase)) + .ToArray(); + + [Fact] + public void TheBuiltRuntimeconfig_NamesOnlyFrameworksThisGuardCanTranslateToADownload() + { + /* The gate that keeps the rest of this file honest. If a new shared framework shows up in the + runtimeconfig, somebody has to decide what an operator downloads for it and say so in the + docs - silently skipping it is how the ASP.NET Core requirement went undocumented for a + release in the first place. Still load-bearing after #2501: the day a publish goes back to + framework-dependent, the sentences this table generates are the ones that have to be written. */ + var requirements = FrameworkDependentRequirements(); + + var unmapped = requirements + .Select(r => r.Name) + .Distinct(StringComparer.Ordinal) + .Where(name => !InstallerNameFormats.ContainsKey(name)) + .ToArray(); + + Assert.True( + unmapped.Length == 0, + $"Lite's runtimeconfig now names framework(s) this guard has no download mapping for: " + + $"{string.Join(", ", unmapped)}. Add the mapping to InstallerNameFormats, and decide whether " + + $"{RootReadmePath}, {LiteReadmePath} and {ShippedNoticePath} have to name it - a prerequisite " + + "nobody documents is #2489."); + + var majors = requirements.Select(r => r.Major).Distinct().ToArray(); + Assert.True( + majors.Length == 1, + $"Lite's runtimeconfig files disagree about the .NET major version ({string.Join(", ", majors)}), " + + "so there is no single number the docs could be correct about."); + } + + [Fact] + public void EveryShippedLiteArtifact_IsPublishedSelfContained() + { + /* #2501. Both publishes are now -r win-x64 --self-contained, and that is what entitles every + other assertion here to say the answer to "what do I install first" is "nothing". It is also + the SMALLER artifact, counter-intuitively: the RID-agnostic publish carried 537 MB of + runtimes\ for platforms Windows can never load, and dropping them beat the cost of bundling + the runtime by roughly two to one - 565 MB tree / 212.7 MB zipped became 277 MB / 114.2 MB, + measured on one commit and one SDK. + + If this goes red, the change that made a publish framework-dependent again also has to put + the prerequisites back into all three docs; the other facts here will say so individually. */ + var frameworkDependent = AllLitePublishShapes() + .Where(kv => !kv.Value) + .Select(kv => kv.Key) + .ToArray(); + + Assert.True( + frameworkDependent.Length == 0, + $"These Lite publishes are framework-dependent: {string.Join(", ", frameworkDependent)}. " + + "Every artifact Lite ships is supposed to carry its own runtime (#2501), and the docs say so " + + "in three places. Either restore --self-contained, or restore the prerequisites sections in " + + $"{RootReadmePath}, {LiteReadmePath} and {ShippedNoticePath} in the same change."); + } + + [Fact] + public void TheDocsNameARuntimeDownload_IfAndOnlyIfSomeArtifactIsFrameworkDependent() + { + /* The assertion #2499 got half right. It checked that the docs DID name the runtimes, so when + #2501 flipped the publish it stayed green while the prose went stale - the exact "docs drift + from the artifact" failure this file exists to stop, just in the other direction. Stated as + an if-and-only-if, both directions are covered by one fact: today every publish is + self-contained, so naming a runtime download in these three places is a demand for something + nobody needs; revert a publish and the same fact demands the sentences come back, with the + version derived from the build rather than remembered. */ + var anyFrameworkDependent = AllLitePublishShapes().Any(kv => !kv.Value); + var expected = DerivedInstallerNames(); + + Assert.True( + expected.Length >= 2, + $"Expected a framework-dependent Lite build to need at least the Desktop and ASP.NET Core " + + $"runtimes; derived only: {string.Join(", ", expected)}"); + + foreach (var (doc, text) in LiteRuntimeDocScopes()) + { + foreach (var installer in expected) + { + var named = text.Contains(installer, StringComparison.OrdinalIgnoreCase); + + if (anyFrameworkDependent) + { + Assert.True( + named, + $"{doc} never names '{installer}', which the built runtimeconfig says a " + + "framework-dependent Lite cannot start without. The .NET host reports only the " + + "FIRST missing framework, so a half-documented pair costs the reader a second " + + "identical failure."); + } + else + { + Assert.False( + named, + $"{doc} tells a reader to install '{installer}', but every Lite publish is " + + "--self-contained (#2501), so both artifacts carry their own runtime and there is " + + "nothing to install. Delete the sentence rather than softening it - a download " + + "instruction nobody needs is the #2489 defect with the sign flipped."); + } + } + } + } + + [Fact] + public void NoDocLineNamingAnArtifact_HangsARuntimeRequirementOnIt() + { + /* THE #2489 defect, generalised. It was README.md:95 reading "(requires .NET 10 Desktop + Runtime)" on the Setup.exe download line - the one artifact that needed nothing - which both + burdened the recommended path with a download and left the impression the requirement had + been handled. After #2501 the same is true of the ZIP line, so the rule is per-ARTIFACT + rather than per artifact-name: no line naming either download may carry a runtime demand, and + every such line must say self-contained so a reader can see why there is not one. */ + Assert.False( + AllLitePublishShapes().Any(kv => !kv.Value), + "A Lite publish is framework-dependent, so this fact's premise no longer holds. " + + "EveryShippedLiteArtifact_IsPublishedSelfContained names which one."); + + foreach (var artifact in new[] { SetupExeName, PortableZipName }) + { + var mentions = LinesMentioning(RootReadmePath, artifact) + .Concat(LinesMentioning(LiteReadmePath, artifact)) + .ToArray(); + + Assert.True(mentions.Length > 0, $"No doc line mentions {artifact}; the download instructions moved."); + + foreach (var line in mentions) + { + Assert.True( + line.Contains("self-contained", StringComparison.OrdinalIgnoreCase), + $"A line naming {artifact} does not say it is self-contained, which is the one fact that " + + $"tells a reader they need install nothing: {line.Trim()}"); + + foreach (var phrase in RuntimeDemandPhrases) + { + Assert.False( + line.Contains(phrase, StringComparison.OrdinalIgnoreCase), + $"A line naming {artifact} says '{phrase}', but it is packed from a --self-contained " + + $"publish and has no runtime prerequisite: {line.Trim()}"); + } + } + } + } + + [Fact] + public void SetupExeIsPackedFromTheVelopackPublish_AndTheZipFromTheOtherOne() + { + /* Both doc claims rest on which publish feeds which artifact, and that wiring is three hops of + YAML away from the sentence it justifies. Derived on both ends rather than pinned as + literals: the Velopack pack source and the release ZIP source each have to keep matching a + publish directory. Before #2501 the two publishes were told apart by their SHAPE; now that + both are self-contained the discriminator is the vpk pack line itself - whatever directory it + packs is the Velopack tree, and the ZIP has to come from the other one. Repoint either and + the docs become wrong silently, which is the failure this whole file exists for. */ + var shapes = LitePublishShapes(BuildWorkflowPath); + + Assert.True( + shapes.Count == 2, + $"Expected exactly two Lite publishes in {BuildWorkflowPath} (the ZIP's and the Velopack one); " + + $"found {shapes.Count}: {string.Join(", ", shapes.Keys)}."); + + var workflow = ParitySource.ReadFile(BuildWorkflowPath); + + var packLine = workflow.Split('\n') + .Select(l => l.Trim()) + .Single(l => l.Contains("vpk pack -u PerformanceMonitorLite", StringComparison.Ordinal)); + + var packTokens = packLine.Split(' ', StringSplitOptions.RemoveEmptyEntries); + var packFlag = Array.IndexOf(packTokens, "-p"); + Assert.True(packFlag >= 0 && packFlag + 1 < packTokens.Length, $"No '-p ' on the vpk pack line: {packLine}"); + + var velopackLeaf = DirectoryLeaf(packTokens[packFlag + 1]); + + Assert.Single(shapes.Keys, d => DirectoryLeaf(d).Equals(velopackLeaf, StringComparison.OrdinalIgnoreCase)); + var zipSourceDir = Assert.Single(shapes.Keys, d => !DirectoryLeaf(d).Equals(velopackLeaf, StringComparison.OrdinalIgnoreCase)); + + var zipLines = workflow.Split('\n') + .Select(l => l.Trim()) + .Where(l => l.Contains("Compress-Archive", StringComparison.Ordinal) && + l.Contains("PerformanceMonitorLite-", StringComparison.Ordinal)) + .ToArray(); + + Assert.True(zipLines.Length > 0, $"Nothing in {BuildWorkflowPath} builds a PerformanceMonitorLite ZIP."); + + foreach (var zipLine in zipLines) + { + var zipTokens = zipLine.Split(' ', StringSplitOptions.RemoveEmptyEntries); + var pathFlag = Array.IndexOf(zipTokens, "-Path"); + Assert.True(pathFlag >= 0 && pathFlag + 1 < zipTokens.Length, $"No '-Path ' on: {zipLine}"); + + /* Release re-zips from signed/Lite, the SignPath round-trip of publish/Lite; same leaf, and + that is the property worth asserting - the zip must never come from the velopack tree. */ + Assert.Equal( + DirectoryLeaf(zipSourceDir), + DirectoryLeaf(zipTokens[pathFlag + 1]), + ignoreCase: true); + } + } + + [Fact] + public void TheNightlyZipHasTheSameShapeAsTheReleaseZip_SoOneAnswerCoversBoth() + { + /* The nightly is the UAT download, and it publishes Lite itself rather than reusing build.yml's + step. If the two workflows ever disagree about the publish shape, the one answer the docs give + is wrong for one of them - and it is the nightly that would be wrong, because it is the + artifact with no Setup.exe alternative to fall back on. */ + var releaseShapes = LitePublishShapes(BuildWorkflowPath); + var nightlyShapes = LitePublishShapes(NightlyWorkflowPath); + + foreach (var (dir, nightlySelfContained) in nightlyShapes) + { + Assert.True( + releaseShapes.TryGetValue(dir, out var releaseSelfContained), + $"{NightlyWorkflowPath} publishes Lite to {dir}, which {BuildWorkflowPath} does not, so the " + + "two artifacts no longer share one documented answer."); + + Assert.Equal(releaseSelfContained, nightlySelfContained); + } + + var zipSourceDir = releaseShapes.Keys + .Single(d => !DirectoryLeaf(d).Contains("velopack", StringComparison.OrdinalIgnoreCase)); + + Assert.True( + nightlyShapes.ContainsKey(zipSourceDir), + $"{NightlyWorkflowPath} does not publish Lite to {zipSourceDir}, so the nightly ZIP is built from " + + "a different tree than the release ZIP."); + } + + [Fact] + public void TheCommittedLockFile_CoversEveryRuntimeIdentifierTheLitePublishesPin() + { + /* The blocker #2501 turned up, and the reason that change is three files rather than one. + + A RID publish restores a RID graph, and restore then REWRITES packages.lock.json to add a + "/" target. Both workflows run `dotnet restore --locked-mode` BEFORE + they publish (build.yml, nightly.yml), and locked mode compares the PROJECT's runtime + identifiers against the LOCK FILE's: with the RID living only on the publish command line + that comparison is "" against "win-x64", and the restore dies NU1004. Reproduced + locally, and it would fire on every PR - not some future --no-restore trap. + + The fix is in the csproj, so the project itself asks for the graph and + one committed lock file satisfies both the RID-less locked-mode restore and the RID publish. + This fact holds the three halves together: pin a new RID in a workflow and it stays red until + the csproj and the lock file follow. */ + var rids = PinnedRuntimeIdentifiers(); + + Assert.True( + rids.Length > 0, + $"No Lite publish in either workflow pins a RID with -r. If that is deliberate, the " + + $" line in {LiteProjectPath} and the RID targets in {LiteLockFilePath} are " + + "now dead weight and should go in the same change."); + + var csproj = ParitySource.ReadFile(LiteProjectPath); + var lockFile = ParitySource.ReadFile(LiteLockFilePath); + + foreach (var rid in rids) + { + Assert.True( + csproj.Contains("", StringComparison.Ordinal) && + csproj.Contains($"{rid}", StringComparison.Ordinal), + $"A Lite publish pins -r {rid}, but {LiteProjectPath} does not declare exactly that in " + + ". Without it the project's RID set is empty while the committed lock " + + $"file's is {rid}, and every --locked-mode restore in CI fails NU1004 before the publish " + + "ever runs."); + + Assert.True( + lockFile.Contains($"/{rid}\"", StringComparison.Ordinal), + $"A Lite publish pins -r {rid}, but {LiteLockFilePath} has no \"/{rid}\" target. " + + "Regenerate it with `dotnet restore Lite/PerformanceMonitorLite.csproj --force-evaluate` " + + "and commit the result; shipping a RID shape the committed lock file does not cover is a " + + "locked-mode failure waiting for the next CI run."); + } + } + + [Fact] + public void TheZipsSection_SaysItCarriesItsOwnRuntimeRatherThanListingDownloads() + { + /* Lite/README.md had no prerequisites section at all before #2499 - zero hits for "prerequisite", + ".NET 10", "Desktop Runtime" or "ASP.NET". Someone reading the Lite folder had nowhere to learn + any of this, which was half of #2489. #2501 did not delete the section, because the root README + links its anchor and because "nothing to install" is itself the answer a reader came for; it + changed what the section has to say. */ + var liteReadme = ParitySource.ReadFile(LiteReadmePath); + var anyFrameworkDependent = AllLitePublishShapes().Any(kv => !kv.Value); + + Assert.True( + liteReadme.Contains("## Prerequisites", StringComparison.Ordinal), + $"{LiteReadmePath} has no '## Prerequisites' section. The root README links to its anchor."); + + if (anyFrameworkDependent) + { + Assert.True( + liteReadme.Contains("framework-dependent", StringComparison.OrdinalIgnoreCase), + $"{LiteReadmePath} does not say which artifact is framework-dependent, which would be the " + + "reason it has prerequisites at all."); + } + else + { + Assert.True( + liteReadme.Contains("self-contained", StringComparison.OrdinalIgnoreCase), + $"{LiteReadmePath} does not say the artifacts are self-contained, which is the whole content " + + "of its Prerequisites section now that neither needs a runtime."); + } + + foreach (var doc in new[] { RootReadmePath, LiteReadmePath }) + { + Assert.True( + LinesMentioning(doc, PortableZipName).Length > 0, + $"{doc} never names the portable ZIP, so what it does and does not need is attached to nothing."); + } + } + + [Fact] + public void TheRuntimeNotice_ShipsInsideTheZipBesideTheExe() + { + /* Kept from #2499 with its subject changed rather than deleted. Lite cannot pre-check the way + Darling's install-darling.ps1 does: there is no install script, and before #2501 the host error + preceded our code. The unzipped folder is still the only surface a reader has once they are + looking at the files rather than the repo, so the notice still has to be COPIED to the publish + output - a file that exists in the repo and never ships is worse than none, because both + READMEs promise it is there. What it says changed with the publish shape; that it ships did not. */ + var noticePath = Path.Combine(ParitySource.RepoRoot(), ShippedNoticePath.Replace('/', Path.DirectorySeparatorChar)); + Assert.True(File.Exists(noticePath), $"{ShippedNoticePath} is missing."); + + var csproj = ParitySource.ReadFile(LiteProjectPath); + var noticeFileName = Path.GetFileName(ShippedNoticePath); + + var itemIndex = csproj.IndexOf($"", StringComparison.Ordinal); + Assert.True( + itemIndex >= 0, + $"PerformanceMonitorLite.csproj has no item, so it never lands " + + "beside the exe and both READMEs promise a file that is not in the ZIP."); + + var itemEnd = csproj.IndexOf("", itemIndex, StringComparison.Ordinal); + Assert.True(itemEnd > itemIndex, $"Unterminated item for {noticeFileName}."); + + Assert.True( + csproj[itemIndex..itemEnd].Contains("", StringComparison.Ordinal), + $"{noticeFileName} is declared but not copied to the output directory."); + + /* And it has to name both artifacts, because it is the copy a reader reaches for when they went + looking for a prerequisite and there is not one. */ + var notice = ParitySource.ReadFile(ShippedNoticePath); + foreach (var artifact in new[] { SetupExeName, PortableZipName }) + { + Assert.True( + notice.Contains(artifact, StringComparison.Ordinal), + $"{ShippedNoticePath} does not name {artifact}, so a reader cannot tell whether it is about " + + "the build they have."); + } + } +} diff --git a/Lite.Tests/LiteSidebarDotRendersTheCardStatusTests.cs b/Lite.Tests/LiteSidebarDotRendersTheCardStatusTests.cs new file mode 100644 index 000000000..acd260bff --- /dev/null +++ b/Lite.Tests/LiteSidebarDotRendersTheCardStatusTests.cs @@ -0,0 +1,224 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2458: the sidebar row's status dot stops deriving its own copy of the four-word status ladder. +/// +/// What was wrong. ServerConnection.DotStatus computed "Unknown"/"Online"/"Warning"/"Offline" +/// from its own (IsOnline, HasCollectorErrors) pair — the same ladder the Overview card derives, on a +/// different type, on a different surface, from a different instance of the same flags. The two could therefore +/// say different things about one server and nothing would notice. #2451 collapsed the card's five renderings onto +/// for exactly this reason and pinned the collapse — but the pin counted one literal +/// in one file, and this fourth copy was in a file that scan never opened. +/// +/// What is pinned here. That the two surfaces agree BY CONSTRUCTION rather than by observation — both +/// render — and that the dot now says what it means. The dot is the thing a +/// reader points at first, which is @ehaar's #2422 complaint one surface over from where it was reported. +/// +/// What is deliberately NOT here. Collection freshness. #2457 kept it out of the status word and gave +/// it its own banded row on the card, because folding it in recreates the #2429/#2422 conflation of a stale +/// collection with a failing one. holds that line at the +/// classifier's signature, and the tooltip tells the reader where the freshness answer actually lives. +/// +public sealed class LiteSidebarDotRendersTheCardStatusTests +{ + private static ServerConnection Connection(bool? isOnline, bool? hasCollectorErrors) => + new() { ServerName = "srv", IsOnline = isOnline, HasCollectorErrors = hasCollectorErrors }; + + private static ServerSummaryItem Card(bool? isOnline, bool hasCollectorErrors) => + new() { ServerName = "srv", IsOnline = isOnline, HasCollectorErrors = hasCollectorErrors }; + + /// + /// The defect itself: the sidebar and the card, handed the same flags, land on the same state and the same + /// word. Sampled over every reachable combination rather than the happy one, because the pairs that drifted on + /// #2429 and #2451 were both edge combinations nobody looked at. + /// + [Fact] + public void TheDot_AndTheCard_RenderOneLadder() + { + foreach (var isOnline in new bool?[] { true, false, null }) + { + foreach (var hasErrors in new[] { true, false }) + { + var dot = Connection(isOnline, hasErrors); + var card = Card(isOnline, hasErrors); + + Assert.Equal(card.CardStatus, dot.CardStatus); + Assert.Equal(card.StatusDisplay, dot.DotStatus); + Assert.StartsWith(dot.DotTooltip.Split('\n')[0], card.StatusTooltip, StringComparison.Ordinal); + } + } + } + + /// + /// The four words are unchanged, so the sidebar's DataTriggers keep painting the dot they always did. This is + /// the compatibility half of the collapse: Word() feeds a XAML DataTrigger Value= match, and a + /// word that stopped matching would fall through to the muted default rather than fail anything. + /// + [Fact] + public void TheDotWords_AreTheOnesTheSidebarPaints() + { + Assert.Equal("Online", Connection(true, false).DotStatus); + Assert.Equal("Warning", Connection(true, true).DotStatus); + Assert.Equal("Offline", Connection(false, false).DotStatus); + Assert.Equal("Unknown", Connection(null, false).DotStatus); + + /* An offline server's collector-error marker must not turn its dot amber — the flags resolve in order. */ + Assert.Equal("Offline", Connection(false, true).DotStatus); + + /* HasCollectorErrors is nullable on this type only. Null means nobody has established collector health, + which is not the same claim as "collectors are failing" — the string ladder folded it to Online and + that reading is kept. */ + Assert.Equal("Online", Connection(true, null).DotStatus); + + var xaml = ReadRepoFile(Path.Combine("Lite", "MainWindow.xaml")); + foreach (var word in new[] { "Online", "Offline", "Warning" }) + { + Assert.Contains( + "", + xaml, StringComparison.Ordinal); + } + } + + /// + /// The dot says what it means, in the card's words. The first line is byte-for-byte the card's, because that + /// is the whole argument for doing every surface: a reader moving between them meets one vocabulary rather + /// than three levels of helpfulness. + /// + [Fact] + public void TheDotTooltip_OpensOnTheSentenceTheCardOpensOn() + { + Assert.StartsWith("Online — the last connection check succeeded", Connection(true, false).DotTooltip, StringComparison.Ordinal); + Assert.StartsWith("Warning — one or more collectors are failing on this server", Connection(true, true).DotTooltip, StringComparison.Ordinal); + Assert.StartsWith("Offline — the last connection check failed", Connection(false, false).DotTooltip, StringComparison.Ordinal); + Assert.StartsWith("Unknown — this server has not been connection-checked yet", Connection(null, false).DotTooltip, StringComparison.Ordinal); + + /* And it ends on the gesture THIS surface supports. The card's line names the card; naming a single click + on either would be naming a no-op, since both handlers act only on a double-click. */ + Assert.EndsWith("Double-click the row to open this server's tab", Connection(true, false).DotTooltip, StringComparison.Ordinal); + Assert.EndsWith("Double-click the card to open this server's tab", Card(true, false).StatusTooltip, StringComparison.Ordinal); + + var xaml = ReadRepoFile(Path.Combine("Lite", "MainWindow.xaml")); + Assert.Contains("MouseDoubleClick=\"ServerListView_MouseDoubleClick\"", xaml, StringComparison.Ordinal); + } + + /// + /// The dot actually carries the tooltip. Removing the attribute compiles perfectly clean and silently returns + /// the sidebar to a coloured circle that will not say what it means, which is why this reads XAML — no + /// assertion about a C# object can reach into an element's attributes. + /// + [Fact] + public void TheSidebarDot_IsBoundToItsTooltip() + { + var xaml = ReadRepoFile(Path.Combine("Lite", "MainWindow.xaml")); + + var at = xaml.IndexOf("{Binding Server.Connection.DotStatus}", StringComparison.Ordinal); + Assert.True(at > 0, "the sidebar status dot is gone — find where it moved before editing this test"); + + /* Walk back to the Ellipse that owns those triggers and assert the ToolTip is on the element itself. */ + var open = xaml.LastIndexOf(" 0, "the sidebar status dot is no longer an Ellipse"); + + var element = xaml[open..xaml.IndexOf(">", open, StringComparison.Ordinal)]; + Assert.Contains("ToolTip=\"{Binding Server.Connection.DotTooltip}\"", element, StringComparison.Ordinal); + } + + /// + /// #2457 kept collection freshness out of the status word on purpose, and this dot is the surface where that + /// separation is easiest to lose: ServerConnection carries no last-collection time, so a green dot here + /// is a connection answer being read somewhere that offers no freshness answer at all. + /// + /// Two things hold the line. The classifier takes exactly two arguments, so freshness cannot be folded in + /// without changing a signature this test names — the failure mode being guarded is a plausible, well-meant + /// edit, not a typo. And the tooltip tells the reader where the freshness answer lives instead of leaving them + /// to infer it from a colour, which is the #2429/#2422 conflation in miniature. + /// + [Fact] + public void TheDot_CannotBandFreshnessEvenByAccident() + { + const string disclaimer = + "It reports the connection check only, not collection freshness — the Overview card's Last Collect row bands that."; + + foreach (var isOnline in new bool?[] { true, false, null }) + { + Assert.Contains(disclaimer, Connection(isOnline, false).DotTooltip, StringComparison.Ordinal); + } + + var rules = ReadRepoFile(Path.Combine("Lite", "Services", "LocalDataService.Overview.cs")); + Assert.Contains( + "public static ServerCardStatus Classify(bool? isOnline, bool hasCollectorErrors) => isOnline switch", + rules, StringComparison.Ordinal); + + /* And the sidebar has nothing to band it FROM, which is why unifying the ladder was the whole of #2458 and + plumbing a collection time to this surface was split off. If that ever changes, this assertion is the + place the decision gets made rather than discovered. */ + var sidebar = ReadRepoFile(Path.Combine("Lite", "Models", "ServerConnection.cs")); + Assert.DoesNotContain("LastCollection", sidebar, StringComparison.Ordinal); + } + + /// + /// The pin that would actually have caught this one. #2451 counted the literal "IsOnline switch" + /// at exactly one occurrence, in exactly one file — and the fourth copy it existed to forbid was already + /// sitting in ANOTHER file, written as a chain of if statements rather than a switch. It evaded + /// that pin on both axes at once, and went on evading it through #2451 and #2457. + /// + /// So the invariant is not "one switch". It is that the four words are WRITTEN once. Scanning for + /// them as string literals is syntax-agnostic: an if chain, a switch, a dictionary and a ternary + /// all have to spell them, and none of them can spell them here any more. Comments are stripped first + /// because DotStatus's own doc comment legitimately names all four while the code writes none. + /// + [Fact] + public void TheFourWords_AreWrittenInExactlyOnePlace() + { + var sidebar = WithoutComments(ReadRepoFile(Path.Combine("Lite", "Models", "ServerConnection.cs"))); + + foreach (var word in new[] { "Online", "Warning", "Offline", "Unknown" }) + { + Assert.DoesNotContain("\"" + word + "\"", sidebar, StringComparison.Ordinal); + } + + /* And they are written in the one function both surfaces render. */ + var rules = ReadRepoFile(Path.Combine("Lite", "Services", "LocalDataService.Overview.cs")); + Assert.Contains( + "public static string Word(this ServerCardStatus status) => status switch", + rules, StringComparison.Ordinal); + } + + /// Drops line comments so a literal scan is neither defeated nor falsely tripped by prose. + /// ServerConnection.cs carries exactly one block comment — the licence header at the top — so + /// dropping //-prefixed lines is sufficient and cannot eat a string literal. + private static string WithoutComments(string source) => + string.Join("\n", source.Split('\n').Where(line => !line.TrimStart().StartsWith("//", StringComparison.Ordinal))); + + /// Locates a repo file by walking up from this file's compile-time path — the same helper + /// LiteOverviewCardExplainsItselfTests uses, and for the same reason: the assertions above are about + /// text that lives in files, not about objects. + private static string ReadRepoFile(string relative, [CallerFilePath] string thisFile = "") + { + for (var dir = new DirectoryInfo(Path.GetDirectoryName(thisFile)!); dir is not null; dir = dir.Parent) + { + var candidate = Path.Combine(dir.FullName, relative); + if (File.Exists(candidate)) + { + return File.ReadAllText(candidate).Replace("\r\n", "\n", StringComparison.Ordinal); + } + } + + throw new FileNotFoundException($"Could not locate {relative} walking up from {thisFile}"); + } +} diff --git a/Lite.Tests/LockWaitTrendToolTests.cs b/Lite.Tests/LockWaitTrendToolTests.cs new file mode 100644 index 000000000..f35843978 --- /dev/null +++ b/Lite.Tests/LockWaitTrendToolTests.cs @@ -0,0 +1,228 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_lock_wait_trend (#2484), the twin of Darling's. The viewer's Blocking-Trends lock-wait lane +/// had no read on either SKU: get_wait_trend charts ONE named wait type, and this is the whole LCK family +/// at once as a per-second rate. +/// +/// Three properties carry the weight, and none of them is "the SQL runs". The empty branch must NOT +/// be filtered the way the read is — a server collected for months that never took a lock wait is the +/// all-clear this branch exists to give, and an LCK-filtered probe would call it uncollected. The rate must +/// survive being fractional, because #2507 shipped an integer rate and a server at 0.4 a second reported +/// zero. And the anchor must reach the query, proven by CONTENT rather than by signature. +/// +public sealed class LockWaitTrendToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "LockWaitSrv"; + + /* Lite DERIVES its server id from the storage name; a hardcoded one would seed rows the tool looks + straight past and pass the never-collected assertion for the wrong reason. */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public LockWaitTrendToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-lockwait-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task NoLockWaits_MeansQuietOnlyWhenWaitStatsWereCollected() + { + var service = new LocalDataService(_duckDb); + + /* 1. nothing sampled at all: NOT "this server has no lock contention". */ + var never = Root(await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 4)); + Assert.Equal("unavailable", never.GetProperty("status").GetString()); + var neverText = never.GetProperty("message").GetString()!; + Assert.Contains("NOT a report of a server without lock contention", neverText, StringComparison.Ordinal); + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + + /* + 2. wait stats collected, and not one LOCK wait among them. The single most important assertion + here: this server is healthy and monitored, and the honest answer is a genuine all-clear. An + existence probe carrying the read's own LIKE 'LCK%' filter would find nothing and report this + server as uncollected, sending someone to fix collection that is working. + */ + await SeedWaitAsync(Truncate(DateTime.UtcNow).AddMinutes(-30), "CXPACKET", 999_999); + + var noLocks = Root(await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 4)); + Assert.Equal("empty", noLocks.GetProperty("status").GetString()); + var noLocksText = noLocks.GetProperty("message").GetString()!; + Assert.Contains("genuinely quiet rather than broken", noLocksText, StringComparison.Ordinal); + + /* Same zero rows as the branch above, and it must NOT reach for the same word. */ + Assert.DoesNotContain("EVER", noLocksText, StringComparison.Ordinal); + } + + [Fact] + public async Task TheRate_IsPerSecond_Fractional_AndDropsCounterResets() + { + var service = new LocalDataService(_duckDb); + var first = Truncate(DateTime.UtcNow).AddMinutes(-30); + var second = first.AddSeconds(60); + + await SeedWaitAsync(first, "LCK_M_X", 1_200); + await SeedWaitAsync(second, "LCK_M_X", 6_000); + + /* Three milliseconds over sixty seconds is 0.05 ms/sec — a real rate that integer division would + report as zero, which is how a quiet server reads as an idle one. */ + await SeedWaitAsync(first, "LCK_M_S", 10); + await SeedWaitAsync(second, "LCK_M_S", 3); + + /* A negative delta is the counter reset across a SQL Server restart, not a negative wait. */ + await SeedWaitAsync(second, "LCK_M_U", -500); + + /* Filtered out by the read even though it is the largest delta in the window. */ + await SeedWaitAsync(second, "CXPACKET", 999_999); + + var root = Root(await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 4)); + Assert.Equal(ServerName, root.GetProperty("server").GetString()); + + var trend = root.GetProperty("trend").EnumerateArray().ToArray(); + Assert.DoesNotContain(trend, r => r.GetProperty("wait_type").GetString() == "CXPACKET"); + Assert.DoesNotContain(trend, r => r.GetProperty("wait_type").GetString() == "LCK_M_U"); + + /* Four rows: two wait types x two collections. The FIRST collection of each type has no prior + sample to difference against, so its interval is NULL and its rate is 0 rather than the raw + delta — the LAG is per wait type, which is what stops one type's cadence describing another. */ + Assert.Equal(4, trend.Length); + + Assert.Equal(100d, RateOf(trend, "LCK_M_X", second), 3); + Assert.Equal(0d, RateOf(trend, "LCK_M_X", first), 3); + + /* The fractional rate. Asserted as > 0 as well as by value, because "0.05" and "0" differ by a cast + and the point of the assertion is that the cast is there. */ + var tinyRate = RateOf(trend, "LCK_M_S", second); + Assert.True(tinyRate > 0, $"a 3 ms delta over 60 s must not truncate to zero, got {tinyRate}"); + Assert.Equal(0.05d, tinyRate, 3); + } + + /// + /// The anchor moves the window and the resolved instant reaches the QUERY. + /// #2495's own failure mode is a tool that takes as_of, validates it, refuses a bad one + /// correctly, and then queries NOW. So this proves it by CONTENT: lock waits 30 hours old are outside + /// every default window on the surface, and only the anchored call can see them. + /// + [Fact] + public async Task TheAnchor_MovesTheWindow_AndTheDefaultAnchorCannotSeeAPastIncident() + { + var service = new LocalDataService(_duckDb); + var incident = Truncate(DateTime.UtcNow).AddHours(-30); + + await SeedWaitAsync(incident, "LCK_M_IX", 600); + await SeedWaitAsync(incident.AddSeconds(60), "LCK_M_IX", 1_800); + + var anchor = DateTime.SpecifyKind(incident.AddSeconds(60), DateTimeKind.Utc).ToString("o"); + var anchored = Root(await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 1, anchor)); + var rows = anchored.GetProperty("trend").EnumerateArray().ToArray(); + + Assert.Equal(2, rows.Length); + Assert.All(rows, r => Assert.Equal("LCK_M_IX", r.GetProperty("wait_type").GetString())); + Assert.Equal(30d, RateOf(rows, "LCK_M_IX", incident.AddSeconds(60)), 3); + + /* The same LENGTH of window at the default anchor cannot reach it — so it is the anchor doing the + work, not hours_back. */ + var unanchored = Root(await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 1)); + Assert.Equal("empty", unanchored.GetProperty("status").GetString()); + + /* An anchor we cannot use is refused, never silently treated as now. */ + var bad = await McpBlockingTools.GetLockWaitTrend(service, _serverManager, ServerName, 1, "last tuesday"); + Assert.Contains("Invalid as_of", bad, StringComparison.Ordinal); + } + + /// The rate the read returned for one (wait type, collection) pair, asserted to exist. + private static double RateOf(JsonElement[] rows, string waitType, DateTime collectionTimeUtc) + { + var stamp = collectionTimeUtc.ToString("yyyy-MM-ddTHH:mm:ss"); + var row = rows.FirstOrDefault(r => + r.GetProperty("wait_type").GetString() == waitType && + r.GetProperty("collection_time").GetString()!.StartsWith(stamp, StringComparison.Ordinal)); + + Assert.True( + row.ValueKind == JsonValueKind.Object, + $"no {waitType} row at {stamp} — the read returned [{string.Join(", ", rows.Select(r => r.GetProperty("wait_type").GetString() + "@" + r.GetProperty("collection_time").GetString()))}]"); + + return row.GetProperty("wait_time_ms_per_second").GetDouble(); + } + + private static JsonElement Root(string json) => JsonDocument.Parse(json).RootElement; + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedWaitAsync(DateTime collectionTime, string waitType, long deltaMs) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO wait_stats + (collection_id, collection_time, server_id, server_name, wait_type, + waiting_tasks_count, wait_time_ms, signal_wait_time_ms, + delta_waiting_tasks, delta_wait_time_ms, delta_signal_wait_time_ms) +VALUES ($1, $2, $3, $4, $5, 0, 0, 0, 1, $6, 0)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTime, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = waitType }); + cmd.Parameters.Add(new DuckDBParameter { Value = deltaMs }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/McpAlertSettingsKeyTests.cs b/Lite.Tests/McpAlertSettingsKeyTests.cs index 2d2a36204..5d1ffba87 100644 --- a/Lite.Tests/McpAlertSettingsKeyTests.cs +++ b/Lite.Tests/McpAlertSettingsKeyTests.cs @@ -11,6 +11,8 @@ using System.IO; using System.Linq; using System.Text.Json; +using System.Text.RegularExpressions; +using PerformanceMonitor.Notifications; using PerformanceMonitorLite; using PerformanceMonitorLite.Mcp; using Xunit; @@ -26,9 +28,11 @@ namespace Lite.Tests; /// just as happily if someone re-added the old key beside the new one, which is the likelier accident than /// deleting the new one — and for an MCP client, two keys meaning the same thing is its own bug. /// -/// Runtime rather than source-parsing: the payload is an anonymous type serialized by -/// JsonSerializer with a naming policy in McpHelpers.JsonOptions, so the C# identifier is not -/// automatically the wire key. Only serializing it actually proves what a client receives. +/// Runtime rather than source-parsing: the payload is an anonymous type handed to +/// JsonSerializer with the SHARED McpHelpers.JsonOptions, and what that turns a C# identifier +/// into is that object's business, not this file's — it carries no naming policy today, and the day it +/// acquires one every key here changes without a line of this payload being touched. Only serializing it +/// actually proves what a client receives. /// /// #1965: because it is runtime, reads whatever the App.Alert* statics hold at /// that instant, so this class shares the "app-alert-statics" collection with the classes that write them — @@ -36,6 +40,12 @@ namespace Lite.Tests; /// App.LoadAlertSettings). xUnit runs separate classes in parallel, and a foreign write of /// CpuAlertMode.Total landing between this class's set and its assert failed the SqlOnly case. The /// finally restore below could not help: the window is before the assert, not after it. +/// +/// #2394 widened it from pinning three renamed keys to pinning the whole SHAPE. Lite reported four +/// groups where Darling reports nineteen, so the drift this class was written to catch had already happened +/// on a scale no per-key assertion would notice. The parity assertions below therefore DERIVE Darling's shape +/// from Darling's source rather than transcribing it — a hand-copied list of nineteen groups is precisely the +/// artifact that stays green on the day a twentieth arrives. /// [Collection("app-alert-statics")] public sealed class McpAlertSettingsKeyTests @@ -122,6 +132,249 @@ private static string FindRepoFile(string relativePath) throw new FileNotFoundException($"Could not locate {relativePath} walking up from {AppContext.BaseDirectory}"); } + /// Darling's one group Lite deliberately does not report. Named once so the omission reads as a + /// decision in both places that reference it. + private const string SelfAlertsGroup = "self_alerts"; + + /// #2417: the MEMBER-level counterpart to — Darling keys Lite + /// genuinely has no equivalent for, exempted BY NAME so the hole is a decision someone justified here + /// rather than a loosened assertion. Each entry is paid for by a test asserting the omission is still + /// real, so an exemption cannot outlive its reason. + /// + /// EMPTY, and the emptiness is asserted in + /// rather than merely + /// being true today. Its one entry was ag.disconnect_refire_minutes, Darling's #1696 / + /// store-V37 knob, which Lite had no equivalent for at all — no static, no settings.json key, and no + /// edge state that could re-announce a still-disconnected replica. #2426 built the re-fire, so the + /// exemption came out with it and all four AG members compare. The seam stays for the next such + /// case: adding an entry means writing the test that pays for it. + private static readonly string[] LiteOmittedMembers = Array.Empty(); + + /// + /// Darling's BuildAlertSettingsPayload shape read out of Darling's SOURCE — each top-level key in + /// document order with its nested keys (an empty list for a scalar like cooldown_minutes). Derived + /// rather than transcribed for the reason in the class summary, and the same reasoning that makes + /// read Darling's file instead of asserting Lite's + /// own constants back at itself. + /// + private static IReadOnlyList>> DarlingPayloadShape() + { + var source = File.ReadAllText(FindRepoFile(Path.Combine( + "Darling", "PerformanceMonitor.Darling.Service", "Mcp", "DarlingMcpAlertTools.cs"))); + + /* Comments come out first: the payload carries several, and the "(0 = off)" inside one would + otherwise scan as a key. Nothing between the anchor and the initializer's closing brace is a + string literal, so comments are the only C# escape this has to understand. */ + source = Regex.Replace(source, @"/\*.*?\*/", " ", RegexOptions.Singleline); + source = Regex.Replace(source, "//[^\r\n]*", " "); + + /* Anchored on the DEFINITION rather than the name: get_alert_settings CALLS + BuildAlertSettingsPayload earlier in the file, so a bare name search would brace-match that call + site's catch block and silently return the wrong object. */ + var start = source.IndexOf("private static object BuildAlertSettingsPayload", StringComparison.Ordinal); + Assert.True(start >= 0, "Darling's BuildAlertSettingsPayload definition could not be located."); + + var shape = new List>>(); + List? nested = null; + var depth = 0; + + for (var i = source.IndexOf('{', start); i >= 0 && i < source.Length; i++) + { + var c = source[i]; + if (c == '{') + { + depth++; + continue; + } + + if (c == '}') + { + depth--; + if (depth == 0) break; + continue; + } + + if ((depth != 1 && depth != 2) || !(char.IsLetter(c) || c == '_')) continue; + + /* The tail of an identifier already consumed, or a member access (s.CpuEnabled) - not a key. */ + var previous = i > 0 ? source[i - 1] : ' '; + if (char.IsLetterOrDigit(previous) || previous == '_' || previous == '.') continue; + + var end = i; + while (end < source.Length && (char.IsLetterOrDigit(source[end]) || source[end] == '_')) end++; + + var after = end; + while (after < source.Length && char.IsWhiteSpace(source[after])) after++; + + /* An identifier followed by a single '=' is an initializer key. Depth 1 opens a group, depth 2 + fills the one it opened; document order makes that association exact without a stack. */ + if (after < source.Length && source[after] == '=' && + (after + 1 >= source.Length || source[after + 1] != '=')) + { + var name = source[i..end]; + if (depth == 1) + { + nested = new List(); + shape.Add(new KeyValuePair>(name, nested)); + } + else + { + nested?.Add(name); + } + } + + i = end - 1; + } + + /* The harness proves itself before anything is trusted to it. A parse that quietly returned nothing + would make every assertion built on it vacuously true, which is worse than having no check at all. */ + var groups = shape.Select(g => g.Key).ToList(); + Assert.InRange(shape.Count, 15, 40); + Assert.Contains("cpu", groups); + Assert.Contains("analysis", groups); + Assert.Contains(SelfAlertsGroup, groups); + Assert.DoesNotContain("smtp", groups); + Assert.Equal(new[] { "enabled", "threshold_percent", "mode" }, shape.Single(g => g.Key == "cpu").Value); + + return shape; + } + + /// + /// #2394: Lite reported four groups — cpu, blocking, deadlocks, smtp — where Darling reports nineteen, + /// even though the SHARED alert engine was already evaluating every one of them here through + /// AppAlertEngineSettings. Nothing about Lite's alerting was narrower; only the MCP surface was, so + /// an agent triaging a Lite instance could not read whether tempdb-space, low-disk, PVS, file-growth, + /// long-running-query/job, failed-job, database-state or analysis alerting was even switched on. + /// Both directions are asserted. A key Lite emits that Darling does not is as much a defect as a + /// missing one: two spellings of the same setting across the two apps is exactly the #1839/#1911 class of + /// bug this file exists to stop, and only smtp is a legitimate Lite addition. + /// + [Fact] + public void GetAlertSettings_ReportsEveryGroupDarlingDoes_SpelledDarlingsWay() + { + var darling = DarlingPayloadShape(); + var root = Settings(); + var problems = new List(); + + foreach (var (group, darlingKeys) in darling) + { + if (group == SelfAlertsGroup) continue; + + if (!root.TryGetProperty(group, out var element)) + { + problems.Add($"missing group '{group}'"); + continue; + } + + /* A scalar (cooldown_minutes, excluded_databases) has no nested keys to compare. */ + if (darlingKeys.Count == 0) continue; + + var liteKeys = KeysOf(element); + problems.AddRange(darlingKeys + .Where(k => !LiteOmittedMembers.Contains($"{group}.{k}", StringComparer.Ordinal)) + .Except(liteKeys) + .Select(k => $"missing '{group}.{k}'")); + problems.AddRange(liteKeys.Except(darlingKeys).Select(k => $"'{group}.{k}' is Lite-only")); + } + + Assert.True( + problems.Count == 0, + "Lite's get_alert_settings has drifted from Darling's shape: " + string.Join("; ", problems)); + + /* #2426: nothing is exempted today, and that is asserted rather than merely true — an entry added + to LiteOmittedMembers without the test that justifies it would silently narrow this comparison, + which is the drift this whole class exists to stop. */ + Assert.Empty(LiteOmittedMembers); + + /* smtp is Lite's ONE addition — Lite delivers its own email where Darling manages delivery + credentials outside the settings row. Pinned as an exact set so a second Lite-only group cannot be + added without this test being the place someone justifies it. */ + Assert.Equal( + new[] { "smtp" }, + KeysOf(root).Except(darling.Select(g => g.Key)).ToArray()); + } + + /// + /// #2426, and the inversion of the exemption that stood here: ag.disconnect_refire_minutes was + /// the one MEMBER Lite could not report, because it had no AG disconnect re-fire at all — a replica + /// disconnected for a week announced itself exactly once, where Darling re-announced it. Both halves + /// of the old exemption are now asserted the other way round. Darling must still emit it (or the + /// comparison is over a key nobody publishes), and Lite must too, carrying its own live value rather + /// than the constant 0 that was rejected for telling an agent it can tune something the app cannot. + /// + /// The value assertion is the half that matters most and the half a key-presence check would + /// miss entirely: it is what distinguishes a real knob from the placeholder this PR exists to avoid + /// shipping. + /// + [Fact] + public void GetAlertSettings_ReportsAgDisconnectRefire_WithLitesOwnValue() + { + var original = App.AgDisconnectRefireMinutes; + try + { + App.AgDisconnectRefireMinutes = 17; + + Assert.Contains("disconnect_refire_minutes", DarlingPayloadShape().Single(g => g.Key == "ag").Value); + + var ag = Settings().GetProperty("ag"); + Assert.Contains("disconnect_refire_minutes", KeysOf(ag)); + Assert.Equal(17, ag.GetProperty("disconnect_refire_minutes").GetInt32()); + } + finally + { + App.AgDisconnectRefireMinutes = original; + } + } + + /// + /// The one group Lite deliberately does NOT report, asserted so the hole reads as a decision rather than + /// the oversight it would otherwise look like. AppAlertEngineSettings returns shipped constants for + /// three of self_alerts' four members precisely because a single-instance WPF app has no headless store + /// volume and no fleet collection loop to self-monitor, and Lite has no concept whatsoever of the fourth, + /// store_job_cadence_warn_percent. Reporting constants under names that read as knobs would tell an agent + /// it can tune something Lite cannot. + /// + [Fact] + public void GetAlertSettings_OmitsSelfAlerts_WhichLiteHasNoEquivalentFor() + { + Assert.Contains(SelfAlertsGroup, DarlingPayloadShape().Select(g => g.Key)); + Assert.DoesNotContain(SelfAlertsGroup, KeysOf(Settings())); + } + + /// + /// The value half for delivery.mode — the second value-level alignment after cpu.mode, and + /// the one place ToString() is the right answer rather than a mapping: + /// is the SHARED enum both SKUs run on, and Darling's store holds literally its ToString(), so the + /// two apps cannot drift the way Lite's app-local CpuAlertMode could. + /// Pinned against Darling's ACCEPTED vocabulary rather than against the enum, because a rename that + /// moved the shared enum would move Lite's emitted value with it and an enum-derived assertion would + /// happily follow — while Darling's validator, which holds the two names as literals, would not. + /// + [Theory] + [InlineData(AlertNotificationMode.Summary, "Summary")] + [InlineData(AlertNotificationMode.PerEvent, "PerEvent")] + public void GetAlertSettings_DeliveryMode_IsDarlingsAcceptedVocabulary(AlertNotificationMode mode, string expected) + { + var original = App.AlertDeliveryMode; + try + { + App.AlertDeliveryMode = mode; + + Assert.Equal(expected, Settings().GetProperty("delivery").GetProperty("mode").GetString()); + + var darlingTools = File.ReadAllText(FindRepoFile(Path.Combine( + "Darling", "PerformanceMonitor.Darling.Service", "Mcp", "DarlingMcpAlertTools.cs"))); + Assert.Contains( + "AddEnum(\"delivery_mode\", n, \"delivery.mode\", \"Summary\", \"PerEvent\")", + darlingTools, + StringComparison.Ordinal); + } + finally + { + App.AlertDeliveryMode = original; + } + } + [Fact] public void GetAlertSettings_TopLevelMasterSwitch_IsAlertsEnabled() { diff --git a/Lite.Tests/McpMissMessageParityPinTests.cs b/Lite.Tests/McpMissMessageParityPinTests.cs new file mode 100644 index 000000000..bccbedae7 --- /dev/null +++ b/Lite.Tests/McpMissMessageParityPinTests.cs @@ -0,0 +1,169 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using Xunit; + +namespace Lite.Tests; + +/// +/// The miss SENTENCES the two SKUs share, pinned against both source trees. +/// +/// #2485 had two halves. One was reads that could not say which kind of nothing they had found; the +/// other was three tools that answered the same question differently depending on which SKU the client was +/// pointed at. The first half is fixable in one place per tool. The second is not: every shared sentence +/// lives twice, once per SKU, and nothing stops one copy being reworded on its own — which is how the +/// divergence this issue exists to close got there in the first place. +/// +/// So the sentences are pinned as SOURCE, in both trees at once. A fragment listed here must appear in +/// Lite/Mcp AND in Darling/PerformanceMonitor.Darling.Service/Mcp; reword one copy and this +/// fails naming the tree that no longer has it. Fragments are chosen to sit BETWEEN interpolation holes, so +/// they are the literal bytes both SKUs emit rather than an approximation of them. +/// +/// This is a per-change pin, not a survey: it holds the sentences this repo has deliberately made +/// shared, and each change that adds one is expected to add it here. It does not claim to enumerate every +/// message either server can produce, and a naive extension that tried to would pass vacuously the day +/// somebody added an unshared one. +/// +public sealed class McpMissMessageParityPinTests +{ + /// + /// #2559 wiring, per SKU. The shared builder guarantees the SENTENCES match; what it cannot guarantee is + /// that both tool bodies actually CALL it, which is the drift the review caught on the first draft of + /// that change — Darling grew the arm and Lite did not, so a Lite login with no msdb access kept getting + /// the affirmative "no jobs running" claim the arm exists to remove. + /// + /// Order matters as much as presence: the gate inference must sit after the collector's own + /// recorded denial (specific evidence beats an inference from an absence) and before the empty miss. + /// Anchored on CODE rather than on the sentences, because both files contain comments that quote them. + /// + [Theory] + [InlineData("Lite/Mcp/McpJobTools.cs", "McpRuntimePrecondition")] + [InlineData("Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingMcpJobTools.cs", "DarlingRuntimePrecondition")] + public void BothSkus_AskAboutTheGate_AfterTheRecordedOutcome_AndBeforeTheEmptyMiss(string relativePath, string helper) + { + var source = File.ReadAllText(Path.Combine(ParitySource.RepoRoot(), relativePath)); + + var recorded = source.IndexOf($"{helper}.StatusAsync", StringComparison.Ordinal); + var gate = source.IndexOf($"{helper}.GatedOffStatusAsync", StringComparison.Ordinal); + var empty = source.IndexOf("McpHelpers.Status(\"empty\"", StringComparison.Ordinal); + + Assert.True(gate > 0, $"{relativePath} never asks whether the collector is gated off (#2559)"); + Assert.True(recorded > 0 && empty > 0, $"{relativePath} no longer has the arms this pin describes"); + Assert.True(gate > recorded, $"{relativePath} lets the gate inference pre-empt a recorded denial"); + Assert.True(empty > gate, $"{relativePath} no longer keeps the empty miss as the last resort"); + } + + private const string LiteMcpDir = "Lite/Mcp"; + private const string DarlingMcpDir = "Darling/PerformanceMonitor.Darling.Service/Mcp"; + + /// + /// Sentence fragments that must read identically on both SKUs. Each sits between interpolation holes, so + /// what is compared is the literal text a caller receives. + /// + public static TheoryData SharedMissFragments() => new() + { + /* get_wait_types */ + "This server HAS collected wait stats before, so this window is genuinely quiet rather than broken — widen hours_back to find the most recent samples.", + "Delta wait stats need a SECOND collection cycle before the first row exists, so on a newly added server this clears itself; otherwise check that collection is running and that the server is enabled.", + + /* get_memory_clerks */ + "This read returns the LATEST snapshot rather than a window, so an empty result is never a quiet period — a live SQL Server always has memory clerks.", + + /* get_mute_rules */ + "No mute rules are configured for this store, so no alert is being suppressed anywhere — a quiet alert history is genuine rather than muted.", + "is disabled or expired, so nothing is being suppressed. Pass enabled_only=false to list them — this is a lapsed mute, not an absent one.", + + /* compare_analysis */ + "in EITHER window, so there is nothing to compare — this is NOT a report that nothing changed.", + "The BASELINE window produced no facts at all, so every fact below counts as a new issue only because there was nothing to compare it against.", + "The COMPARISON window produced no facts at all, so every fact below counts as a resolved issue only because there is nothing in the recent window to compare against.", + + /* get_query_heatmap (#2484) — the three empty branches, one of which (a collected but IDLE + window) no other read has. */ + "so this is NOT a report of a quiet server — there is nothing to draw. query_stats is a PERIODIC table rather than an edge table: the collector writes rows every cycle for whatever is in the plan cache, so an empty history means nobody looked. Check get_collection_health for this server.", + " hour(s), so the grid has no columns rather than no hot cells. Widen hours_back, or check get_collection_health — a collector that stopped looks exactly like this.", + " hour(s), but no capture recorded an execution: every row carried a zero execution delta, so nothing lands on the grid. A server that is up and idle looks exactly like this, and so does a database_name filter matching nothing collected. Delta-based collection also needs a SECOND cycle before the first non-zero row exists.", + + /* get_lock_wait_trend (#2484). The all-clear sentence is get_wait_types' own, reused deliberately: + both reads are looking at the same PERIODIC table and the advice is identical, so two spellings + of it would be a divergence for nothing. Only the never-collected half needs its own words, + because "no lock contention" is the wrong thing to hear about a server nothing was stored for. */ + "No lock waits recorded for ", + ", so this is NOT a report of a server without lock contention — nothing has been stored for it at all. ", + + /* get_daily_summary_range (#2484) — the two empty branches. The first is the one that is easy to + get wrong: a day with ANY collection appears even when quiet, so no days at all cannot mean a + quiet stretch. */ + ". A day with ANY collection appears here even when every signal was quiet, so this range is outside what the store holds for this server rather than a stretch of quiet days — widen days_back, or move as_of.", + ", so the calendar is empty because nothing has been collected — not because those days were quiet. Check that the service is running and that the server is enabled for collection.", + + /* The instructions' miss-vocabulary paragraph (#2511). The engine-gap MESSAGE itself is built by + CollectorEngineCapability and is byte-identical by construction rather than by pinning; what lives + twice, and therefore belongs here, is the paragraph that teaches a caller how to read it. */ + "`not_collected` means this server does not collect that at all — and when the reason is the ENGINE, the gap is PERMANENT", + + /* The PostgreSQL half of the same paragraph (#2532). An agent over MCP has no tabs: it asks a + read by name, so the instructions are the only place it can be told which family answers on + this engine. */ + "a PostgreSQL target collects none of the SQL Server signals at all, and the `get_pg_*` reads are the ones that answer there", + + /* The FOURTH word (#2546). The precondition MESSAGES themselves are built by + CollectorRuntimePrecondition and are byte-identical by construction; what lives twice, and + therefore belongs here, is the paragraph teaching a caller that this one is the one they can act + on — and specifically that it is re-derived per read, so doing what it asks is enough. An agent + told to restart the monitoring service instead would be back at the defect the word exists to + close. */ + "`precondition` is the one that IS worth acting on: this server could have that data, the collector is running, and a setup step on the monitored server is in the way", + "It is re-derived on EVERY read rather than decided when the connection was made, so once somebody does the thing it asked for the next call answers with data", + + /* #2559. The paragraph above USED to end "there is nothing to restart on the monitoring side" as a + flat promise, and for connect-scoped preconditions that is false — the fact that gates them is read + once at connect and cached for the connection's life, so an agent following the general case sends + the user round a loop that never terminates. The correction has to be identical on both SKUs or one + of them keeps teaching the wrong rule, which is exactly what this pin is for. */ + "A few preconditions are the exception and SAY SO IN THEIR OWN MESSAGE: the fact that gates them is read once when the service connects to that server and cached for the connection's life", + "telling somebody to retry a connect-scoped one without reconnecting sends them round a loop that never terminates", + + /* The gated-off arm's GATE CANDIDATES (#2559). The message body itself comes from the shared + CollectorRuntimePrecondition and is byte-identical by construction, so it does not belong here — + but this sentence is supplied by each tool body at its own call site, lives twice, and is exactly + what drifts. A tree missing it is a tree whose get_running_jobs never grew the arm at all. */ + "tables are not reachable to a monitoring login at all and no grant changes that.", + }; + + [Theory] + [MemberData(nameof(SharedMissFragments))] + public void EverySharedMissSentence_ReadsIdenticallyOnBothSkus(string fragment) + { + Assert.True( + AppearsIn(LiteMcpDir, fragment), + $"Lite/Mcp no longer contains the shared sentence: \"{fragment}\""); + + Assert.True( + AppearsIn(DarlingMcpDir, fragment), + $"Darling's MCP tools no longer contain the shared sentence: \"{fragment}\""); + } + + /// + /// A non-vacuous floor. Without it, a fragment list that drifted into naming text neither tree contains + /// would still be a green suite the day somebody emptied it — the guard-that-stopped-guarding shape. + /// + [Fact] + public void ThePinCoversTheSentencesThisChangeMadeShared() + { + Assert.True(SharedMissFragments().Count() >= 11, + "the shared-sentence pin lost entries; a sentence that stops being pinned can drift between the SKUs unnoticed"); + } + + private static bool AppearsIn(string relativeDir, string fragment) => + ParitySource.EnumerateCsFiles(relativeDir) + .Any(f => File.ReadAllText(f).Contains(fragment, StringComparison.Ordinal)); +} diff --git a/Lite.Tests/McpSettingsGuardTests.cs b/Lite.Tests/McpSettingsGuardTests.cs new file mode 100644 index 000000000..1858c3c51 --- /dev/null +++ b/Lite.Tests/McpSettingsGuardTests.cs @@ -0,0 +1,345 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.RegularExpressions; +using PerformanceMonitorLite.Mcp; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2431. McpSettings.Load ended in a bare catch whose fallback is MCP OFF, so the same +/// trailing comma #2425 fixed elsewhere also revoked an endpoint the user deliberately configured — and +/// said nothing, anywhere, about that endpoint specifically. +/// +/// Off is pinned on purpose, and so is the silence being gone. Two of these tests assert the +/// endpoint stays disabled on an unreadable file. That is not the defect; it is the decision. These two +/// keys are consent and address for a TCP listener, an unreadable file supplies neither, and starting on +/// guesses opens a port nobody asked for at an address no client is aimed at. The defect is that nothing +/// distinguished that from a user who never wanted MCP, which is why every other test here is about the +/// Problem string existing and naming what went wrong. +/// +/// Absent stays silent. A first run has no settings.json, off is correct, and there is no +/// endpoint to have lost. If that case ever starts producing a Problem, the reporting becomes noise on +/// every clean install and gets ignored exactly when it matters. +/// +public sealed class McpSettingsGuardTests +{ + private static string NewTempDir(string tag) + { + var dir = Path.Combine(Path.GetTempPath(), $"pmlite_mcp_{tag}_{Guid.NewGuid():N}"); + Directory.CreateDirectory(dir); + return dir; + } + + private static void WriteSettings(string dir, string content) => + File.WriteAllText(Path.Combine(dir, "settings.json"), content); + + private static void InTempConfig(string tag, string? content, Action check) + { + var dir = NewTempDir(tag); + try + { + if (content is not null) + { + WriteSettings(dir, content); + } + + check(McpSettings.Load(dir)); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The legitimate first run: no file, no endpoint, nothing to say. The control that stops every other + /// test here from being satisfied by a Load that simply complains about everything. + /// + [Fact] + public void Load_IsSilentlyOff_WhenThereIsNoFile() + { + InTempConfig("absent", null, settings => + { + Assert.False(settings.Enabled); + Assert.Equal(McpSettings.DefaultPort, settings.Port); + Assert.Null(settings.Problem); + Assert.False(settings.DisabledByUnreadableSettings); + }); + } + + /// + /// The other control: an ordinary file still configures the endpoint, and still says nothing. Without + /// this one, a Load that returned a Problem unconditionally would pass all the reporting tests. + /// + [Fact] + public void Load_ReadsBothKeys_AndStaysSilent_ForAnOrdinaryFile() + { + InTempConfig("ordinary", @"{ + ""alerts_enabled"": true, + ""mcp_enabled"": true, + ""mcp_port"": 5199 +}", settings => + { + Assert.True(settings.Enabled); + Assert.Equal(5199, settings.Port); + Assert.Null(settings.Problem); + }); + } + + /// + /// The reported case. A hand-edited file with one trailing comma — the shape #2418's sample makes + /// likelier, not rarer — used to return defaults and no signal at all. + /// + [Fact] + public void Load_SaysWhyTheEndpointIsOff_WhenTheFileIsUnparseable() + { + InTempConfig("comma", @"{ + ""mcp_enabled"": true, + ""mcp_port"": 5199, +}", settings => + { + Assert.True(settings.DisabledByUnreadableSettings); + Assert.NotNull(settings.Problem); + }); + } + + /// + /// The Problem has to be worth reading, not just non-null. "settings.json is broken" sends someone + /// through a file they believe is correct; a line and a position is a minute's work, and it is the + /// reason this reuses #2425's guard rather than growing a second read path with its own message. + /// + [Fact] + public void Load_NamesTheLineAndPosition_SoTheCommaCanBeFound() + { + InTempConfig("position", @"{ + ""mcp_enabled"": true, + ""mcp_port"": 5199, +}", settings => + { + Assert.NotNull(settings.Problem); + Assert.Contains("line", settings.Problem, StringComparison.Ordinal); + Assert.Contains("position", settings.Problem, StringComparison.Ordinal); + }); + } + + /// + /// The JSON literal null parses, so it never reached the old catch as a parse failure — the + /// first TryGetProperty threw on it instead, and landed in the same silent fallback. #2425 routes it to + /// Unreadable because it is the shape that used to make a save replace the whole document. + /// + [Fact] + public void Load_SaysWhyTheEndpointIsOff_WhenTheRootIsTheJsonLiteralNull() + { + InTempConfig("nullroot", "null", settings => + { + Assert.False(settings.Enabled); + Assert.True(settings.DisabledByUnreadableSettings); + }); + } + + /// + /// A quoted boolean is the other hand-edit that turns the endpoint off in silence, and the document + /// parses fine, so the file guard alone does not catch it. The old code reached GetBoolean, threw, and + /// fell into the same bare catch. Naming the key is the whole value: nothing else in Lite can tell + /// someone that mcp_enabled in particular is the reason a connection is being refused. + /// + [Fact] + public void Load_NamesTheKey_WhenMcpEnabledHoldsTheWrongKindOfValue() + { + InTempConfig("quotedbool", @"{ ""mcp_enabled"": ""true"" }", settings => + { + Assert.False(settings.Enabled); + Assert.True(settings.DisabledByUnreadableSettings); + Assert.Contains("mcp_enabled", settings.Problem!, StringComparison.Ordinal); + }); + } + + /// + /// Same for the port, and the case matters more than it looks: the endpoint was explicitly enabled + /// here, and the only thing wrong is the address. Starting anyway on the 5151 fallback would bind a + /// port the operator's firewall rule (#2414) does not follow and no client is aimed at, so this stays + /// off — loudly. + /// + [Fact] + public void Load_StaysOffAndNamesTheKey_WhenMcpPortHoldsTheWrongKindOfValue() + { + InTempConfig("quotedport", @"{ ""mcp_enabled"": true, ""mcp_port"": ""5199"" }", settings => + { + Assert.False(settings.Enabled); + Assert.True(settings.DisabledByUnreadableSettings); + Assert.Contains("mcp_port", settings.Problem!, StringComparison.Ordinal); + }); + } + + /// + /// The decision, pinned so a later "be helpful and start it anyway" cannot land quietly: a file that + /// says the endpoint is enabled but cannot be read leaves it OFF. Consent that cannot be parsed is not + /// consent, and this is the assertion that would go red if someone read the enabled flag out of a + /// half-parsed document. + /// + [Fact] + public void Load_FailsClosed_WhenAFileThatEnablesTheEndpointCannotBeParsed() + { + InTempConfig("failclosed", @"{ + ""mcp_enabled"": true, + ""mcp_port"": 5199 + ""alerts_enabled"": true +}", settings => + { + Assert.False(settings.Enabled); + Assert.True(settings.DisabledByUnreadableSettings); + }); + } +} + +/// +/// The wiring half of #2431. McpSettings.Load can now say the endpoint was lost, but a signal +/// nobody reads is the same silence with more code in it — and the three readers all live in WPF windows +/// that no test in this suite can instantiate, so the source is the only place the wiring is visible. +/// +/// These are shape pins, not text pins: they require that the branch exists and reports, not that +/// it reports any particular wording. +/// +public sealed class McpEndpointLossReportingTests +{ + /// + /// The load path must have no arm that answers a failure with defaults and nothing else. This is the + /// literal shape that shipped — catch { return new McpSettings(); } — and banning it stops the + /// fix being undone by someone tidying an unreachable-looking arm back into a bare catch. + /// + [Fact] + public void McpSettingsLoad_HasNoCatchThatReturnsDefaults() + { + var source = File.ReadAllText(FindRepoFile(Path.Combine("Lite", "Mcp", "McpSettings.cs"))); + + Assert.DoesNotMatch( + new Regex(@"catch\s*(\([^)]*\))?\s*\{\s*return new McpSettings\(\);"), + source); + } + + /// + /// The one moment Lite ever knows the endpoint is gone. Before this, the loss took the same + /// if (!mcpSettings.Enabled) return; as a user who never wanted MCP, and the symptom surfaced + /// as a refused connection on a different machine against an app reporting itself healthy. + /// + [Fact] + public void StartMcpServer_ReportsTheLostEndpoint_BeforeItGivesUpQuietly() + { + var body = MethodBody( + File.ReadAllText(FindRepoFile(Path.Combine("Lite", "MainWindow.xaml.cs"))), + "private async Task StartMcpServerAsync()"); + + var lossBranch = body.IndexOf("DisabledByUnreadableSettings", StringComparison.Ordinal); + Assert.True(lossBranch >= 0, + "StartMcpServerAsync never asks whether the endpoint is off because settings.json could not be " + + "read, so an unreadable file is still indistinguishable from a user who left MCP off."); + + var quietReturn = body.IndexOf("if (!mcpSettings.Enabled) return;", StringComparison.Ordinal); + Assert.True(quietReturn >= 0, + "The quiet not-enabled return moved; this guard's anchor is stale and it is pinning nothing."); + Assert.True(lossBranch < quietReturn, + "The unreadable-settings branch has to come BEFORE the quiet return, or the loss is swallowed " + + "by the branch that legitimately says nothing."); + + Assert.Contains("AppLogger.Error", body.Substring(lossBranch, quietReturn - lossBranch), + StringComparison.Ordinal); + } + + /// + /// #2425's startup report is what the person at the keyboard actually sees, and "every setting is at + /// its default" does not tell them an endpoint on another machine just stopped answering. The report + /// has to name the capability, not the settings. + /// + [Fact] + public void TheUnreadableSettingsReport_NamesTheMcpEndpoint() + { + var source = File.ReadAllText(FindRepoFile(Path.Combine("Lite", "App.xaml.cs"))); + + Assert.Contains("MCP", MethodBody(source, "private static void ReportUnreadableSettings(string? problem)"), + StringComparison.Ordinal); + Assert.Contains("MCP", MethodBody(source, "private static void ReportUnreadableSettingsToUser()"), + StringComparison.Ordinal); + } + + /// + /// The Settings window is the one place in the UI that claims to show the endpoint's configuration. + /// On an unreadable file it shows an unticked box and port 5151, which is a fallback and not a reading + /// of anything, so the status line beside them must not call that "Disabled". + /// + [Fact] + public void TheSettingsWindow_DoesNotCallAnUnreadableEndpointDisabled() + { + var body = MethodBody( + File.ReadAllText(FindRepoFile(Path.Combine("Lite", "Windows", "SettingsWindow.xaml.cs"))), + "private void UpdateMcpStatus()"); + + Assert.Contains("_mcpSettingsProblem", body, StringComparison.Ordinal); + + var problemBranch = body.IndexOf("_mcpSettingsProblem != null", StringComparison.Ordinal); + var disabled = body.IndexOf("\"Status: Disabled\"", StringComparison.Ordinal); + Assert.True(problemBranch >= 0, + "UpdateMcpStatus never checks whether settings.json could be read, so an endpoint lost to a " + + "parse error is reported with the same words as one nobody ever turned on."); + Assert.True(disabled < 0 || problemBranch < disabled, + "The unreadable branch has to be reached before the \"Disabled\" wording, or the window agrees " + + "with the file that the endpoint was never wanted."); + } + + /// + /// A warning that cannot go away is its own defect. Every Save in this window copies an unreadable + /// settings.json aside and writes a fresh one, and the window stays open afterwards — so a status + /// line set once at construction would keep telling the user the file cannot be read over a file + /// they have just fixed, for the rest of the session. Found in review on the first round of #2431. + /// + [Fact] + public void TheSettingsWindow_ReconsidersTheEndpointStateAfterASave() + { + var body = MethodBody( + File.ReadAllText(FindRepoFile(Path.Combine("Lite", "Windows", "SettingsWindow.xaml.cs"))), + "private async void SaveButton_Click("); + + Assert.Contains("_mcpSettingsProblem", body, StringComparison.Ordinal); + Assert.Contains("UpdateMcpStatus()", body, StringComparison.Ordinal); + } + + /// + /// Everything between a method's signature and the first line that closes it at method indentation. + /// Crude on purpose: it is enough to keep an assertion from being satisfied by a match somewhere else + /// in a two-thousand-line window class, which is the only thing that would make these pins vacuous. + /// + private static string MethodBody(string source, string signature) + { + var start = source.IndexOf(signature, StringComparison.Ordinal); + Assert.True(start >= 0, $"'{signature}' was not found; this guard's anchor is stale."); + + var end = source.IndexOf("\n }", start, StringComparison.Ordinal); + Assert.True(end > start, $"No close found for '{signature}'."); + + return source.Substring(start, end - start); + } + + private static string FindRepoFile(string relativePath) + { + var dir = AppContext.BaseDirectory; + for (var i = 0; i < 8 && dir is not null; i++) + { + var candidate = Path.Combine(dir, relativePath); + if (File.Exists(candidate)) + { + return candidate; + } + dir = Path.GetDirectoryName(dir); + } + + throw new FileNotFoundException($"Could not locate {relativePath} walking up from {AppContext.BaseDirectory}"); + } +} diff --git a/Lite.Tests/McpStatusEnvelopeTests.cs b/Lite.Tests/McpStatusEnvelopeTests.cs index b3349cc9e..d9f373030 100644 --- a/Lite.Tests/McpStatusEnvelopeTests.cs +++ b/Lite.Tests/McpStatusEnvelopeTests.cs @@ -206,4 +206,112 @@ public async Task GetDeadlocks_NoDeadlocks_ReturnsEmpty() Assert.Equal("empty", root.GetProperty("status").GetString()); } + + /// + /// #2546, Lite's side, through the real tool: the deadlock table is empty in BOTH states, and the only + /// thing that differs is what the last running_jobs run recorded. That single difference has to + /// turn "no running SQL Agent jobs found" — an affirmative claim about the server's Agent — into "we + /// were denied msdb, here is the grant". + /// + [Fact] + public async Task GetRunningJobs_AfterADeniedRun_ReturnsPreconditionNamingTheDenial() + { + await SeedCollectionLogAsync("running_jobs", "PERMISSIONS", + "The server principal is not able to access the database \"msdb\" under the current security context."); + + var root = Parse(await McpJobTools.GetRunningJobs(_dataService, _serverManager)); + + Assert.Equal("precondition", root.GetProperty("status").GetString()); + var message = root.GetProperty("message").GetString()!; + Assert.Contains("msdb", message, StringComparison.Ordinal); + Assert.Contains("the SQL Agent running-job snapshot", message, StringComparison.Ordinal); + Assert.Contains("re-derives it on EVERY call", message, StringComparison.Ordinal); + } + + /// + /// The other direction, and the one that keeps the branch from becoming a blanket rule: a collector that + /// ran and found nothing keeps the read's own empty. Without this, a precondition answer that had + /// stopped distinguishing anything would still pass the assertion above. + /// + [Fact] + public async Task GetRunningJobs_AfterASuccessfulRun_KeepsItsEmptyMiss() + { + await SeedCollectionLogAsync("running_jobs", "SUCCESS", null); + + var root = Parse(await McpJobTools.GetRunningJobs(_dataService, _serverManager)); + + Assert.Equal("empty", root.GetProperty("status").GetString()); + } + + /// + /// The Query Store half, whose evidence is a collected SNAPSHOT rather than a log status. Both SKUs used + /// to say "Query Store may not be enabled on target databases" — a guess, and equally true of a server + /// where it IS enabled. The hourly health collector has recorded the answer all along. + /// + [Fact] + public async Task GetQueryStoreTop_WithQueryStoreOff_StatesItFromTheSnapshot() + { + await SeedQueryStoreHealthAsync("OffDb", "OFF"); + + var root = Parse(await McpQueryTools.GetQueryStoreTop(_dataService, _serverManager, database_name: "OffDb")); + + Assert.Equal("precondition", root.GetProperty("status").GetString()); + var message = root.GetProperty("message").GetString()!; + Assert.Contains("OffDb", message, StringComparison.Ordinal); + Assert.Contains("SET QUERY_STORE = ON", message, StringComparison.Ordinal); + } + + /// + /// Scope is part of the evidence: a read narrowed to a database whose Query Store IS collecting must keep + /// its own miss, or the branch would claim a precondition about every empty window on the server. + /// + [Fact] + public async Task GetQueryStoreTop_WithQueryStoreOn_KeepsItsUnavailableMiss() + { + await SeedQueryStoreHealthAsync("OnDb", "READ_WRITE"); + + var root = Parse(await McpQueryTools.GetQueryStoreTop(_dataService, _serverManager, database_name: "OnDb")); + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + } + + private async Task SeedCollectionLogAsync(string collectorName, string status, string? errorMessage) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + + using var cmd = conn.CreateCommand(); + cmd.CommandText = @"INSERT INTO collection_log + (log_id, server_id, server_name, collector_name, collection_time, duration_ms, status, + error_message, rows_collected, sql_duration_ms, duckdb_duration_ms) + VALUES ($1, $2, 'TestServer', $3, $4, 0, $5, $6, 0, 0, 0)"; + void P(object? v) => cmd.Parameters.Add(new DuckDBParameter { Value = v ?? DBNull.Value }); + P(_nextId--); + P(_serverId); + P(collectorName); + P(DateTime.UtcNow.AddMinutes(-1)); + P(status); + P(errorMessage); + await cmd.ExecuteNonQueryAsync(); + } + + private async Task SeedQueryStoreHealthAsync(string databaseName, string actualState) + { + using var readLock = _duckDb.AcquireReadLock(); + var conn = await SeedConnectionAsync(); + + using var cmd = conn.CreateCommand(); + cmd.CommandText = @"INSERT INTO query_store_health + (config_id, capture_time, server_id, server_name, database_name, actual_state, desired_state, + readonly_reason, current_storage_size_mb, max_storage_size_mb, size_based_cleanup_mode, + stale_query_threshold_days, max_plans_per_query, interval_length_minutes) + VALUES ($1, $2, $3, 'TestServer', $4, $5, 'READ_WRITE', 0, 0, 1000, 'AUTO', 30, 200, 60)"; + void P(object v) => cmd.Parameters.Add(new DuckDBParameter { Value = v }); + P(_nextId--); + P(DateTime.UtcNow.AddMinutes(-5)); + P(_serverId); + P(databaseName); + P(actualState); + await cmd.ExecuteNonQueryAsync(); + } } diff --git a/Lite.Tests/PartialDatabaseFailureNoteTests.cs b/Lite.Tests/PartialDatabaseFailureNoteTests.cs new file mode 100644 index 000000000..93b763ac1 --- /dev/null +++ b/Lite.Tests/PartialDatabaseFailureNoteTests.cs @@ -0,0 +1,180 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2623: a per-database collector that fails in SOME databases and succeeds in the rest must say so. +/// +/// +/// Both runners tolerate a per-database failure by design — one offline database must not cost the other +/// twenty-nine — and both escalate only when EVERY database failed. In between sat a hole with no +/// evidence in it at all: the cycle recorded SUCCESS, whatever the survivors produced, and a note +/// composed solely from probe failures, which a thrown exception is not. +/// +/// +/// +/// That hole is how #2622 stayed alive. Three collectors were writing one fewer payload value than they +/// declared, so every row they produced was rejected. All three are per-database. All three failed +/// identically in the one database that had data. Only pg_extension_availability surfaced as an +/// ERROR, and only because it returns rows in every database including postgres, so all-failed +/// tripped the escalation. The other two had a second database with legitimately nothing to report, it +/// succeeded, and the cycle logged SUCCESS with zero rows — indistinguishable from a target that has no +/// large tables, which is what I assumed it was. +/// +/// +/// +/// So the assertions here are about the ABSENCE being explained, not about the failure being prevented. +/// Skipping the database is still correct. Skipping it quietly is what turned a fixable bug into three +/// schema versions of a plausible-looking empty table. +/// +/// +public class PartialDatabaseFailureNoteTests +{ + [Fact] + public void APartialLossComposesANoteNamingWhatWasSkipped() + { + var note = EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: 1, + attempted: 2, + failedDatabases: new[] { "appdb" }, + firstError: "Collector wrote 9 payload values but declares 10 payload columns"); + + Assert.NotNull(note); + Assert.Contains("1 of 2", note, StringComparison.Ordinal); + Assert.Contains("appdb", note, StringComparison.Ordinal); + Assert.Contains("declares 10 payload columns", note, StringComparison.Ordinal); + } + + /// + /// The note's whole job. A row count is the first thing an operator reads off this cycle, and on a + /// partial loss it is a number about the survivors wearing the shape of a number about the server. + /// + [Fact] + public void TheNoteWarnsThatTheRowCountIsAboutTheSurvivorsOnly() + { + var note = EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: 1, attempted: 2, failedDatabases: new[] { "appdb" }, firstError: "boom"); + + Assert.NotNull(note); + Assert.Contains("survivors ONLY", note, StringComparison.Ordinal); + } + + [Fact] + public void NothingFailedComposesNoNote() + => Assert.Null(EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: 0, attempted: 5, failedDatabases: Array.Empty(), firstError: null)); + + /// + /// EVERYTHING failing is not this note's case: that path rethrows the first failure so the run is + /// classified (PERMISSIONS / SESSION_MISSING / ERROR), and a note on a row about to carry an error + /// message would only compete with it. + /// + [Fact] + public void EverythingFailingComposesNoNoteBecauseItRethrowsInstead() + => Assert.Null(EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: 4, attempted: 4, failedDatabases: new[] { "a", "b", "c", "d" }, firstError: "boom")); + + /// + /// One login problem can fail every database on a busy server. The note column is a one-line summary, + /// so the names are capped and the remainder counted — the opposite trade from the probe-failure note, + /// which carries no names at all, because there the names are the same problem repeated. + /// + [Fact] + public void ManyFailedDatabasesAreCappedAndCounted() + { + var many = Enumerable.Range(1, 30).Select(i => $"db{i}").ToArray(); + + var note = EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: many.Length, attempted: many.Length + 1, failedDatabases: many, firstError: "boom"); + + Assert.NotNull(note); + Assert.Contains("db1", note, StringComparison.Ordinal); + Assert.Contains($"and {many.Length - EnumeratedCollectorDriver.MaxNamedFailedDatabases} more", note, StringComparison.Ordinal); + Assert.DoesNotContain("db30", note, StringComparison.Ordinal); + } + + /// + /// A cycle can lose databases to thrown exceptions AND report probe failures. Two independent sources, + /// one note column: whichever host wrote its own join would be the one that silently dropped one. + /// + [Fact] + public void BothNoteSourcesSurviveTheMerge() + { + var probe = EnumeratedCollectorDriver.BuildNote(enumerationWasEmpty: false, probeFailureCount: 3); + var partial = EnumeratedCollectorDriver.BuildPartialFailureNote( + failed: 1, attempted: 4, failedDatabases: new[] { "appdb" }, firstError: "boom"); + + var merged = EnumeratedCollectorDriver.MergeNotes(probe, partial); + + Assert.NotNull(merged); + Assert.Contains("enumeration probe", merged, StringComparison.Ordinal); + Assert.Contains("appdb", merged, StringComparison.Ordinal); + } + + [Theory] + [InlineData(null, null, null)] + [InlineData("a", null, "a")] + [InlineData(null, "b", "b")] + [InlineData("a", "b", "a; b")] + public void MergeNotesHandlesEitherSideMissing(string? first, string? second, string? expected) + => Assert.Equal(expected, EnumeratedCollectorDriver.MergeNotes(first, second)); + + /// + /// The source pin. Both runners compose this note in their per-database loop, and the two have already + /// drifted apart on this exact loop once — Darling carried its own copy of the fallback predicate, + /// which is how #857's fix missed it. Anchored on the CALL, not on a message string. + /// + [Theory] + [InlineData("Darling/PerformanceMonitor.Darling.Service/DarlingCollectorRunner.cs")] + [InlineData("Lite/Services/RemoteCollectorService.DefinitionRunner.cs")] + public void BothRunnersComposeThePartialFailureNote(string relativePath) + { + var source = File.ReadAllText(Path.Combine(RepoRoot(), relativePath)); + + Assert.Contains("EnumeratedCollectorDriver.BuildPartialFailureNote(", source, StringComparison.Ordinal); + Assert.Contains("EnumeratedCollectorDriver.MergeNotes(", source, StringComparison.Ordinal); + + /* Every arm that counts a failure must also NAME it, or the note reports a count with the wrong + set of names beside it — worse than no names, because it reads as complete. */ + var counted = CountOccurrences(source, "failed++;"); + var named = CountOccurrences(source, "failedDatabases.Add(databaseName);"); + Assert.Equal(counted, named); + } + + private static int CountOccurrences(string haystack, string needle) + { + var count = 0; + var index = 0; + while ((index = haystack.IndexOf(needle, index, StringComparison.Ordinal)) >= 0) + { + count++; + index += needle.Length; + } + + return count; + } + + private static string RepoRoot() + { + var directory = new DirectoryInfo(AppContext.BaseDirectory); + while (directory is not null && !File.Exists(Path.Combine(directory.FullName, "PerformanceMonitor.sln"))) + { + directory = directory.Parent; + } + + return directory?.FullName ?? throw new InvalidOperationException("PerformanceMonitor.sln not found above the test output directory."); + } +} diff --git a/Lite.Tests/PerDatabaseCollectorAttributionTests.cs b/Lite.Tests/PerDatabaseCollectorAttributionTests.cs new file mode 100644 index 000000000..5f0621168 --- /dev/null +++ b/Lite.Tests/PerDatabaseCollectorAttributionTests.cs @@ -0,0 +1,89 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System.Collections.Generic; +using System.Linq; +using System.Reflection; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the one invariant that ties a per-database collector's rows back to the database they describe: +/// if a definition ever returns true from , its payload +/// must carry a database_name column. +/// +/// Why this is a build-time assertion and not a review habit. A per-database collector without +/// that column still compiles, still collects, and still looks correct in a single-database test rig. It +/// fails only on a cluster that has the same schema in two databases — the ordinary multi-tenant shape — +/// where the rows collide on (server_id, schema, table, column) and nothing distinguishes them. Two +/// collectors shipped in that state (#2599): pg_column_stats, which declared +/// RunsPerDatabase => true and no database_name, and pg_extension_availability, +/// which read the per-database pg_extension catalog from a single connection. +/// +/// The targets below are representative rather than exhaustive on purpose: the property takes a +/// target because a few SQL Server definitions only go per-database on Azure SQL DB, so a single target +/// would silently skip them. Any target for which the answer is true triggers the requirement. +/// +public class PerDatabaseCollectorAttributionTests +{ + private static readonly CollectorTargetInfo[] s_targets = + { + new() { Engine = CollectorTargetEngine.SqlServer, SqlMajorVersion = 16 }, + new() { Engine = CollectorTargetEngine.SqlServer, SqlMajorVersion = 12, IsAzureSqlDb = true }, + new() { Engine = CollectorTargetEngine.SqlServer, SqlMajorVersion = 15, IsAzureManagedInstance = true }, + new() { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17, PostgresVersionNum = 170000 }, + new() { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 14, PostgresVersionNum = 140000, IsAurora = true }, + }; + + [Fact] + public void EveryPerDatabaseCollectorDeclaresDatabaseName() + { + var offenders = new List(); + + foreach (var schema in CollectorCatalog.All) + { + /* RunsPerDatabase is declared on the generic ICollectorDefinition, and the catalog is + typed as the row-agnostic ICollectorSchemaInfo, so there is no non-generic surface to call + it through. Reflection over the concrete type is the only route that stays generic across + every row shape; a definition that somehow lacks the method is not per-database. */ + var method = schema.GetType().GetMethod( + "RunsPerDatabase", + BindingFlags.Public | BindingFlags.Instance, + binder: null, + types: new[] { typeof(CollectorTargetInfo) }, + modifiers: null); + + if (method is null) + { + continue; + } + + var runsPerDatabase = s_targets.Any(target => (bool)method.Invoke(schema, new object[] { target })!); + + if (!runsPerDatabase) + { + continue; + } + + var hasDatabaseName = schema.PayloadColumns + .Any(column => column.Name == "database_name"); + + if (!hasDatabaseName) + { + offenders.Add(schema.Name); + } + } + + Assert.True( + offenders.Count == 0, + "These collectors run per database but do not declare a database_name payload column, so their " + + "rows cannot be attributed to the database they came from: " + string.Join(", ", offenders)); + } +} diff --git a/Lite.Tests/PerformanceTrendsToolTests.cs b/Lite.Tests/PerformanceTrendsToolTests.cs new file mode 100644 index 000000000..512c82e84 --- /dev/null +++ b/Lite.Tests/PerformanceTrendsToolTests.cs @@ -0,0 +1,230 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's Performance-Trends siblings (#2484) - get_procedure_duration_trend and +/// get_query_store_duration_trend - and the two-branch empty answer all three siblings now share. +/// +/// Written at the TOOL level: the empty branches and the payload field names live in the tool, and a +/// reader-level test would see neither. The SKUs promise each other the same sentences for the same state, +/// which is a promise only a test can hold. +/// +public sealed class PerformanceTrendsToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "PerfTrendsSrv"; + private const string Db = "AppDb"; + + /* Lite DERIVES its server id from the storage name; seeding under a hardcoded one would write rows the + tool looks straight past and pass the never-sampled assertion for the wrong reason. */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public PerformanceTrendsToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-perftrends-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task AQuietWindow_AndNeverSampled_AreDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + /* Registered, never sampled: not a quiet server, and it must not be described as one. */ + var never = await McpQueryTools.GetProcedureDurationTrend(service, _serverManager, ServerName, 24); + var neverRoot = JsonDocument.Parse(never).RootElement; + Assert.Equal("unavailable", neverRoot.GetProperty("status").GetString()); + var neverText = neverRoot.GetProperty("message").GetString()!; + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + Assert.Contains("NOT a quiet server", neverText, StringComparison.Ordinal); + Assert.DoesNotContain("widen", neverText, StringComparison.OrdinalIgnoreCase); + + /* Sampled, but outside the asked-for window: widening IS the move. */ + await SeedProcedureAsync(Truncate(DateTime.UtcNow.AddHours(-48)), executions: 5, elapsedUs: 1_000_000); + + var quiet = await McpQueryTools.GetProcedureDurationTrend(service, _serverManager, ServerName, 1); + var quietRoot = JsonDocument.Parse(quiet).RootElement; + Assert.Equal("empty", quietRoot.GetProperty("status").GetString()); + var quietText = quietRoot.GetProperty("message").GetString()!; + Assert.Contains("widen", quietText, StringComparison.OrdinalIgnoreCase); + Assert.DoesNotContain("EVER", quietText, StringComparison.Ordinal); + } + + [Fact] + public async Task QueryStoreEmpty_NamesTheCauseThePlanCacheTrendsDoNotHave() + { + var service = new LocalDataService(_duckDb); + + var never = await McpQueryTools.GetQueryStoreDurationTrend(service, _serverManager, ServerName, 24); + var root = JsonDocument.Parse(never).RootElement; + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + + /* + Query Store can simply be OFF on every database, which no amount of collector health fixes. A + message that led with the collector would send the reader to the wrong place. + */ + Assert.Contains("Query Store may be OFF", root.GetProperty("message").GetString()!, StringComparison.Ordinal); + } + + [Fact] + public async Task TheExecutionRate_SurvivesBeingBelowOnePerSecond() + { + var service = new LocalDataService(_duckDb); + + /* Two snapshots five minutes apart, two executions between them: 0.0067/sec. */ + var baseNow = Truncate(DateTime.UtcNow); + await SeedProcedureAsync(baseNow.AddMinutes(-20), executions: 0, elapsedUs: 0); + await SeedProcedureAsync(baseNow.AddMinutes(-15), executions: 2, elapsedUs: 600_000); + + var hit = await McpQueryTools.GetProcedureDurationTrend(service, _serverManager, ServerName, 4); + var trend = JsonDocument.Parse(hit).RootElement.GetProperty("trend"); + Assert.Equal(2, trend.GetArrayLength()); + + var second = trend[1]; + Assert.True(second.GetProperty("value").GetDouble() > 0, "elapsed ms/sec must be a real rate"); + + /* The shipped integer field rounds this to an idle server. The double is why it is here. */ + Assert.Equal(0, second.GetProperty("execution_count").GetInt64()); + Assert.True( + second.GetProperty("executions_per_second").GetDouble() > 0, + "executions_per_second must survive a rate below 1/sec that execution_count truncates to zero"); + } + + [Fact] + public async Task EachQueryStoreInterval_IsCountedOnce_AtTheHourTheWorkRan() + { + var service = new LocalDataService(_duckDb); + + /* + Two runtime intervals, each fetched twice while open, the second fetch carrying the higher + cumulative count. Charging every fetch to its collection time would give four points and double + the work; dedup + placement gives two, at the interval starts. + */ + var baseNow = Truncate(DateTime.UtcNow); + var intervalA = baseNow.AddHours(-3); + var intervalB = baseNow.AddHours(-2); + await SeedQueryStoreAsync(baseNow.AddMinutes(-150), 41, intervalA, executions: 10, avgDurationUs: 2000); + await SeedQueryStoreAsync(baseNow.AddMinutes(-145), 41, intervalA, executions: 40, avgDurationUs: 2000); + await SeedQueryStoreAsync(baseNow.AddMinutes(-90), 42, intervalB, executions: 5, avgDurationUs: 3000); + await SeedQueryStoreAsync(baseNow.AddMinutes(-85), 42, intervalB, executions: 25, avgDurationUs: 3000); + + var hit = await McpQueryTools.GetQueryStoreDurationTrend(service, _serverManager, ServerName, 6); + var trend = JsonDocument.Parse(hit).RootElement.GetProperty("trend"); + + Assert.Equal(2, trend.GetArrayLength()); + + /* + The surviving snapshot is the FINAL one (25 executions over the 3600 seconds between the two + interval starts), not the first (5) and not their sum (30). A dedup keeping the wrong row would + still return two points and would still look right. + */ + Assert.Equal(25d / 3600d, trend[1].GetProperty("executions_per_second").GetDouble(), 6); + } + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedProcedureAsync(DateTime collectionTimeUtc, long executions, long elapsedUs) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO procedure_stats + (collection_id, collection_time, server_id, server_name, + database_name, schema_name, object_name, + delta_execution_count, delta_worker_time, delta_elapsed_time, delta_logical_reads) +VALUES ($1, $2, $3, $4, $5, 'dbo', $6, $7, $8, $9, $10)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = collectionTimeUtc }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = Db }); + cmd.Parameters.Add(new DuckDBParameter { Value = "usp_Trend" }); + cmd.Parameters.Add(new DuckDBParameter { Value = executions }); + cmd.Parameters.Add(new DuckDBParameter { Value = elapsedUs / 2 }); + cmd.Parameters.Add(new DuckDBParameter { Value = elapsedUs }); + cmd.Parameters.Add(new DuckDBParameter { Value = 0L }); + await cmd.ExecuteNonQueryAsync(); + } + + private async Task SeedQueryStoreAsync( + DateTime collectionTime, long intervalId, DateTime intervalStart, long executions, long avgDurationUs) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO query_store_stats + (collection_id, collection_time, server_id, server_name, database_name, + query_id, plan_id, execution_type_desc, execution_count, avg_duration_us, + runtime_stats_interval_id, interval_start_time_utc) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = collectionTime }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = Db }); + cmd.Parameters.Add(new DuckDBParameter { Value = 7L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 9L }); + cmd.Parameters.Add(new DuckDBParameter { Value = "Regular" }); + cmd.Parameters.Add(new DuckDBParameter { Value = executions }); + cmd.Parameters.Add(new DuckDBParameter { Value = avgDurationUs }); + cmd.Parameters.Add(new DuckDBParameter { Value = intervalId }); + cmd.Parameters.Add(new DuckDBParameter { Value = intervalStart }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/PgBufferUsageCollectorDefinitionTests.cs b/Lite.Tests/PgBufferUsageCollectorDefinitionTests.cs new file mode 100644 index 000000000..1307170ec --- /dev/null +++ b/Lite.Tests/PgBufferUsageCollectorDefinitionTests.cs @@ -0,0 +1,153 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2544, the buffers slice. Every assertion here pins a correctness trap that was measured rather than +/// reasoned about — the filenode join in particular, where the query every published example writes loses +/// any table that has ever been rewritten. +/// +public class PgBufferUsageCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgBufferUsageCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// The filenode join must go through pg_relation_filenode(). Measured: after one + /// VACUUM FULL, the naive pg_class.oid = relfilenode join reported ZERO buffers for a + /// table holding 6,667 — a table silently vanishes from the report once rewritten. Mapped catalogs + /// carry relfilenode = 0, so joining the raw column drops those too. + /// + [Fact] + public void TheFilenodeJoin_GoesThroughPgRelationFilenode() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + Assert.Contains("pg_catalog.pg_relation_filenode(c.oid)", sql, StringComparison.Ordinal); + + /* And never the two wrong keys. Comments are stripped first because the query explains exactly + these mistakes and would otherwise satisfy its own prohibition. */ + Assert.DoesNotMatch(new Regex(@"c\.oid\s*=\s*\S*relfilenode"), sql); + Assert.DoesNotMatch(new Regex(@"c\.relfilenode\s*="), sql); + } + + /// + /// The relation name may be resolved ONLY for buffers belonging to the connected database. The pool is + /// cluster-wide and pg_class is not, so a filenode from another database can collide with a local + /// OID and produce a confidently WRONG name — measured. + /// + [Fact] + public void TheRelationJoin_IsScopedToTheConnectedDatabase() + => Assert.Matches(new Regex(@"reldatabase\s*=\s*\(SELECT oid FROM pg_catalog\.pg_database WHERE datname = current_database\(\)\)"), Sql()); + + /// + /// Buffers from other databases are KEPT, not filtered out. Dropping them would understate how full the + /// pool is, which is the one number this collector exists to report. + /// + [Fact] + public void ForeignDatabaseBuffers_AreKeptRatherThanFiltered() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + /* The only WHERE on the outer query excludes UNUSED buffers (relfilenode IS NULL), never foreign + ones — a predicate on reldatabase in WHERE would silently drop them. */ + Assert.Matches(new Regex(@"WHERE\s+p\.relfilenode IS NOT NULL"), sql); + Assert.DoesNotMatch(new Regex(@"WHERE[^)]*p\.reldatabase\s*="), sql); + } + + /// + /// The relation join must be LEFT. An inner join would drop every buffer this instance cannot name, + /// which is precisely the foreign-database occupancy above. + /// + [Fact] + public void TheRelationJoin_IsLeft() + => Assert.Matches(new Regex(@"LEFT JOIN pg_catalog\.pg_class"), Sql()); + + /// + /// Pool totals ride on every row so a share never has to be computed against a second read of a pool + /// that has moved on since. + /// + [Fact] + public void PoolTotals_TravelOnEveryRow() + { + var names = PgBufferUsageCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("pool_buffers_total", names); + Assert.Contains("pool_buffers_used", names); + } + + /// + /// Both totals come from a window over the SAME scan. A second query, or + /// pg_buffercache_summary() alongside the view, would sample a moving pool twice and the + /// disagreement would land in the percentage a reader computes. + /// + [Fact] + public void PoolTotals_ComeFromTheSameScan() + { + var sql = Sql(); + + Assert.Equal(2, Regex.Matches(sql, @"OVER \(\)").Count); + Assert.DoesNotContain("pg_buffercache_summary", sql, StringComparison.Ordinal); + } + + [Fact] + public void AppliesTo_EveryPostgresTarget() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + Assert.True(PgBufferUsageCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + } + } + + /// One SELECT alias per payload column, in order — a mismatch is a silently shifted binary + /// COPY, which writes every value into the wrong column rather than failing. + [Fact] + public void SelectAliases_MatchThePayloadOrder() + { + var expected = PgBufferUsageCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + var lines = Sql().Split('\n'); + var outerSelect = Array.FindLastIndex(lines, l => l.TrimEnd() == "SELECT"); + + Assert.True(outerSelect >= 0, "the outer SELECT is no longer where this test looks for it"); + + var selected = lines + .Skip(outerSelect) + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal) + && !line.TrimStart().StartsWith("LEFT JOIN", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(expected, selected); + } +} diff --git a/Lite.Tests/PgColumnStatsCollectorDefinitionTests.cs b/Lite.Tests/PgColumnStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..8096b5ad6 --- /dev/null +++ b/Lite.Tests/PgColumnStatsCollectorDefinitionTests.cs @@ -0,0 +1,179 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2543: per-column planner statistics. The load-bearing assertion here is a NEGATIVE one — that the two +/// value-bearing columns are never collected — because those hold raw customer data and a later refactor +/// that "completed" the column set would be a data-handling regression rather than an improvement. +/// +public class PgColumnStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgColumnStatsCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// The value-bearing columns must never be selected. Measured on a realistic table, they return customer + /// names, identifier fragments and live email addresses — collecting them copies customer data into the + /// monitoring store under our retention, which is the same exposure as the auto_explain literals + /// on #2538. + /// + /// Asserted against the SELECT list specifically, so the comment explaining the exclusion cannot + /// satisfy it — the source-pin trap this repo has hit repeatedly. + /// + [Fact] + public void TheValueBearingColumns_AreNeverSelected() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + Assert.DoesNotContain("most_common_vals", sql.Replace("cardinality(s.most_common_vals)", " "), StringComparison.Ordinal); + Assert.DoesNotContain("histogram_bounds", sql, StringComparison.Ordinal); + } + + /// + /// most_common_vals may be touched ONLY through cardinality, which reads the array's + /// length and never its contents. That distinction is the whole reason a "how skewed" answer is + /// available without a "skewed toward what" answer. + /// + [Fact] + public void MostCommonVals_IsUsedOnlyForItsLength() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + foreach (Match match in Regex.Matches(sql, @"(\S*)most_common_vals")) + { + Assert.Equal("cardinality(s.", match.Groups[1].Value); + } + } + + /// + /// Only the HEAD of the frequency array is taken. The first element is the skew signal; storing the + /// whole array would be storing a distribution nobody reads past its head — and unlike the values array, + /// frequencies are safe, so the restraint here is about volume rather than exposure. + /// + [Fact] + public void OnlyTheTopFrequency_IsTaken() + => Assert.Matches(new Regex(@"most_common_freqs\[1\]"), Sql()); + + /// + /// n_distinct must not be stored as an integer. Negative values are a RATIO of row count, so an + /// integer column would let a read render "-1 distinct values" for a unique key — the commonest column + /// shape there is. + /// + [Fact] + public void NDistinct_IsNotAnIntegerColumn() + { + var column = PgColumnStatsCollector.Instance.PayloadColumns.Single(c => c.Name == "n_distinct"); + + Assert.Equal(CollectorColumnType.Double, column.Type); + } + + /// + /// A size floor, because statistics on a tiny table cannot produce a misestimate anyone notices — the + /// planner is choosing between two cheap paths — and this is the widest fan-out of any collector here + /// (columns x tables x databases). + /// + [Fact] + public void ThereIsASizeFloor_SoTheLongTailIsNotCollected() + => Assert.Matches(new Regex(@"c\.relpages\s*>=\s*\d+"), Sql()); + + /// + /// System schemas are excluded. Their statistics describe the catalog rather than the user's data, and + /// they would swamp the result on any database with few user tables. + /// + [Fact] + public void SystemSchemas_AreExcluded() + { + var sql = Sql(); + + Assert.Contains("pg_catalog", sql, StringComparison.Ordinal); + Assert.Contains("information_schema", sql, StringComparison.Ordinal); + } + + /// + /// Per-database, because pg_stats describes the CONNECTED database only. A cluster-wide claim + /// built from one database's statistics would be silently missing every table in every other one — the + /// same scope mistake that has now appeared three times in this effort. + /// + [Fact] + public void ItRunsPerDatabase() + => Assert.True(PgColumnStatsCollector.Instance.RunsPerDatabase( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + + [Fact] + public void AppliesTo_EveryPostgresTarget() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + Assert.True(PgColumnStatsCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + } + } + + /// + /// Catalog reads are schema-qualified: pg_catalog is searched implicitly but not necessarily + /// FIRST, so an unqualified read can resolve to an object a user created in a schema earlier in the + /// monitoring login's search_path. + /// + [Fact] + public void EveryCatalogRead_IsSchemaQualified() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + foreach (var view in new[] { "pg_stats", "pg_class", "pg_namespace" }) + { + /* Excludes the string-literal mentions of pg_catalog in the schema-exclusion predicate, which + are data rather than object references. */ + foreach (Match match in Regex.Matches(sql, $@"(\S*)\b{Regex.Escape(view)}\b")) + { + Assert.Equal("pg_catalog.", match.Groups[1].Value); + } + } + } + + /// One SELECT alias per payload column, in order — a mismatch is a silently shifted binary + /// COPY, which writes every value into the wrong column rather than failing. + [Fact] + public void SelectAliases_MatchThePayloadOrder() + { + var expected = PgColumnStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + var selected = Sql() + .Split('\n') + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal) + && !line.TrimStart().StartsWith("JOIN", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(expected, selected); + } +} diff --git a/Lite.Tests/PgDatabaseStatsCollectorDefinitionTests.cs b/Lite.Tests/PgDatabaseStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..e5eed1733 --- /dev/null +++ b/Lite.Tests/PgDatabaseStatsCollectorDefinitionTests.cs @@ -0,0 +1,334 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the pg_stat_database collector (#2539): the table name that must NOT shadow the catalog view, the +/// gate that deliberately gates on nothing, the UTC conversion, and the two shapes of NULL this view really +/// produces — the shared-relations row with no database name, and a stats_reset that is NULL until +/// the first reset ever. +/// +public class PgDatabaseStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext(int major = 17, ICollectorDeltaCalculator? deltas = null) + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc), + Deltas = deltas ?? s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + }; + + /// + /// The table name is the one thing here that cannot be changed later without a migration, and it must + /// not be pg_stat_database: pg_catalog is searched before search_path, so a store table by that + /// name makes every unqualified read resolve to the MONITORING store's own view. + /// + [Fact] + public void Identity_Pinned_AndTheTableDoesNotShadowTheCatalogView() + { + Assert.Equal("pg_database_stats", PgDatabaseStatsCollector.Instance.Name); + Assert.Equal("pg_database_stats", PgDatabaseStatsCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgDatabaseStatsCollector.Instance.TargetEngine); + + Assert.NotEqual("pg_stat_database", PgDatabaseStatsCollector.Instance.TargetTable); + } + + /// + /// Every PostgreSQL major, Aurora or not, writer or replica. Each of those is a decision: + /// + /// No version floor — the newest column selected (temp_files / temp_bytes / deadlocks) arrived in + /// PostgreSQL 9.2, so copying pg_stat_io's PG16 gate would have gated off a working view. + /// No Aurora gate — this is core PostgreSQL, and it is the ONLY temp-file evidence a stock + /// PostgreSQL target has, because pg_statement_stats reads aurora_stat_statements(). + /// No recovery gate — a sort on a read replica spills exactly the way it does on a writer, and a + /// replica is where the heavy reporting queries usually are. + /// + /// + [Theory] + [InlineData(13, false, false)] + [InlineData(15, false, true)] + [InlineData(16, true, false)] + [InlineData(17, true, true)] + [InlineData(0, false, false)] + public void AppliesToEveryPostgresTarget_IncludingStockAndReplicas(int major, bool isAurora, bool inRecovery) + { + Assert.True(PgDatabaseStatsCollector.Instance.AppliesTo(new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + IsAurora = isAurora, + IsInRecovery = inRecovery, + })); + } + + /// + /// The composed gate still refuses a SQL Server target: the collector's own AppliesTo says yes to + /// everything, so the ENGINE half is the only thing standing between this PostgreSQL query text and a + /// SQL Server connection. Asserting the collector's gate alone would not see that. + /// + [Fact] + public void TheEngineHalfOfTheDispatchGateStillRefusesASqlServerTarget() + { + Assert.False(CollectorCatalog.AppliesTo(PgDatabaseStatsCollector.Instance, new CollectorTargetInfo())); + Assert.True(CollectorCatalog.AppliesTo( + PgDatabaseStatsCollector.Instance, + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql })); + } + + /// + /// Cluster-wide, so no per-database fan-out — which is what makes a per-minute cadence affordable and + /// is the third leg of the per-database STORAGE argument (the grain is free, not paid for). + /// + [Fact] + public void DoesNotRunPerDatabase() + { + Assert.False(PgDatabaseStatsCollector.Instance.RunsPerDatabase(MakeContext().Target)); + } + + /// + /// The four questions, each present by name, plus the reset timestamp. A column dropped from the query + /// is a silent capability loss — the payload column would still exist and would simply always be NULL. + /// + [Theory] + [InlineData("d.temp_files")] + [InlineData("d.temp_bytes")] + [InlineData("d.blks_read")] + [InlineData("d.blks_hit")] + [InlineData("d.deadlocks")] + [InlineData("d.xact_commit")] + [InlineData("d.xact_rollback")] + public void SelectsEveryCounterTheFourQuestionsNeed(string column) + { + Assert.Contains(column, PgDatabaseStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + /// + /// timestamptz::text renders in the SESSION's TimeZone and the store contract is naive UTC, so + /// the conversion has to be explicit. Byte-identical to UTC on every instance in the fleet, which is + /// exactly why it survives every probe and has to be pinned instead. + /// + [Fact] + public void ConvertsStatsResetToUtcExplicitly() + { + var sql = PgDatabaseStatsCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("(d.stats_reset AT TIME ZONE 'UTC')", sql, StringComparison.Ordinal); + Assert.DoesNotContain("stats_reset::timestamp", sql, StringComparison.Ordinal); + } + + /// + /// The query text is version-invariant, deliberately: every column it names has existed since 9.2, so + /// there is nothing to branch on and a branch would be a source of drift rather than of correctness. + /// + [Fact] + public void TheQueryIsIdenticalOnEveryMajor() + { + var majors = new[] { 13, 14, 15, 16, 17, 18 } + .Select(m => PgDatabaseStatsCollector.Instance.BuildQuery(MakeContext(m)).Text) + .Distinct(StringComparer.Ordinal) + .ToArray(); + + Assert.Single(majors); + } + + /// + /// The columns PostgreSQL added AFTER 9.2 are absent, and that is what makes the gate-free AppliesTo + /// honest: selecting any of them would need a version floor, and a floor written without one would fail + /// the whole collection every cycle on an older major with "column does not exist". + /// + [Theory] + [InlineData("checksum_failures")] // PostgreSQL 12 + [InlineData("session_time")] // PostgreSQL 14 + [InlineData("sessions_killed")] // PostgreSQL 14 + [InlineData("parallel_workers_launched")] // PostgreSQL 18 + public void SelectsNoColumnThatWouldNeedAVersionFloor(string column) + { + Assert.DoesNotContain(column, PgDatabaseStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + [Fact] + public void PayloadColumns_OrderAndKeyTypes_Pinned() + { + var columns = PgDatabaseStatsCollector.Instance.PayloadColumns; + + Assert.Equal(9, columns.Count); + Assert.Equal( + new[] + { + "database_name", "xact_commit", "xact_rollback", "blks_read", "blks_hit", + "temp_files", "temp_bytes", "deadlocks", "stats_reset", + }, + columns.Select(c => c.Name).ToArray()); + + Assert.Equal(CollectorColumnType.Varchar, columns[0].Type); + Assert.Equal(CollectorColumnType.BigInt, columns[5].Type); // temp_files + Assert.Equal(CollectorColumnType.Timestamp, columns[8].Type); // stats_reset + } + + /// + /// A normal row: every counter populated, and a reset timestamp that has been set at some point. + /// + [Fact] + public async Task ReadsAFullyPopulatedRow() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "appdb", + 18_442_113L, 90_115L, // xact_commit, xact_rollback + 4_112_889L, 10_150_937_570L, // blks_read, blks_hit + 214L, 9_663_676_416L, // temp_files, temp_bytes + 3L, // deadlocks + new DateTime(2026, 5, 18, 7, 4, 22, DateTimeKind.Unspecified), + }); + + var rows = await PgDatabaseStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal("appdb", row.DatabaseName); + Assert.Equal(214L, row.TempFiles); + Assert.Equal(9_663_676_416L, row.TempBytes); + Assert.Equal(3L, row.Deadlocks); + Assert.Equal(90_115L, row.XactRollback); + Assert.Equal(new DateTime(2026, 5, 18, 7, 4, 22), row.StatsReset); + } + + /// + /// PostgreSQL's own shared-relations row: datname is NULL, and it must arrive as null rather + /// than being dropped or turned into a sentinel string. Its block counters are real, and the read keys + /// its series on that NULL. + /// + [Fact] + public async Task PreservesTheNullDatabaseNameOfTheSharedRelationsRow() + { + var reader = new FakeCollectorDataReader( + new object[] + { + DBNull.Value, + 0L, 0L, + 8_412L, 1_004_991L, + 0L, 0L, + 0L, + DBNull.Value, + }); + + var rows = await PgDatabaseStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Null(row.DatabaseName); + Assert.Equal(1_004_991L, row.BlksHit); + + /* NULL stats_reset is the ordinary state of a server nobody has ever reset, and it means exactly + "never reset" — not "unknown", and certainly not zero. */ + Assert.Null(row.StatsReset); + + /* A real zero stays a zero, so "no temp files" is distinguishable from "not reported". */ + Assert.Equal(0L, row.TempFiles); + } + + [Fact] + public async Task ReturnsNoRowsWhenTheViewReturnsNone() + { + var rows = await PgDatabaseStatsCollector.Instance.ReadAsync( + new FakeCollectorDataReader(), MakeContext(), CancellationToken.None); + + Assert.Empty(rows); + } + + /// + /// No stored deltas: the raw cumulative counter is stored and the window is differenced at read time, + /// so a reset stays visible in the data rather than being smoothed away at write time by a delta the + /// reader cannot re-derive. + /// + [Fact] + public void TakesNoDeltas() + { + var deltas = new RecordingCollectorDeltaCalculator(); + + PgDatabaseStatsCollector.Instance.WritePayload( + new PgDatabaseStatsCollector.Row("appdb", 1, 0, 2, 3, 0, 0, 0, null), + new RecordingCollectorRowWriter(), + MakeContext(deltas: deltas)); + + Assert.Empty(deltas.Calls); + } + + /// + /// Every payload column is written, in order. WritePayload is positional, so a column added without a + /// matching Value() shifts everything after it and stores data that is silently wrong. + /// + [Fact] + public void WritesEveryPayloadColumnInOrder() + { + var writer = new RecordingCollectorRowWriter(); + var reset = new DateTime(2026, 5, 18, 7, 4, 22); + + PgDatabaseStatsCollector.Instance.WritePayload( + new PgDatabaseStatsCollector.Row(null, 10, 1, 20, 30, 4, 5_000, 2, reset), + writer, + MakeContext()); + + Assert.Equal(PgDatabaseStatsCollector.Instance.PayloadColumns.Count, writer.Values.Count); + Assert.Null(writer.Values[0]); // database_name — the shared-relations row + Assert.Equal(10L, writer.Values[1]); // xact_commit + Assert.Equal(4L, writer.Values[5]); // temp_files + Assert.Equal(5_000L, writer.Values[6]); // temp_bytes + Assert.Equal(2L, writer.Values[7]); // deadlocks + Assert.Equal(reset, writer.Values[8]); + } + + [Fact] + public void RegisteredInBothTheCatalogAndTheSchedule() + { + Assert.Contains(CollectorCatalog.All, d => d.Name == "pg_database_stats"); + + var schedule = CollectorScheduleDefaults.All["pg_database_stats"]; + + /* Per-minute is affordable because pg_stat_database is cluster-wide — one query on the connection + the collector already holds, returning one row per database. The cadence is also the value: a + spill correlated to the minute names the deploy or the job that caused it. */ + Assert.Equal(1, schedule.FrequencyMinutes); + Assert.Equal(30, schedule.RetentionDays); + Assert.True(schedule.DefaultEnabled); + } + + /// + /// The capability vocabulary has a noun phrase for this collector, so a SQL Server target asked + /// get_pg_database_stats is told what it does not collect rather than "the data this read is served + /// from", which explains nothing. + /// + [Fact] + public void TheCapabilityMessageNamesWhatIsNotCollected() + { + var message = CollectorEngineCapability.NotCollectedMessage( + "sql-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.SqlServer, "pg_database_stats"); + + Assert.NotNull(message); + Assert.Contains("pg_stat_database", message, StringComparison.Ordinal); + Assert.Contains("temp-file spills", message, StringComparison.Ordinal); + Assert.DoesNotContain("the data this read is served from", message, StringComparison.Ordinal); + } +} diff --git a/Lite.Tests/PgDeadlockLogParserTests.cs b/Lite.Tests/PgDeadlockLogParserTests.cs new file mode 100644 index 000000000..89f73d4e8 --- /dev/null +++ b/Lite.Tests/PgDeadlockLogParserTests.cs @@ -0,0 +1,216 @@ +// Copyright (c) Erik Darling Data. All rights reserved. +// Licensed under the terms in the LICENSE file in the repository root. + +using System; +using System.Linq; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// The deadlock log parser (#2661). Every fixture here is REAL output captured from a PostgreSQL 17.11 +/// target, not a hand-written approximation of one — the two traps below are exactly the kind that a +/// plausible-looking fake would not have. +/// +public sealed class PgDeadlockLogParserTests +{ + /* Captured verbatim, including the %Q query id running straight into ERROR: with no separator. */ + private const string RealBlock = + "2026-08-26 22:25:24.100 UTC [1549] 322048460535975151ERROR: deadlock detected\n" + + "2026-08-26 22:25:24.100 UTC [1549] 322048460535975151DETAIL: Process 1549 waits for ShareLock on transaction 809; blocked by process 1556.\n" + + "\tProcess 1556 waits for ShareLock on transaction 808; blocked by process 1549.\n" + + "\tProcess 1549: \n" + + "\tBEGIN; UPDATE dl SET v=v+1 WHERE id=1; SELECT pg_sleep(2); UPDATE dl SET v=v+1 WHERE id=2; COMMIT;\n" + + "\tProcess 1556: \n" + + "\tBEGIN; UPDATE dl SET v=v+1 WHERE id=2; SELECT pg_sleep(2); UPDATE dl SET v=v+1 WHERE id=1; COMMIT;\n" + + "2026-08-26 22:25:24.100 UTC [1549] 322048460535975151HINT: See server log for query details.\n"; + + [Fact] + public void ParsesARealReport_WholeGraphAndBothStatements() + { + var found = PgDeadlockLogParser.Extract(RealBlock); + + var deadlock = Assert.Single(found); + + Assert.Equal(new DateTime(2026, 8, 26, 22, 25, 24, 100, DateTimeKind.Utc), deadlock.OccurredAtUtc); + Assert.Equal(1549, deadlock.VictimPid); + Assert.Equal(2, deadlock.ParticipantCount); + Assert.Equal("ShareLock", deadlock.LockModes); + Assert.Equal("transaction 808, transaction 809", deadlock.Resources); + + /* The victim's statement, not merely SOME statement: the report names both, and attributing the + wrong one to the cancelled session would point an investigation at the surviving query. */ + Assert.StartsWith("BEGIN; UPDATE dl SET v=v+1 WHERE id=1;", deadlock.VictimStatement, StringComparison.Ordinal); + + var statements = PgDeadlockLogParser.ParseStatements(deadlock.GraphText); + Assert.Equal(2, statements.Count); + Assert.Contains("WHERE id=1", statements[1549], StringComparison.Ordinal); + Assert.Contains("WHERE id=2", statements[1556], StringComparison.Ordinal); + } + + /// + /// %Q writes the query id with NO separator before the severity — the captured line really does + /// read [1549] 322048460535975151ERROR:. A pattern requiring whitespace there matches nothing, + /// and matches nothing in the way that looks like "this server has no deadlocks", which is the worst + /// possible failure for this collector. + /// + [Fact] + public void ParsesWithAndWithoutTheQueryIdInThePrefix() + { + Assert.Single(PgDeadlockLogParser.Extract(RealBlock)); + + var noQueryId = RealBlock.Replace("322048460535975151", string.Empty, StringComparison.Ordinal); + Assert.Single(PgDeadlockLogParser.Extract(noQueryId)); + } + + /// + /// A participant's statement is arbitrary user SQL and can span lines, each arriving tab-indented under + /// the DETAIL block. Taking one line per participant would truncate every multi-line statement to its + /// first — silently, because the row still looks fine. + /// + [Fact] + public void KeepsAMultiLineStatementWhole() + { + var multiline = + "2026-08-26 22:25:24.100 UTC [1549] ERROR: deadlock detected\n" + + "2026-08-26 22:25:24.100 UTC [1549] DETAIL: Process 1549 waits for ShareLock on transaction 809; blocked by process 1556.\n" + + "\tProcess 1556 waits for ShareLock on transaction 808; blocked by process 1549.\n" + + "\tProcess 1549: \n" + + "\tUPDATE orders\n" + + "\t SET status = 'shipped'\n" + + "\t WHERE id = 1;\n" + + "\tProcess 1556: \n" + + "\tUPDATE orders SET status = 'held' WHERE id = 2;\n" + + "2026-08-26 22:25:24.100 UTC [1549] HINT: See server log for query details.\n"; + + var deadlock = Assert.Single(PgDeadlockLogParser.Extract(multiline)); + + Assert.Contains("SET status = 'shipped'", deadlock.VictimStatement, StringComparison.Ordinal); + Assert.Contains("WHERE id = 1;", deadlock.VictimStatement, StringComparison.Ordinal); + } + + /// + /// Participants are counted from the wait EDGES, not from the Process N: statement headers: the + /// server omits a header when it could not recover the text, and a participant with no statement is + /// still in the cycle. Counting headers would under-report it. + /// + [Fact] + public void CountsParticipantsFromEdges_NotFromStatementHeaders() + { + var noStatements = + "2026-08-26 22:25:24.100 UTC [1549] ERROR: deadlock detected\n" + + "2026-08-26 22:25:24.100 UTC [1549] DETAIL: Process 1549 waits for ShareLock on transaction 809; blocked by process 1556.\n" + + "\tProcess 1556 waits for ShareLock on transaction 810; blocked by process 1601.\n" + + "\tProcess 1601 waits for ShareLock on transaction 808; blocked by process 1549.\n" + + "2026-08-26 22:25:24.100 UTC [1549] HINT: See server log for query details.\n"; + + var deadlock = Assert.Single(PgDeadlockLogParser.Extract(noStatements)); + + Assert.Equal(3, deadlock.ParticipantCount); + Assert.Null(deadlock.VictimStatement); + } + + /// + /// The resource is captured WHOLE rather than decomposed. It is transaction N in the common case + /// but also tuple (b,o) of relation N, relation N of database N and advisory locks, and an + /// enumeration would silently drop the shapes it did not anticipate. + /// + [Fact] + public void CarriesLockResourcesItWasNotWrittenAgainst() + { + var tupleLock = + "2026-08-26 22:25:24.100 UTC [1549] ERROR: deadlock detected\n" + + "2026-08-26 22:25:24.100 UTC [1549] DETAIL: Process 1549 waits for ExclusiveLock on tuple (0,2) of relation 16385 of database 16384; blocked by process 1556.\n" + + "\tProcess 1556 waits for ShareLock on transaction 808; blocked by process 1549.\n" + + "2026-08-26 22:25:24.100 UTC [1549] HINT: See server log for query details.\n"; + + var deadlock = Assert.Single(PgDeadlockLogParser.Extract(tupleLock)); + + Assert.Contains("tuple (0,2) of relation 16385 of database 16384", deadlock.Resources, StringComparison.Ordinal); + Assert.Contains("ExclusiveLock", deadlock.LockModes, StringComparison.Ordinal); + Assert.Contains("ShareLock", deadlock.LockModes, StringComparison.Ordinal); + } + + /// + /// The hash is identity across the overlapping reads the collector makes on purpose: the same report is + /// seen every cycle while it stays in the log tail, and must be stored once. Two DIFFERENT deadlocks + /// must not collide — verified on the rig, where a repeat of the same query pair produced a different + /// hash because the process IDs differed. + /// + [Fact] + public void HashIsStableForOneReportAndDistinctBetweenReports() + { + var first = Assert.Single(PgDeadlockLogParser.Extract(RealBlock)); + var again = Assert.Single(PgDeadlockLogParser.Extract(RealBlock)); + + Assert.Equal(first.DeadlockHash, again.DeadlockHash); + + var otherPids = RealBlock.Replace("1549", "1827", StringComparison.Ordinal) + .Replace("1556", "1830", StringComparison.Ordinal); + var other = Assert.Single(PgDeadlockLogParser.Extract(otherPids)); + + Assert.NotEqual(first.DeadlockHash, other.DeadlockHash); + } + + /// + /// A block with no wait edge is not a deadlock report, and a half-read one is ordinary rather than a + /// fault: both transports read a bounded window, so a report cut at the edge is whole on the next + /// overlapping pass. Skipped, never thrown. + /// + [Fact] + public void SkipsWhatItCannotParse_RatherThanThrowing() + { + Assert.Empty(PgDeadlockLogParser.Extract(null)); + Assert.Empty(PgDeadlockLogParser.Extract(string.Empty)); + Assert.Empty(PgDeadlockLogParser.Extract("nothing to see here")); + + var truncated = RealBlock[..RealBlock.IndexOf("blocked by process", StringComparison.Ordinal)]; + Assert.Empty(PgDeadlockLogParser.Extract(truncated)); + } + + /// + /// The collector declares what it writes. A mismatch here is the #2622 class of defect: the runtime + /// check only fires when the collector RUNS, and this one only runs where the log is readable. + /// + [Fact] + public void CollectorDeclaresEveryColumnItWrites() + { + var writer = new RecordingCollectorRowWriter(); + + PgDeadlocksCollector.Instance.WritePayload( + new PgDeadlocksCollector.Row( + new DateTime(2026, 8, 26, 22, 25, 24, DateTimeKind.Utc), + 1549, 2, "ABC123", "ShareLock", "transaction 808", "UPDATE dl ...", "graph"), + writer, + MakeContext()); + + Assert.Equal(PgDeadlocksCollector.Instance.PayloadColumns.Count, writer.Values.Count); + } + + private static CollectorContext MakeContext() => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 26, 12, 0, 0, DateTimeKind.Utc), + Deltas = new NoDeltas(), + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }; + + private sealed class NoDeltas : ICollectorDeltaCalculator + { + public long CalculateDelta(int serverId, string key, string metric, long current, DateTime? at = null, int i = 0) => 0; + + public long CalculateDeltaWithInterval(int serverId, string key, string metric, long current, out int seconds, DateTime? at = null, int i = 0) + { + seconds = 60; + return 0; + } + } +} diff --git a/Lite.Tests/PgExtensionAvailabilityCollectorDefinitionTests.cs b/Lite.Tests/PgExtensionAvailabilityCollectorDefinitionTests.cs new file mode 100644 index 000000000..f513cbe65 --- /dev/null +++ b/Lite.Tests/PgExtensionAvailabilityCollectorDefinitionTests.cs @@ -0,0 +1,209 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2545: the extension capability axis. The assertions are about the DISTINCTIONS the collector has to +/// draw — four states rather than a boolean, and the two catalogs' different scopes — not about wording. +/// +public class PgExtensionAvailabilityCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgExtensionAvailabilityCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// All four states must be expressible. A boolean would collapse available into absent, + /// and available is the only ACTIONABLE one — the entire reason this axis exists. + /// + [Fact] + public void AllFourStates_AreExpressible() + { + var sql = Sql(); + + foreach (var state in new[] { "absent", "available", "outdated", "installed" }) + { + Assert.Contains($"'{state}'", sql, StringComparison.Ordinal); + } + } + + /// + /// outdated must be decided by comparing the installed version against the server's default, and + /// by EQUALITY only. Extension versions are free-form strings, so ordering them needs a parser this has + /// no business carrying — and the server already names which one is default. + /// + [Fact] + public void Outdated_ComparesVersionsForInequality_NeverOrdersThem() + { + var sql = Sql(); + + Assert.Matches(new Regex(@"installed_version\s*<>\s*p?\.?default_version"), sql); + Assert.DoesNotMatch(new Regex(@"installed_version\s*[<>]\s*(?!>)"), sql.Replace("<>", "!=")); + } + + /// + /// The scope trap. pg_extension is PER-DATABASE and pg_available_extensions is + /// cluster-wide — measured on one cluster reporting an extension installed in one database and not in + /// another while the available list said yes in both. Both catalogs must be read, because either alone + /// answers a narrower question than it appears to. + /// + [Fact] + public void BothCatalogs_AreRead_BecauseTheirScopesDiffer() + { + var sql = Sql(); + + Assert.Contains("pg_catalog.pg_available_extensions", sql, StringComparison.Ordinal); + Assert.Contains("pg_catalog.pg_extension", sql, StringComparison.Ordinal); + } + + /// + /// A FULL OUTER join, not an inner or a left one. An extension installed in this database but no longer + /// offered by the server (a downgraded binary, a dropped contrib package) has a row in one catalog and + /// not the other, and dropping it would hide the case most worth seeing. + /// + [Fact] + public void TheCatalogsAreFullOuterJoined_SoNeitherSideIsLost() + => Assert.Equal(2, Regex.Matches(Sql(), @"FULL OUTER JOIN").Count); + + /// + /// Preload-only modules must NOT be in the relevant roster. auto_explain and + /// pg_wait_sampling never appear in pg_available_extensions on ANY server — including ones + /// actively running them — so listing them here would manufacture a permanent false absent. + /// + /// This exact defect shipped once already (#2564, fixed in #2584), which is why it is pinned + /// rather than left to the comment that explains it. + /// + [Fact] + public void PreloadOnlyModules_AreNotInTheRelevantRoster() + { + /* Scoped to the VALUES roster rather than the whole query, so the explanatory comment naming these + modules does not satisfy the assertion — the source-pin trap this repo has hit repeatedly. */ + var roster = Regex.Match(Sql(), @"WITH relevant \(name, comment\) AS \(\s*VALUES(.*?)\n\)", RegexOptions.Singleline); + + Assert.True(roster.Success, "the relevant-extension roster is no longer where this test looks for it"); + Assert.DoesNotContain("auto_explain", roster.Groups[1].Value, StringComparison.Ordinal); + Assert.DoesNotContain("pg_wait_sampling", roster.Groups[1].Value, StringComparison.Ordinal); + } + + /// + /// The roster exists ONLY so absence is reportable — absence is not a row in any catalog, so it cannot + /// be derived. Everything else about the result set is derived from the catalogs, which is why the + /// roster can stay small and why a server's other extensions still get recorded. + /// + [Fact] + public void TheRoster_NamesTheExtensionsThisProductCanActuallyUse() + { + var sql = Sql(); + + foreach (var name in new[] { "pg_stat_statements", "pgstattuple", "pg_buffercache", "pg_stat_kcache" }) + { + Assert.Contains($"('{name}'", sql, StringComparison.Ordinal); + } + } + + /// + /// Catalog reads are schema-qualified: pg_catalog is searched implicitly but not necessarily + /// FIRST, so an unqualified read can resolve to an object a user created in a schema earlier in the + /// monitoring login's search_path — which here would mean fabricating the capability answer. + /// + [Fact] + public void EveryCatalogRead_IsSchemaQualified() + { + /* COMMENTS STRIPPED FIRST. The query explains the per-database/cluster-wide scope split in a comment + that necessarily names both catalogs, and an unstripped scan matches those mentions and fails on a + perfectly qualified query. This repo has now hit that trap several times — a guard that greps for + an identifier finds the comment ABOUT the identifier — so the assertion is made against code. */ + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + foreach (Match match in Regex.Matches(sql, @"(\S*)pg_available_extensions")) + { + Assert.Equal("pg_catalog.", match.Groups[1].Value); + } + + /* Word-boundary matched, so the pg_extension_availability table name is not read as an unqualified + catalog reference. */ + foreach (Match match in Regex.Matches(sql, @"(\S*)\bpg_extension\b")) + { + Assert.Equal("pg_catalog.", match.Groups[1].Value); + } + } + + [Fact] + public void AppliesTo_EveryPostgresTarget() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + Assert.True(PgExtensionAvailabilityCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + } + } + + /// + /// Version columns are TEXT. Typing them as numbers is how a collector starts failing on a server whose + /// extension is versioned 2.0-beta. + /// + [Fact] + public void VersionColumns_AreText() + { + var columns = PgExtensionAvailabilityCollector.Instance.PayloadColumns; + + foreach (var name in new[] { "installed_version", "default_version" }) + { + Assert.Equal(CollectorColumnType.Varchar, columns.Single(c => c.Name == name).Type); + } + } + + /// One SELECT alias per payload column, in order — a mismatch is a silently shifted binary + /// COPY, which writes every value into the wrong column rather than failing. + [Fact] + public void SelectAliases_MatchThePayloadOrder() + { + var expected = PgExtensionAvailabilityCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + /* Scoped to the OUTER select. The CTEs above it alias their own columns with the same keyword, so + scanning the whole query collects three extra names and reads as an ordering bug in the collector + rather than as the test looking in the wrong place. The outer SELECT is the only one at column + zero, which is what this finds. */ + var lines = Sql().Split('\n'); + var outerSelect = Array.FindLastIndex(lines, l => l.TrimEnd() == "SELECT"); + + Assert.True(outerSelect >= 0, "the outer SELECT is no longer where this test looks for it"); + + var selected = lines + .Skip(outerSelect) + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal) + && !line.TrimStart().StartsWith("FULL OUTER JOIN", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(expected, selected); + } +} diff --git a/Lite.Tests/PgIndexBloatCollectorDefinitionTests.cs b/Lite.Tests/PgIndexBloatCollectorDefinitionTests.cs new file mode 100644 index 000000000..58c7f9cc0 --- /dev/null +++ b/Lite.Tests/PgIndexBloatCollectorDefinitionTests.cs @@ -0,0 +1,187 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2561: measured b-tree index bloat. The two assertions that matter are that non-btree indexes can never +/// reach pgstatindex (it RAISES on them, so one GIN index would take the whole collection down) and +/// that no derived bloat percentage is stored. +/// +public class PgIndexBloatCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgIndexBloatCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// Only b-trees may reach the function. Verified against a live server: GIN, BRIN and hash each raise + /// relation "x" is not a btree index, so a single one of them would fail the collection every + /// cycle. + /// + [Fact] + public void OnlyBtreeIndexes_AreCandidates() + => Assert.Matches(new Regex(@"am\.amname\s*=\s*'btree'"), Sql()); + + /// + /// The btree filter sits behind an OFFSET 0 fence so it is applied BEFORE the function call + /// rather than alongside it. In testing the planner did filter first without the fence — but when the + /// failure mode is the entire collection erroring, correctness should not depend on plan shape. + /// + [Fact] + public void TheCandidateFilter_IsFencedFromTheFunctionCall() + { + var sql = Sql(); + + Assert.Contains("OFFSET 0", sql, StringComparison.Ordinal); + + var fence = sql.IndexOf("OFFSET 0", StringComparison.Ordinal); + var call = sql.IndexOf("pgstatindex", StringComparison.Ordinal); + + Assert.True(call > fence, "pgstatindex is called before the candidate fence, so the btree filter may not have been applied yet"); + } + + /// + /// A LEFT join, so an index over the measurement ceiling still produces a row. A cross join would drop + /// exactly the indexes most likely to be holding reclaimable space. + /// + [Fact] + public void TheFunctionJoin_IsLeft_SoSkippedIndexesStillAppear() + => Assert.Matches(new Regex(@"LEFT JOIN LATERAL\s+public\.pgstatindex"), Sql()); + + /// + /// Skipping is recorded, never silent. A size cap that made indexes disappear would read as "no bloat + /// here" on precisely the biggest ones. + /// + [Fact] + public void SkippedIndexes_CarryAReason() + { + Assert.Contains("skipped_reason", PgIndexBloatCollector.Instance.PayloadColumns.Select(c => c.Name)); + Assert.Contains("skipped_reason", Sql(), StringComparison.Ordinal); + } + + /// + /// The density is stored raw. A freshly built index measures near 90 — between 89.98 and 91.48 across + /// the seven measured while designing this — so a stored "bloat percent" would bake in a false floor + /// that also varies per index. + /// + [Fact] + public void TheRawDensityIsStored_AndNoDerivedBloatPercent() + { + var names = PgIndexBloatCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("avg_leaf_density", names); + Assert.DoesNotContain(names, n => n.Contains("bloat_pct", StringComparison.Ordinal) + || n.Contains("bloat_percent", StringComparison.Ordinal)); + } + + /// + /// pgstatindex is an EXTENSION function and lives where pgstattuple was created, so it is + /// qualified public. — qualifying it pg_catalog. does not resolve at all (verified), and + /// leaving it unqualified would let an object earlier in search_path shadow it. + /// + [Fact] + public void TheExtensionFunction_IsQualifiedPublic_NotPgCatalog() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + Assert.Contains("public.pgstatindex", sql, StringComparison.Ordinal); + Assert.DoesNotContain("pg_catalog.pgstatindex", sql, StringComparison.Ordinal); + } + + /// Primaries only: a standby's index files are byte-identical by replication, so measuring + /// both pays the full-index read twice for one answer. + [Fact] + public void AppliesTo_PrimariesOnly() + { + Assert.True(PgIndexBloatCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + Assert.False(PgIndexBloatCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17, IsInRecovery = true })); + } + + /// Only valid, ready indexes — an invalid one from a failed CREATE INDEX CONCURRENTLY has no + /// meaningful bloat and is a different finding entirely. + [Fact] + public void InvalidIndexes_AreExcluded() + { + var sql = Sql(); + + Assert.Contains("x.indisvalid", sql, StringComparison.Ordinal); + Assert.Contains("x.indisready", sql, StringComparison.Ordinal); + } + + /// + /// The CYCLE has a work budget, not just each index (#2617). + /// + /// This is the assertion that would have caught a collector which never returned a row. + /// pgstatindex reads every page it is pointed at, and the 20 GB ceiling bounds one index while + /// nothing bounded the statement. Measured on a live Aurora target: 1,517 indexes totalling 461 GB in a + /// single statement, which never finished and dropped the connection mid-read — rows_ever = 0 for + /// the collector's entire life. The local rig had two indexes, so the question never arose there. + /// + [Fact] + public void TheCycleHasAWorkBudget_NotJustAPerIndexCeiling() + { + var sql = Sql(); + + /* Ranked by size so the measured ones are where bloat is worth reclaiming. */ + Assert.Contains("row_number() OVER (ORDER BY k.index_bytes DESC", sql, StringComparison.Ordinal); + + /* And the budget gates the LATERAL, which is the thing that costs pages. Gating only the + skipped_reason would label rows correctly while still reading every index. */ + Assert.Matches(new Regex(@"LEFT JOIN LATERAL[\s\S]*?size_rank\s*<=\s*\d+"), sql); + } + + /// + /// Over-budget indexes are RETURNED with a reason, never dropped. An index missing from the result + /// reads as one that does not exist; an index present with a stated reason cannot be mistaken for + /// healthy — the same argument that put skipped_reason on the size ceiling. + /// + [Fact] + public void OverBudgetIndexesAreReturnedWithAReason() + { + var sql = Sql(); + + Assert.Contains("not measured this cycle (work budget)", sql, StringComparison.Ordinal); + + /* No WHERE that would remove them from the result set. */ + Assert.DoesNotMatch(new Regex(@"WHERE[^;]*size_rank\s*<="), sql); + } + + /// + /// A command-timeout override, so a slow single index yields a CLASSIFIED timeout rather than the + /// unclassified Exception while reading from stream that #2617 actually surfaced as. + /// + [Fact] + public void ItOverridesTheCommandTimeout() + { + Assert.NotNull(PgIndexBloatCollector.Instance.CommandTimeoutSecondsOverride); + Assert.True(PgIndexBloatCollector.Instance.CommandTimeoutSecondsOverride >= 120); + } +} diff --git a/Lite.Tests/PgIndexUsageStatsCollectorDefinitionTests.cs b/Lite.Tests/PgIndexUsageStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..cdc7e7074 --- /dev/null +++ b/Lite.Tests/PgIndexUsageStatsCollectorDefinitionTests.cs @@ -0,0 +1,532 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the per-index usage collector (#2541): the ONE column that carries a version floor, the size floor +/// and the escape hatch that lets an invalid index through it, the writers-only gate, and — the substance of +/// the collector — the catalog facts that stop "unused" being read as "droppable". +/// +public class PgIndexUsageStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext( + int major = 17, ICollectorDeltaCalculator? deltas = null, string? currentDatabase = "appdb") + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc), + Deltas = deltas ?? s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + CurrentDatabaseName = currentDatabase, + }; + + /// + /// The table name cannot be changed later without a migration, and it must not be + /// pg_stat_user_indexes: pg_catalog is searched before search_path, so a store table by a catalog + /// name makes every unqualified read resolve to the MONITORING store's own view. + /// + [Fact] + public void Identity_Pinned_AndTheTableDoesNotShadowACatalogView() + { + Assert.Equal("pg_index_usage_stats", PgIndexUsageStatsCollector.Instance.Name); + Assert.Equal("pg_index_usage_stats", PgIndexUsageStatsCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgIndexUsageStatsCollector.Instance.TargetEngine); + + Assert.NotEqual("pg_stat_user_indexes", PgIndexUsageStatsCollector.Instance.TargetTable); + Assert.NotEqual("pg_stat_all_indexes", PgIndexUsageStatsCollector.Instance.TargetTable); + Assert.NotEqual("pg_statio_user_indexes", PgIndexUsageStatsCollector.Instance.TargetTable); + } + + /// + /// WRITERS ONLY, and the gate reads IsInRecovery and nothing else — no version floor and no + /// Aurora gate, so a stock PostgreSQL 13 writer is collected exactly like an Aurora 17 one. + /// + /// The recovery half is the decision worth pinning. On a standby idx_scan counts the + /// STANDBY's own scans, so an index the primary's workload uses a million times an hour reads as zero + /// there — and this collector's whole output is a judgement about whether a zero means "drop it". A + /// confidently wrong "unused" sends someone to DROP INDEX; a missing answer only sends them looking. + /// + [Theory] + [InlineData(13, false, false, true)] + [InlineData(15, false, false, true)] + [InlineData(16, true, false, true)] + [InlineData(17, true, false, true)] + [InlineData(18, false, false, true)] + [InlineData(13, false, true, false)] + [InlineData(16, true, true, false)] + [InlineData(17, true, true, false)] + [InlineData(18, false, true, false)] + public void AppliesToWritersOnly_OnEveryMajorAndBothAuroraAndStock( + int major, bool isAurora, bool inRecovery, bool expected) + { + Assert.Equal(expected, PgIndexUsageStatsCollector.Instance.AppliesTo(new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + IsAurora = isAurora, + IsInRecovery = inRecovery, + })); + } + + /// + /// The composed gate still refuses a SQL Server target. The collector's own AppliesTo only asks about + /// recovery, so the ENGINE half is the only thing standing between this PostgreSQL query text and a SQL + /// Server connection — the #2213 class of defect, which asserting the collector's gate alone cannot see. + /// + [Fact] + public void TheEngineHalfOfTheDispatchGateStillRefusesASqlServerTarget() + { + Assert.False(CollectorCatalog.AppliesTo(PgIndexUsageStatsCollector.Instance, new CollectorTargetInfo())); + Assert.True(CollectorCatalog.AppliesTo( + PgIndexUsageStatsCollector.Instance, + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql })); + } + + /// + /// pg_stat_user_indexes is scoped to the connected database and PostgreSQL has no cross-database + /// read, so this is necessarily a fan-out — and on PostgreSQL a fan-out is one CONNECTION per database + /// per cycle, which is what the daily cadence pays for. + /// + [Fact] + public void RunsPerDatabase() + { + Assert.True(PgIndexUsageStatsCollector.Instance.RunsPerDatabase(MakeContext().Target)); + } + + /// + /// MEASURED, not read from the documentation: last_idx_scan arrived in PostgreSQL 16. The gated-ON + /// form was executed against live PostgreSQL 13, 14 and 15 and each returned + /// ERROR: column i.last_idx_scan does not exist — which fails the WHOLE collection for that + /// database, every cycle, not just that column. Hence the substitution rather than a bare select. + /// + [Theory] + [InlineData(13)] + [InlineData(14)] + [InlineData(15)] + public void OmitsLastIdxScanBelowPostgres16_SubstitutingATypedNull(int major) + { + var sql = PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(major)).Text; + + Assert.DoesNotContain("last_idx_scan", sql, StringComparison.Ordinal); + + /* A TYPED null, so the row SHAPE does not change across a mixed-version fleet. NULL is also the + honest value: on PostgreSQL 15 the server genuinely does not know when the index was last + scanned, and a sentinel timestamp would have to be a real instant and would read as one. */ + Assert.Contains("NULL::timestamp", sql, StringComparison.Ordinal); + } + + /// + /// And on 16 and above it IS selected, converted to UTC. Without this half the pin would pass on a + /// collector that had quietly stopped collecting the column everywhere. + /// + [Theory] + [InlineData(16)] + [InlineData(17)] + [InlineData(18)] + public void SelectsLastIdxScanOnPostgres16AndAbove(int major) + { + var sql = PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(major)).Text; + + Assert.Contains("i.last_idx_scan", sql, StringComparison.Ordinal); + Assert.DoesNotContain("NULL::timestamp", sql, StringComparison.Ordinal); + } + + /// + /// timestamptz::text renders in the SESSION's TimeZone and the store contract is naive UTC, so + /// both timestamps convert explicitly. Byte-identical to UTC on every instance in the fleet, which is + /// exactly why it survives every probe and has to be pinned instead. + /// + [Fact] + public void ConvertsBothTimestampsToUtcExplicitly() + { + var sql = PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(17)).Text; + + Assert.Contains("(i.last_idx_scan AT TIME ZONE 'UTC')", sql, StringComparison.Ordinal); + Assert.Contains("AT TIME ZONE 'UTC') AS stats_reset", sql, StringComparison.Ordinal); + + Assert.DoesNotContain("last_idx_scan::timestamp", sql, StringComparison.Ordinal); + Assert.DoesNotContain("stats_reset::timestamp", sql, StringComparison.Ordinal); + } + + /// + /// The size floor is the fan-out cost control, and the OR NOT s.is_valid beside it is the escape + /// hatch that keeps it honest. MEASURED live: a 16,384-byte INVALID index came through a 65,536-byte + /// floor because of that clause — an invalid index is a finding at any size, since the planner will + /// never use it while writes still maintain it. + /// + [Fact] + public void FiltersOnTheSizeFloor_WithAnEscapeHatchForInvalidIndexes() + { + var sql = PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(17)).Text; + + Assert.Contains("WHERE s.index_bytes >= 65536", sql, StringComparison.Ordinal); + Assert.Contains("OR NOT s.is_valid", sql, StringComparison.Ordinal); + } + + /// + /// The droppability facts, each present by name. THIS IS THE COLLECTOR: "nobody scans it" is the easy + /// half, and a column dropped from here turns a safe read into one that tells somebody to drop their + /// primary key. A uniqueness check is not a scan of the kind idx_scan counts, so the single most + /// common zero-scan index on any schema is one that must never be dropped — and nothing in the usage + /// counters can say so. + /// + [Theory] + [InlineData("x.indisunique")] + [InlineData("x.indisprimary")] + [InlineData("x.indisvalid")] + [InlineData("x.indisready")] + [InlineData("x.indisreplident")] + [InlineData("x.indpred IS NOT NULL")] + [InlineData("x.indexprs IS NOT NULL")] + [InlineData("con.conindid = i.indexrelid")] + [InlineData("pg_get_indexdef(i.indexrelid)")] + public void SelectsEveryFactThatDecidesWhetherAnIndexCanActuallyGo(string fragment) + { + Assert.Contains(fragment, PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(17)).Text, StringComparison.Ordinal); + } + + /// + /// The usage counters and the block counters. idx_blks_hit matters next to idx_scan + /// because it accrues on WRITES too: an index with zero scans and millions of block accesses is being + /// maintained by the write path and read by nobody, stated in the server's own units. + /// + [Theory] + [InlineData("i.idx_scan::bigint")] + [InlineData("i.idx_tup_read::bigint")] + [InlineData("i.idx_tup_fetch::bigint")] + [InlineData("io.idx_blks_read")] + [InlineData("io.idx_blks_hit")] + [InlineData("pg_relation_size(i.indexrelid)")] + [InlineData("pg_relation_size(i.relid)")] + public void SelectsTheUsageAndCostCounters(string fragment) + { + Assert.Contains(fragment, PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(17)).Text, StringComparison.Ordinal); + } + + /// + /// Every catalog reference is schema-qualified. pg_catalog is searched implicitly and FIRST, so + /// an unqualified reference is not wrong today — it is a hostage to whatever a future search_path or a + /// user object named after a catalog view does to it. + /// + [Theory] + [InlineData("pg_catalog.pg_stat_user_indexes")] + [InlineData("pg_catalog.pg_statio_user_indexes")] + [InlineData("pg_catalog.pg_index")] + [InlineData("pg_catalog.pg_class")] + [InlineData("pg_catalog.pg_am")] + [InlineData("pg_catalog.pg_constraint")] + [InlineData("pg_catalog.pg_stat_database")] + public void SchemaQualifiesEveryCatalogReference(string reference) + { + Assert.Contains(reference, PgIndexUsageStatsCollector.Instance.BuildQuery(MakeContext(17)).Text, StringComparison.Ordinal); + } + + [Fact] + public void PayloadColumns_OrderAndKeyTypes_Pinned() + { + var columns = PgIndexUsageStatsCollector.Instance.PayloadColumns; + + Assert.Equal(24, columns.Count); + Assert.Equal( + new[] + { + "database_name", "schema_name", "table_name", "index_name", + "index_scans", "tuples_read", "tuples_fetched", + "blocks_read", "blocks_hit", + "index_bytes", "table_bytes", + "is_unique", "is_primary_key", "is_valid", "is_ready", + "is_replica_identity", "is_partial", "is_expression", "supports_constraint", + "index_method", "column_count", "index_definition", + "last_scan", "stats_reset", + }, + columns.Select(c => c.Name).ToArray()); + + Assert.Equal(CollectorColumnType.Varchar, columns[0].Type); // database_name + Assert.Equal(CollectorColumnType.BigInt, columns[4].Type); // index_scans + Assert.Equal(CollectorColumnType.BigInt, columns[9].Type); // index_bytes + Assert.Equal(CollectorColumnType.Boolean, columns[18].Type); // supports_constraint + Assert.Equal(CollectorColumnType.Integer, columns[20].Type); // column_count + Assert.Equal(CollectorColumnType.Varchar, columns[21].Type); // index_definition + Assert.Equal(CollectorColumnType.Timestamp, columns[22].Type); // last_scan + Assert.Equal(CollectorColumnType.Timestamp, columns[23].Type); // stats_reset + } + + /// + /// Every field mapped to its own ordinal. The values are deliberately all DIFFERENT and the booleans + /// deliberately alternate, so a transposed pair fails rather than passing on two equal values. + /// + [Fact] + public async Task ReadsAFullyPopulatedRow_WithEveryFieldOnItsOwnOrdinal() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "widget", "widget_status_idx", + 418L, 9_042L, 8_113L, // index_scans, tuples_read, tuples_fetched + 276L, 852_057L, // blocks_read, blocks_hit + 6_758_400L, 56_516_608L, // index_bytes, table_bytes + true, false, true, false, true, false, true, false, // the eight droppability booleans + "btree", 2, + "CREATE INDEX widget_status_idx ON public.widget USING btree (status, qty)", + new DateTime(2026, 8, 19, 6, 30, 0, DateTimeKind.Unspecified), + new DateTime(2026, 5, 18, 7, 4, 22, DateTimeKind.Unspecified), + }); + + var rows = await PgIndexUsageStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal("public", row.SchemaName); + Assert.Equal("widget", row.TableName); + Assert.Equal("widget_status_idx", row.IndexName); + Assert.Equal(418L, row.IndexScans); + Assert.Equal(9_042L, row.TuplesRead); + Assert.Equal(8_113L, row.TuplesFetched); + Assert.Equal(276L, row.BlocksRead); + Assert.Equal(852_057L, row.BlocksHit); + Assert.Equal(6_758_400L, row.IndexBytes); + Assert.Equal(56_516_608L, row.TableBytes); + + /* The alternating booleans, each on its own ordinal. Ordinals 10-17 in select order. */ + Assert.True(row.IsUnique); + Assert.False(row.IsPrimaryKey); + Assert.True(row.IsValid); + Assert.False(row.IsReady); + Assert.True(row.IsReplicaIdentity); + Assert.False(row.IsPartial); + Assert.True(row.IsExpression); + Assert.False(row.SupportsConstraint); + + Assert.Equal("btree", row.IndexMethod); + Assert.Equal(2, row.ColumnCount); + Assert.Equal( + "CREATE INDEX widget_status_idx ON public.widget USING btree (status, qty)", + row.IndexDefinition); + Assert.Equal(new DateTime(2026, 8, 19, 6, 30, 0), row.LastScan); + Assert.Equal(new DateTime(2026, 5, 18, 7, 4, 22), row.StatsReset); + } + + /// + /// The two NULL shapes this collector really produces, and they mean different things. + /// + /// The COUNTERS are NOT NULL in the catalog, so an absent one is a shape change rather than a + /// value and 0 is the correct reading of "this cumulative count has never moved". The TIMESTAMPS are + /// genuinely nullable and must stay null: last_scan is NULL on PostgreSQL 15 and below (the + /// server does not record it) and on 16+ for an index never scanned since the reset; + /// stats_reset is NULL until the first reset ever, which is the ordinary state and means + /// "the counters run back to the beginning", NOT "unknown". + /// + [Fact] + public async Task CountersFallBackToZero_WhileTheTimestampsStayNull() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "widget", "widget_never_scanned_idx", + DBNull.Value, DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, + DBNull.Value, + DBNull.Value, + }); + + var rows = await PgIndexUsageStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal(0L, row.IndexScans); + Assert.Equal(0L, row.TuplesRead); + Assert.Equal(0L, row.BlocksHit); + Assert.Equal(0L, row.IndexBytes); + + /* An absent boolean must never read TRUE: is_valid = false is what the read reports as "safe to + drop", and is_unique / supports_constraint = true is what STOPS it recommending a drop. */ + Assert.False(row.IsUnique); + Assert.False(row.IsPrimaryKey); + Assert.False(row.SupportsConstraint); + + Assert.Equal(string.Empty, row.IndexMethod); + Assert.Equal(0, row.ColumnCount); + Assert.Equal(string.Empty, row.IndexDefinition); + + Assert.Null(row.LastScan); + Assert.Null(row.StatsReset); + } + + /// + /// A PostgreSQL 15 target: last_scan arrives NULL because the server does not record it, and the + /// row must still be complete in every other respect rather than being dropped. + /// + [Fact] + public async Task ReadsARowFromAMajorThatDoesNotRecordLastScan() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "widget", "widget_pkey", + 200_000L, 0L, 0L, + 2L, 800_429L, + 4_513_792L, 49_233_920L, + true, true, true, true, false, false, false, true, + "btree", 1, + "CREATE UNIQUE INDEX widget_pkey ON public.widget USING btree (id)", + DBNull.Value, + DBNull.Value, + }); + + var rows = await PgIndexUsageStatsCollector.Instance.ReadAsync( + reader, MakeContext(major: 15), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Null(row.LastScan); + Assert.Equal(200_000L, row.IndexScans); + Assert.True(row.IsPrimaryKey); + Assert.True(row.SupportsConstraint); + } + + [Fact] + public async Task ReturnsNoRowsWhenTheViewReturnsNone() + { + var rows = await PgIndexUsageStatsCollector.Instance.ReadAsync( + new FakeCollectorDataReader(), MakeContext(), CancellationToken.None); + + Assert.Empty(rows); + } + + /// + /// No stored deltas: the cumulative counter is stored RAW so the window is differenced at read time. + /// That is what lets the read answer BOTH questions the data supports — the lifetime count and the + /// windowed one — and keeps a statistics reset visible instead of smoothed away at write time. + /// + [Fact] + public void TakesNoDeltas() + { + var deltas = new RecordingCollectorDeltaCalculator(); + + PgIndexUsageStatsCollector.Instance.WritePayload( + SampleRow(), + new RecordingCollectorRowWriter(), + MakeContext(deltas: deltas)); + + Assert.Empty(deltas.Calls); + } + + /// + /// Every payload column is written, in order. WritePayload is positional, so a column added without a + /// matching Value() shifts everything after it and stores data that is silently wrong rather than + /// failing. + /// + [Fact] + public void WritesEveryPayloadColumnInOrder() + { + var writer = new RecordingCollectorRowWriter(); + + PgIndexUsageStatsCollector.Instance.WritePayload(SampleRow(), writer, MakeContext()); + + Assert.Equal(PgIndexUsageStatsCollector.Instance.PayloadColumns.Count, writer.Values.Count); + + /* The connection's database, NOT a value parsed from the result set: the per-database loop sets it + and the connection's database IS the row's database, which a parsed value could not guarantee. */ + Assert.Equal("appdb", writer.Values[0]); + + Assert.Equal("public", writer.Values[1]); + Assert.Equal("widget", writer.Values[2]); + Assert.Equal("widget_status_idx", writer.Values[3]); + Assert.Equal(418L, writer.Values[4]); // index_scans + Assert.Equal(852_057L, writer.Values[8]); // blocks_hit + Assert.Equal(6_758_400L, writer.Values[9]); // index_bytes + Assert.Equal(true, writer.Values[11]); // is_unique + Assert.Equal(false, writer.Values[18]); // supports_constraint + Assert.Equal("btree", writer.Values[19]); + Assert.Equal(2, writer.Values[20]); // column_count + Assert.Equal(new DateTime(2026, 8, 19, 6, 30, 0), writer.Values[22]); // last_scan + Assert.Equal(new DateTime(2026, 5, 18, 7, 4, 22), writer.Values[23]); // stats_reset + } + + private static PgIndexUsageStatsCollector.Row SampleRow() => new( + SchemaName: "public", + TableName: "widget", + IndexName: "widget_status_idx", + IndexScans: 418, + TuplesRead: 9_042, + TuplesFetched: 8_113, + BlocksRead: 276, + BlocksHit: 852_057, + IndexBytes: 6_758_400, + TableBytes: 56_516_608, + IsUnique: true, + IsPrimaryKey: false, + IsValid: true, + IsReady: false, + IsReplicaIdentity: true, + IsPartial: false, + IsExpression: true, + SupportsConstraint: false, + IndexMethod: "btree", + ColumnCount: 2, + IndexDefinition: "CREATE INDEX widget_status_idx ON public.widget USING btree (status, qty)", + LastScan: new DateTime(2026, 8, 19, 6, 30, 0), + StatsReset: new DateTime(2026, 5, 18, 7, 4, 22)); + + /// + /// DAILY, and the cadence is inherited from index_object_stats — the SQL Server collector answering the + /// same question — rather than from the PostgreSQL fan-out sibling. "Has anything scanned this index" is + /// a structural question, not a rate, so an hourly sample would record the same catalog facts 24 times a + /// day at 24x the fan-out connections. + /// + /// 90 days is the number that actually matters: the retention window IS the evidence, because an + /// index can only be called unused for as long as we have been watching it. 30 days cannot clear a + /// monthly report. + /// + [Fact] + public void RegisteredInBothTheCatalogAndTheSchedule() + { + Assert.Contains(CollectorCatalog.All, d => d.Name == "pg_index_usage_stats"); + + var schedule = CollectorScheduleDefaults.All["pg_index_usage_stats"]; + + Assert.Equal(1440, schedule.FrequencyMinutes); + Assert.Equal(90, schedule.RetentionDays); + Assert.True(schedule.DefaultEnabled); + } + + /// + /// The capability vocabulary has a noun phrase for this collector, so a SQL Server target asked + /// get_pg_index_usage is told what it does not collect — and pointed at the fact that get_index_usage is + /// the tool there — rather than getting the generic fallback, which explains nothing. + /// + [Fact] + public void TheCapabilityMessageNamesWhatIsNotCollected() + { + var message = CollectorEngineCapability.NotCollectedMessage( + "sql-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.SqlServer, "pg_index_usage_stats"); + + Assert.NotNull(message); + Assert.Contains("pg_stat_user_indexes", message, StringComparison.Ordinal); + Assert.Contains("droppability", message, StringComparison.Ordinal); + Assert.DoesNotContain("the data this read is served from", message, StringComparison.Ordinal); + } +} diff --git a/Lite.Tests/PgIoStatsCollectorDefinitionTests.cs b/Lite.Tests/PgIoStatsCollectorDefinitionTests.cs index 319c4882c..ff3dc6392 100644 --- a/Lite.Tests/PgIoStatsCollectorDefinitionTests.cs +++ b/Lite.Tests/PgIoStatsCollectorDefinitionTests.cs @@ -1,4 +1,4 @@ -/* +/* * Copyright (c) 2026 Erik Darling, Darling Data LLC * * This file is part of the SQL Server Performance Monitor Lite. @@ -98,9 +98,14 @@ public void NeverRunsPerDatabase() } /// - /// PG18 REMOVED op_bytes (replaced by read_bytes / write_bytes / extend_bytes). Selecting it there - /// would fail with "column does not exist" and take the whole collection down, so it is substituted — - /// and the substitution must keep the column count identical so the stored shape cannot drift. + /// PG18 REMOVED op_bytes. Selecting it there would fail with "column does not exist" and take the + /// whole collection down, so it is substituted — and the substitution must keep the column count + /// identical so the stored shape cannot drift. + /// + /// Since V101 (#2655) the exchange runs BOTH ways, which is the point of pinning both arms: 18 + /// loses op_bytes and gains the three measured byte columns, while 17 and below select op_bytes and + /// NULL for the three. Same shape either way, and each server generation carries whichever quantity it + /// actually has. /// [Fact] public void SubstitutesOpBytesOnPg18WithoutChangingShape() @@ -111,6 +116,12 @@ public void SubstitutesOpBytesOnPg18WithoutChangingShape() Assert.Contains(" op_bytes ", pg17, StringComparison.Ordinal); Assert.Contains("NULL::bigint", pg18, StringComparison.Ordinal); Assert.Equal(pg17.Split(" AS ").Length, pg18.Split(" AS ").Length); + + /* The measured columns are real on 18 and NULL below it — never the other way round, which is the + failure that would make a pre-18 store claim byte totals it never had. */ + Assert.Contains(" read_bytes ", pg18, StringComparison.Ordinal); + Assert.Contains("NULL::numeric", pg17, StringComparison.Ordinal); + Assert.DoesNotContain("NULL::numeric", pg18, StringComparison.Ordinal); } /// @@ -161,12 +172,22 @@ public void PayloadColumns_CountAndKeyTypes_Pinned() { var columns = PgIoStatsCollector.Instance.PayloadColumns; - Assert.Equal(18, columns.Count); + Assert.Equal(21, columns.Count); Assert.Equal("backend_type", columns[0].Name); Assert.Equal("context", columns[2].Name); Assert.Equal("reads", columns[3].Name); Assert.Equal(CollectorColumnType.Double, columns[4].Type); // read_time_ms Assert.Equal(CollectorColumnType.Timestamp, columns[17].Type); // stats_reset + + /* V101 (#2655): PostgreSQL 18's measured byte totals, appended so no stored ordinal moved. + Decimal, not BigInt — PostgreSQL declares them `numeric` while the counts beside them are + `bigint`, and a byte total is exactly the quantity that outgrows a narrowing nobody promised. */ + Assert.Equal("read_bytes", columns[18].Name); + Assert.Equal("write_bytes", columns[19].Name); + Assert.Equal("extend_bytes", columns[20].Name); + Assert.Equal(CollectorColumnType.Decimal, columns[18].Type); + Assert.Equal(CollectorColumnType.Decimal, columns[19].Type); + Assert.Equal(CollectorColumnType.Decimal, columns[20].Type); } /// @@ -188,6 +209,9 @@ public async Task PreservesNullWriteCountersFromAurora() 10_150_937_570L, 0L, 0L, // hits, evictions, reuses DBNull.Value, DBNull.Value, // fsyncs, fsync_time — NULL on Aurora new DateTime(2026, 5, 18, 7, 4, 22, DateTimeKind.Unspecified), + /* A pre-18 server, so the measured byte columns do not exist and the collector selects + NULL for them (#2655). op_bytes above is the byte answer here. */ + DBNull.Value, DBNull.Value, DBNull.Value, // read_bytes, write_bytes, extend_bytes }); var rows = await PgIoStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); @@ -225,6 +249,7 @@ public async Task PreservesNullForCombinationsWhereACounterDoesNotApply() DBNull.Value, // hits does not apply either DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, // read_bytes, write_bytes, extend_bytes }); var rows = await PgIoStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); @@ -248,7 +273,7 @@ public void TakesNoDeltas() PgIoStatsCollector.Instance.WritePayload( new PgIoStatsCollector.Row( "client backend", "relation", "normal", 1, 1.0, null, null, null, null, - 0, 0.0, 8192, 2, 0, 0, null, null, null), + 0, 0.0, 8192, 2, 0, 0, null, null, null, null, null, null), new RecordingCollectorRowWriter(), MakeContext(deltas: deltas)); @@ -264,7 +289,7 @@ public void WritesEveryPayloadColumn() PgIoStatsCollector.Instance.WritePayload( new PgIoStatsCollector.Row( "client backend", "relation", "bulkread", 5, 2.5, null, null, null, null, - null, null, 8192, 7, null, 3, null, null, null), + null, null, 8192, 7, null, 3, null, null, null, null, null, null), writer, MakeContext()); diff --git a/Lite.Tests/PgKernelStatsCollectorDefinitionTests.cs b/Lite.Tests/PgKernelStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..43d43c366 --- /dev/null +++ b/Lite.Tests/PgKernelStatsCollectorDefinitionTests.cs @@ -0,0 +1,141 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the kcache collector on the four things that produce plausible, wrong output when they are missed. +/// +public class PgKernelStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext() + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + PostgresVersionNum = 170000, + }, + }; + + private static string Sql => PgKernelStatsCollector.Instance.BuildQuery(MakeContext()).Text; + + [Fact] + public void Identity_IsTheTableAndEngineTheStoreExpects() + { + Assert.Equal("pg_kernel_stats", PgKernelStatsCollector.Instance.Name); + Assert.Equal("pg_kernel_stats", PgKernelStatsCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgKernelStatsCollector.Instance.TargetEngine); + } + + /// + /// Only top-level statements. With pg_stat_statements.track = 'all' a nested statement appears + /// both on its own row and inside its caller's, so summing them double-counts every function body on + /// the server. The rig ran track = 'top', where the filter is inert — which is exactly why it + /// needs a test rather than a memory. + /// + [Fact] + public void OnlyTopLevelStatements_AreCounted() + { + Assert.Matches(new Regex(@"WHERE\s+k\.top\b"), Sql); + } + + /// + /// An extension function lives in the schema the extension was created in, not pg_catalog — + /// pg_catalog.pgstatindex did not resolve at all when #2561 tried it. + /// + [Fact] + public void TheExtensionFunction_IsQualifiedWhereItActuallyLives() + { + Assert.Contains("public.pg_stat_kcache()", Sql, StringComparison.Ordinal); + Assert.DoesNotContain("pg_catalog.pg_stat_kcache", Sql, StringComparison.Ordinal); + } + + /// + /// Seconds in, milliseconds out, and the column names say which. The function reports seconds; storing + /// them under a name ending in _ms without the conversion would be off by a thousand and look + /// entirely plausible. + /// + [Fact] + public void TimesAreConvertedToMilliseconds_AndNamedThatWay() + { + foreach (var column in new[] { "exec_user_time_ms", "exec_system_time_ms", "plan_cpu_time_ms" }) + { + Assert.Contains(column, PgKernelStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + } + + Assert.Contains("* 1000.0", Sql, StringComparison.Ordinal); + } + + /// + /// I/O columns are named _bytes. These are bytes that reached the device, and a column called + /// exec_reads sitting in a SQL Server codebase would be read as a logical-read count — a + /// different measure entirely, and one this cannot be compared against. + /// + [Fact] + public void IoColumns_AreNamedAsBytes() + { + var names = PgKernelStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("exec_read_bytes", names); + Assert.Contains("exec_write_bytes", names); + Assert.DoesNotContain("exec_reads", names); + Assert.DoesNotContain("exec_writes", names); + } + + /// + /// Ranked by CPU. Elapsed time already lives in pg_statement_stats; CPU is what this adds, and + /// ranking by bytes would answer a different question than the panel asks. + /// + [Fact] + public void ItRanksByCpu_NotByBytes() + { + Assert.Matches(new Regex(@"ORDER BY\s+sum\(k\.exec_user_time \+ k\.exec_system_time\) DESC"), Sql); + } + + /// + /// stats_since is collected so a reset is a recorded fact rather than an inference from a + /// counter that moved backwards — the improvement this table can afford over the wait profile. + /// + [Fact] + public void TheResetStamp_IsCollected() + { + Assert.Contains("stats_since", PgKernelStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + } + + /// + /// Per-database attribution is legitimate here and the collector takes it: the function exposes + /// dbid and pg_database is a SHARED catalog, so the name resolves from whichever database + /// the collector connected to. Contrast pg_wait_sampling, which has no such column and therefore + /// claims no such thing (#2599). + /// + [Fact] + public void ItAttributesToADatabase_BecauseTheCatalogSupportsIt() + { + Assert.Contains("database_name", PgKernelStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + Assert.Contains("pg_catalog.pg_database", Sql, StringComparison.Ordinal); + + /* Cluster-wide from one connection, so running per database would collect it N times over. */ + Assert.False(PgKernelStatsCollector.Instance.RunsPerDatabase( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + } +} diff --git a/Lite.Tests/PgLockStatsCollectorDefinitionTests.cs b/Lite.Tests/PgLockStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..acb423835 --- /dev/null +++ b/Lite.Tests/PgLockStatsCollectorDefinitionTests.cs @@ -0,0 +1,204 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2544, the locks slice. The assertions are about what this collector must capture that +/// structurally cannot — the lock MODE and the RELATION — and about the +/// scope mismatch between the two catalogs it reads. +/// +public class PgLockStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgLockStatsCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// The whole reason this collector exists. PgBlockingCollector reads pg_blocking_pids() + /// and never touches pg_locks, so the mode and the relation are unavailable today — and the mode + /// is what decides the remedy. + /// + [Fact] + public void ItReadsPgLocks_WhichTheBlockingCollectorDoesNot() + { + Assert.Contains("pg_catalog.pg_locks", Sql(), StringComparison.Ordinal); + + var blocking = PgBlockingCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /* Asserted rather than assumed: if the blocking collector ever grows a pg_locks read, these two + overlap and that is a decision to make deliberately, not to discover. */ + Assert.DoesNotContain("pg_locks", blocking, StringComparison.Ordinal); + } + + /// + /// Mode and granted must both be captured and grouped on. Either alone is uninterpretable: a mode + /// without granted cannot distinguish a queue from ordinary held locks, and granted without the mode + /// cannot distinguish a DDL queue from write concurrency. + /// + [Fact] + public void ModeAndGranted_AreBothGroupedOn() + { + var sql = Sql(); + + Assert.Matches(new Regex(@"GROUP BY[^;]*l\.mode"), sql); + Assert.Matches(new Regex(@"GROUP BY[^;]*l\.granted"), sql); + } + + /// + /// Both relation columns are stored. pg_locks is cluster-wide and pg_class is + /// per-database, so a lock in another database has an OID and no name — measured, not assumed. Keeping + /// only the name would drop the row; keeping only the OID would make the common case unreadable. + /// + [Fact] + public void BothRelationColumns_AreStored_BecauseTheNameCanBeUnresolvable() + { + var names = PgLockStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("relation_oid", names); + Assert.Contains("relation_name", names); + } + + /// + /// The relation join must be LEFT. An inner join silently drops every non-relation lock — + /// transactionid, virtualxid, advisory — and every lock held in another database, + /// which is the contention most worth seeing. + /// + [Fact] + public void TheRelationJoin_IsLeft_SoNonRelationLocksSurvive() + => Assert.Matches(new Regex(@"LEFT JOIN pg_catalog\.pg_class"), Sql()); + + /// + /// relation_oid is bigint, not integer. PostgreSQL OIDs are UNSIGNED 32-bit, so one + /// past 2^31 lands negative in a signed int — rare, and silently wrong when it happens. + /// + [Fact] + public void RelationOid_IsBigInt_BecauseOidsAreUnsigned32Bit() + => Assert.Equal( + CollectorColumnType.BigInt, + PgLockStatsCollector.Instance.PayloadColumns.Single(c => c.Name == "relation_oid").Type); + + /// + /// The collector excludes its own backend. Without it every snapshot reports the AccessShareLocks the + /// collector itself holds on the catalogs it is reading — the same self-marker rule the statement + /// collector applies. + /// + [Fact] + public void ItExcludesItsOwnBackend() + => Assert.Matches(new Regex(@"l\.pid\s*<>\s*pg_catalog\.pg_backend_pid\(\)"), Sql()); + + /// + /// Wait time is measured from state_change, not query_start. A backend waiting on a lock + /// has been in its current STATE since it began waiting, whereas query_start also covers the work + /// it did before hitting the lock — which would overstate every wait by however long the statement had + /// already been running. + /// + [Fact] + public void WaitTime_ComesFromStateChange_NotQueryStart() + { + var sql = Sql(); + + Assert.Contains("a.state_change", sql, StringComparison.Ordinal); + Assert.DoesNotContain("a.query_start", sql, StringComparison.Ordinal); + } + + /// + /// The wait is only computed for ungranted rows. A granted lock has no wait, and reporting 0 would read + /// as "granted instantly" — a measurement — where NULL is the absence of one. + /// + [Fact] + public void WaitIsComputedOnlyForUngrantedRows() + => Assert.Matches(new Regex(@"CASE WHEN NOT l\.granted"), Sql()); + + /// + /// Ungranted sorts first. A granted lock is not a finding — every working server holds thousands — so a + /// grid ordered any other way buries the only rows worth reading. + /// + [Fact] + public void UngrantedSortsFirst() + => Assert.Matches(new Regex(@"ORDER BY l\.granted"), Sql()); + + /// + /// Catalog reads are schema-qualified: pg_catalog is searched implicitly but not necessarily + /// FIRST, so an unqualified read can resolve to an object a user created in a schema earlier in the + /// monitoring login's search_path. + /// + [Fact] + public void EveryCatalogRead_IsSchemaQualified() + { + /* Comments stripped first: the query explains the cluster-wide/per-database split in prose that + necessarily names the catalogs, and an unstripped scan matches the explanation rather than the + code. This repo has hit that trap repeatedly. */ + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + foreach (var view in new[] { "pg_locks", "pg_database", "pg_class", "pg_stat_activity" }) + { + foreach (Match match in Regex.Matches(sql, $@"(\S*)\b{Regex.Escape(view)}\b")) + { + Assert.Equal("pg_catalog.", match.Groups[1].Value); + } + } + } + + [Fact] + public void AppliesTo_EveryPostgresTarget() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + Assert.True(PgLockStatsCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + } + } + + /// One SELECT alias per payload column, in order — a mismatch is a silently shifted binary + /// COPY, which writes every value into the wrong column rather than failing. + [Fact] + public void SelectAliases_MatchThePayloadOrder() + { + var expected = PgLockStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + var selected = Sql() + .Split('\n') + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal) + && !line.TrimStart().StartsWith("LEFT JOIN", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(expected, selected); + } +} diff --git a/Lite.Tests/PgPlanCaptureCollectorDefinitionTests.cs b/Lite.Tests/PgPlanCaptureCollectorDefinitionTests.cs new file mode 100644 index 000000000..10a31db38 --- /dev/null +++ b/Lite.Tests/PgPlanCaptureCollectorDefinitionTests.cs @@ -0,0 +1,278 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.Json.Nodes; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins plan capture, and above all the redaction — this is the one collector that reads text a customer +/// wrote, so a regression here publishes their data rather than merely reporting a wrong number. +/// +public class PgPlanCaptureCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext() + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + PostgresVersionNum = 170000, + }, + }; + + private static string Sql => PgPlanCaptureCollector.Instance.BuildQuery(MakeContext()).Text; + + /// + /// Through the PUBLIC parser entry point rather than reflection into a private method. The redaction + /// is shared with the RDS log-API transport (#2538), so testing it through the seam both callers use + /// is the only way this guard covers both of them. + /// + private static JsonObject Redact(string json) + { + var parsed = PgPlanLogParser.FromBlock(1, 1.0, json); + Assert.NotNull(parsed); + return (JsonObject)JsonNode.Parse(parsed!.Value.PlanJson)!; + } + + [Fact] + public void Identity_IsTheTableAndEngineTheStoreExpects() + { + Assert.Equal("pg_plan_capture", PgPlanCaptureCollector.Instance.Name); + Assert.Equal("pg_plan_capture", PgPlanCaptureCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgPlanCaptureCollector.Instance.TargetEngine); + } + + /// + /// Query text never becomes a column. auto_explain emits it verbatim — literals and all — and + /// log_parameter_max_length = 0 does not suppress it, because that setting covers bind parameters + /// only (measured on #2565). The statement identity is query_id. + /// + [Fact] + public void ThereIsNoQueryTextColumn() + { + var names = PgPlanCaptureCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.DoesNotContain("query_text", names); + Assert.DoesNotContain("statement", names); + Assert.Contains("query_id", names); + } + + /// + /// The header is dropped before anything else touches the tree, and the value that was in it must not + /// survive anywhere else either — it also appears inside Filter. + /// + [Fact] + public void QueryTextAndItsLiterals_AreBothRemoved() + { + var root = Redact(""" + { + "Query Text": "SELECT 1 FROM accounts WHERE email = 'someone@example.com';", + "Plan": { "Node Type": "Seq Scan", "Relation Name": "accounts", + "Filter": "(email = 'someone@example.com'::text)" } + } + """); + + var json = root.ToJsonString(); + + Assert.Null(root["Query Text"]); + Assert.DoesNotContain("someone@example.com", json, StringComparison.Ordinal); + Assert.DoesNotContain("SELECT 1 FROM accounts", json, StringComparison.Ordinal); + } + + /// + /// Redaction must not mangle IDENTITY. A blanket numeric strip would rewrite a relation genuinely named + /// transactionitems1 into transactionitems?, destroying the name to hide a value that was + /// never there — so bare numbers are stripped only inside condition fields. + /// + [Fact] + public void RedactionStripsValues_WithoutDestroyingNames() + { + var root = Redact(""" + { + "Plan": { "Node Type": "Seq Scan", "Relation Name": "transactionitems1", "Alias": "t1", + "Plan Rows": 100, "Filter": "((id > 100) AND (v = 'secret'::text))" } + } + """); + + var plan = (JsonObject)root["Plan"]!; + + Assert.Equal("transactionitems1", plan["Relation Name"]!.GetValue()); + Assert.Equal("t1", plan["Alias"]!.GetValue()); + + var filter = plan["Filter"]!.GetValue(); + Assert.DoesNotContain("100", filter, StringComparison.Ordinal); + Assert.DoesNotContain("secret", filter, StringComparison.Ordinal); + + /* Numeric JSON values are estimates, not customer data, and stay untouched. */ + Assert.Equal(100, plan["Plan Rows"]!.GetValue()); + } + + /// + /// Nested plans and string arrays carry values too — Output and the various key lists are arrays, + /// and a child node's Filter is where most real literals live. + /// + [Fact] + public void RedactionReachesNestedPlansAndArrays() + { + var root = Redact(""" + { + "Plan": { "Node Type": "Limit", "Output": ["id", "'inline-literal'"], + "Plans": [ { "Node Type": "Seq Scan", + "Filter": "(name = 'nested-secret'::text)" } ] } + } + """); + + var json = root.ToJsonString(); + + Assert.DoesNotContain("nested-secret", json, StringComparison.Ordinal); + Assert.DoesNotContain("inline-literal", json, StringComparison.Ordinal); + } + + /// + /// The log read is BOUNDED. #2565 measured 772 MB of log in twenty seconds at capture-everything, and a + /// collector that read the file whole would become the server's largest reader. + /// + [Fact] + public void TheLogReadIsBounded() + { + Assert.Contains("greatest(n.size -", Sql, StringComparison.Ordinal); + Assert.Contains("pg_catalog.pg_read_file", Sql, StringComparison.Ordinal); + + /* The current file is discovered, not configured: log_filename is a strftime pattern. */ + Assert.Contains("pg_catalog.pg_ls_logdir()", Sql, StringComparison.Ordinal); + } + + /// + /// Server-wide, and it claims no database attribution: the log line prefix is not guaranteed to carry + /// %d, so a database column would be a claim the source cannot support (#2599). + /// + [Fact] + public void ItClaimsNoDatabaseAttribution() + { + Assert.False(PgPlanCaptureCollector.Instance.RunsPerDatabase( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + + Assert.DoesNotContain("database_name", + PgPlanCaptureCollector.Instance.PayloadColumns.Select(c => c.Name)); + } +} + +/// +/// The parser is shared by both transports — pg_read_file against a self-hosted log, and the RDS +/// DownloadDBLogFilePortion API for managed PostgreSQL, which has no filesystem (#2538). These pin +/// the entry point the API transport uses, which no collector test reaches. +/// +/// The redaction living in one place is the point. Every other divergence in this codebase has cost a +/// wrong number; this one would cost a customer's data. +/// +public class PgPlanLogParserTests +{ + /* Built with explicit \t escapes rather than a raw string literal. auto_explain indents the JSON + block with real TABS and the parser keys on them, but inside a raw literal \t is two characters + rather than one - so a raw-literal fixture silently stops looking like a log and the extraction + finds nothing. Caught by this test failing 1 != 2 after passing in a scratch harness that used + escaped strings. */ + private const string RawLog = + "2026-08-25 14:50:05.299 UTC [58] -3560200806914842915 LOG: duration: 0.006 ms plan:\n" + + "\t{\n" + + "\t \"Query Text\": \"SELECT 1 FROM accounts WHERE email = 'someone@example.com';\",\n" + + "\t \"Plan\": {\n" + + "\t \"Node Type\": \"Seq Scan\",\n" + + "\t \"Relation Name\": \"accounts1\",\n" + + "\t \"Filter\": \"((email = 'someone@example.com'::text) AND (id > 100))\"\n" + + "\t }\n" + + "\t}\n" + + "2026-08-25 14:50:05.300 UTC [48] 0 LOG: received fast shutdown request\n" + + "2026-08-25 14:50:11.111 UTC [81] 510393640047350727 LOG: duration: 11.877 ms plan:\n" + + "\t{\n" + + "\t \"Query Text\": \"SELECT 2;\",\n" + + "\t \"Plan\": { \"Node Type\": \"Result\" }\n" + + "\t}\n"; + + [Fact] + public void ExtractPullsEveryPlanAndIgnoresOrdinaryLogLines() + { + var plans = PgPlanLogParser.Extract(RawLog); + + Assert.Equal(2, plans.Count); + Assert.Equal(-3560200806914842915, plans[0].QueryId); + Assert.Equal(510393640047350727, plans[1].QueryId); + Assert.Equal(0.006, plans[0].DurationMs, 3); + + /* The shutdown line between them is not a plan and must not become one. */ + Assert.DoesNotContain(plans, p => p.TopNodeType is null); + } + + [Fact] + public void ExtractRedactsEveryPlanItReturns() + { + var plans = PgPlanLogParser.Extract(RawLog); + + foreach (var plan in plans) + { + Assert.DoesNotContain("Query Text", plan.PlanJson, StringComparison.Ordinal); + Assert.DoesNotContain("someone@example.com", plan.PlanJson, StringComparison.Ordinal); + Assert.DoesNotMatch(new Regex(@"'(?!\?')[^']+'"), plan.PlanJson); + } + + /* And the identity survives: a relation named with a trailing digit is not a redacted number. */ + Assert.Contains("accounts1", plans[0].PlanJson, StringComparison.Ordinal); + } + + /// + /// The hash is of the REDACTED plan, so the same shape recurs to the same hash whatever values it ran + /// with — which is the whole basis of dedup. Hashing raw text would defeat it exactly where it matters. + /// + [Fact] + public void TheSameShapeWithDifferentValues_HashesTheSame() + { + var a = PgPlanLogParser.FromBlock(1, 1.0, + """{"Plan":{"Node Type":"Seq Scan","Relation Name":"t","Filter":"(id = 1)"}}"""); + var b = PgPlanLogParser.FromBlock(1, 9.0, + """{"Plan":{"Node Type":"Seq Scan","Relation Name":"t","Filter":"(id = 99999)"}}"""); + + Assert.NotNull(a); + Assert.NotNull(b); + Assert.Equal(a!.Value.PlanHash, b!.Value.PlanHash); + + /* A genuinely different SHAPE must not collide with it. */ + var c = PgPlanLogParser.FromBlock(1, 1.0, + """{"Plan":{"Node Type":"Index Scan","Relation Name":"t","Filter":"(id = 1)"}}"""); + Assert.NotEqual(a.Value.PlanHash, c!.Value.PlanHash); + } + + /// + /// Both transports read a BOUNDED window, so a block cut in half at the edge is ordinary rather than + /// exceptional. It is skipped, never stored half-parsed, and never throws. + /// + [Fact] + public void ATruncatedBlockIsSkippedRatherThanThrowing() + { + var truncated = "2026-08-25 14:50:05.299 UTC [58] 123 LOG: duration: 0.006 ms plan:\n\t{\n\t \"Plan\": {\n"; + + var plans = PgPlanLogParser.Extract(truncated); + + Assert.Empty(plans); + Assert.Null(PgPlanLogParser.FromBlock(1, 1.0, "{not json")); + Assert.Null(PgPlanLogParser.FromBlock(1, 1.0, null)); + } +} diff --git a/Lite.Tests/PgPlanCaptureReadinessCollectorDefinitionTests.cs b/Lite.Tests/PgPlanCaptureReadinessCollectorDefinitionTests.cs new file mode 100644 index 000000000..321f2893a --- /dev/null +++ b/Lite.Tests/PgPlanCaptureReadinessCollectorDefinitionTests.cs @@ -0,0 +1,352 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2564: the plan-capture readiness collector. The assertions are about the DISTINCTIONS it has to draw, +/// not its wording — a test that pinned the sentences would fail on every improvement to them and say +/// nothing about the only property that matters, which is that a reader can tell which remedy applies. +/// +public class PgPlanCaptureReadinessCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext(int major = 16) + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 24, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + }; + + /// + /// The store table must not collide with a pg_catalog object: pg_catalog is searched first, so a + /// colliding name breaks CREATE INDEX with 42809 and makes unqualified reads resolve to the MONITORING + /// store's own copy. + /// + [Fact] + public void NameAndTargetTable_AreTheContract_AndShadowNoCatalogObject() + { + Assert.Equal("pg_plan_capture_readiness", PgPlanCaptureReadinessCollector.Instance.Name); + Assert.Equal("pg_plan_capture_readiness", PgPlanCaptureReadinessCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgPlanCaptureReadinessCollector.Instance.TargetEngine); + Assert.NotEqual("pg_settings", PgPlanCaptureReadinessCollector.Instance.TargetTable); + Assert.NotEqual("pg_available_extensions", PgPlanCaptureReadinessCollector.Instance.TargetTable); + } + + /// + /// Every PostgreSQL target including standbys. A replica can carry a different parameter group from its + /// writer, and gating this to writers would hide exactly that divergence. + /// + [Theory] + [InlineData(13)] + [InlineData(16)] + [InlineData(17)] + public void AppliesTo_EveryPostgresTarget(int major) + { + var target = new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major }; + Assert.True(PgPlanCaptureReadinessCollector.Instance.AppliesTo(target)); + Assert.True(CollectorCatalog.AppliesTo(PgPlanCaptureReadinessCollector.Instance, target)); + } + + /// And never at a SQL Server target — the composed gate stops the dialect being sent at all. + [Fact] + public void AppliesTo_NeverASqlServerTarget() + => Assert.False(CollectorCatalog.AppliesTo( + PgPlanCaptureReadinessCollector.Instance, + new CollectorTargetInfo { Engine = CollectorTargetEngine.SqlServer, SqlMajorVersion = 16 })); + + /// + /// Every current_setting must use the two-argument MISSING_OK form. Reading a GUC that does not + /// exist because the library was never loaded is the NORMAL case for this collector, and the + /// one-argument form raises 42704 — which would turn its entire purpose into an error every cycle. + /// + [Fact] + public void EveryCurrentSetting_UsesTheMissingOkForm() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + var calls = Regex.Matches(sql, @"current_setting\(([^)]*)\)"); + Assert.NotEmpty(calls); + foreach (Match call in calls) + { + Assert.Contains(", true", call.Groups[1].Value, StringComparison.Ordinal); + } + } + + /// + /// The GUC value is never cast to a number. current_setting renders some settings WITH their + /// unit, so a cast is how a collector starts failing on one major version and not another — the reason + /// observed is text and interpretation happens downstream. + /// + [Fact] + public void TheObservedValue_IsNeverCastToANumber() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.DoesNotMatch(new Regex(@"current_setting\([^)]*\)\s*::\s*(int|integer|bigint|numeric|double)"), sql); + Assert.DoesNotMatch(new Regex(@"pg_size_bytes\s*\(\s*current_setting"), sql); + } + + /// + /// Every facet is a separate row, and that IS the feature. Collapsing them would produce the one thing + /// this collector exists to prevent: a single "plans unavailable" that tells nobody what to do, when the + /// remedies are a parameter-group change plus a reboot, a single setting, and a log_line_prefix edit + /// respectively. + /// + [Fact] + public void EveryFacet_IsASeparateRow() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + var facets = new[] { "library_loaded", "capture_threshold", "extension_available", "plan_text_setting", "plan_attribution" }; + foreach (var facet in facets) + { + Assert.Contains($"'{facet}'::text", sql, StringComparison.Ordinal); + } + + /* DERIVED from the roster above, not spelled as a literal. This assertion was `Assert.Equal(3, ...)` + and adding a fifth facet left it stale — the roster was updated, the count was not, and a guard + whose whole job is to catch a collapsed facet failed for arithmetic instead. N facets are joined by + N-1 UNION ALLs by construction, so saying that is both the real invariant and un-staleable. */ + Assert.Equal(facets.Length - 1, Regex.Matches(sql, @"\bUNION ALL\b").Count); + } + + /// + /// The trap the issue was filed about: log_min_duration = -1 is loaded-and-capturing-nothing, + /// which from outside looks identical to not-loaded but has a completely different remedy. The query + /// must treat it as UNSATISFIED and must distinguish it from the setting being absent entirely. + /// + [Fact] + public void ACaptureThresholdOfMinusOne_IsNotSatisfied_AndIsDistinctFromAbsent() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + /* -1 fails the satisfied test rather than passing as "a value is set". */ + Assert.Contains("<> '-1'", sql, StringComparison.Ordinal); + + /* And an absent GUC reports as absent rather than being folded into the -1 case, so the reader can + tell "loaded but switched off" from "never loaded". */ + Assert.Contains("(setting absent - library not loaded)", sql, StringComparison.Ordinal); + Assert.Contains("IS NOT NULL", sql, StringComparison.Ordinal); + } + + /// + /// It must not claim plans ARE captured. Library loaded, capture configured, and plans being readable by + /// us are three separate facts (#2566/#2567 own the last one), and conflating them is how a readiness + /// probe starts lying about a capability. + /// + [Fact] + public void ItReportsReadiness_NeverThatPlansAreBeingCaptured() + { + var columns = PgPlanCaptureReadinessCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("is_satisfied", columns); + Assert.DoesNotContain("plans_captured", columns); + Assert.DoesNotContain("plan_count", columns); + Assert.DoesNotContain("plan_xml", columns); + } + + /// + /// The remedy travels with the observation. A read that reconstructed it later would drift from what the + /// collector actually saw, and the remedy is specific to both the facet and the platform. + /// + [Fact] + public void EachFacet_CarriesItsOwnRemedy() + { + var columns = PgPlanCaptureReadinessCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + Assert.Contains("detail", columns); + Assert.Contains("observed", columns); + + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + /* The Aurora/RDS instruction specifically — it is a CLUSTER parameter group and a writer reboot, not + a SET, and that misconception is the one worth heading off. */ + Assert.Contains("CLUSTER", sql, StringComparison.Ordinal); + Assert.Contains("REBOOT", sql, StringComparison.Ordinal); + } + + /// + /// The remedy must depend on what was OBSERVED, not be one fixed sentence per facet. Measured against a + /// live PostgreSQL 17: with the threshold at 0 the first draft still said "log_min_duration = -1 means + /// ... capturing NOTHING" beside is_satisfied = true, so the row contradicted itself. Three of the + /// four states that facet reaches were being given the wrong advice. + /// + [Fact] + public void TheCaptureThresholdRemedy_DependsOnWhatWasObserved() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + /* A CASE over the observed value, not a literal. Anchored on the threshold COLUMN rather than on an + inline current_setting call: the settings are now read once into a CTE (#2605), so matching the + call text here pinned the plumbing rather than the behaviour and broke on a refactor that kept + every branch intact. */ + Assert.Contains("CASE", sql, StringComparison.Ordinal); + Assert.Contains("WHEN p.threshold IS NULL", sql, StringComparison.Ordinal); + Assert.Contains("WHEN p.threshold = '-1'", sql, StringComparison.Ordinal); + Assert.Contains("WHEN p.threshold = '0'", sql, StringComparison.Ordinal); + + /* The state that made this facet wrong on a live target: a threshold set with no library behind it + (#2605). It must reach its own branch and not fall through to the satisfied wording. */ + Assert.Contains("WHEN NOT p.loaded AND p.threshold IS NOT NULL", sql, StringComparison.Ordinal); + + /* And the satisfied arm must not repeat the -1 sentence, which is the exact contradiction found. */ + var elseArm = sql[sql.IndexOf("ELSE 'auto_explain is loaded and capturing", StringComparison.Ordinal)..]; + Assert.DoesNotContain("capturing NOTHING", elseArm[..200], StringComparison.Ordinal); + } + + /// + /// observed is stored as the server's own rendering, and this is why: measured on PostgreSQL 17, + /// auto_explain.log_min_duration = 250 comes back as 250ms — the GUC carries its + /// unit. A numeric column or a cast would raise invalid-input-syntax and fail the whole collection, which + /// is the track_activity_query_size defect this codebase has already paid for once. + /// + [Fact] + public void TheObservedColumn_IsText_BecauseGucsRenderWithTheirUnit() + { + var observed = PgPlanCaptureReadinessCollector.Instance.PayloadColumns.Single(c => c.Name == "observed"); + Assert.Equal(CollectorColumnType.Varchar, observed.Type); + } + + /// + /// The library check is boundary-aware, not a substring. shared_preload_libraries is a + /// comma-separated list, and LIKE '%auto_explain%' reports true for any library whose name merely + /// CONTAINS it — measured against PostgreSQL 17, both my_auto_explain_shim and + /// auto_explain_extra false-positived under the substring form and are correctly rejected under + /// the boundary form. This facet's whole job is to be trustworthy about loaded-versus-not. + /// + [Fact] + public void TheLibraryCheck_IsBoundaryAware_NotASubstringMatch() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.DoesNotContain("LIKE '%auto_explain%'", sql, StringComparison.Ordinal); + Assert.Contains(@"~ '(^|,)\s*auto_explain\s*(,|$)'", sql, StringComparison.Ordinal); + } + + /// + /// Catalog reads are schema-qualified, matching every other PostgreSQL collector here. pg_catalog is + /// searched implicitly but not necessarily FIRST, so an unqualified read can resolve to an object a user + /// created in a schema earlier in the monitoring login's search_path — which for this collector would + /// mean fabricating an answer about whether plan capture is possible at all. + /// + [Fact] + public void CatalogReads_AreSchemaQualified() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.DoesNotMatch(new Regex(@"FROM\s+pg_available_extensions"), sql); + Assert.Contains("pg_catalog.pg_available_extensions", sql, StringComparison.Ordinal); + } + + /// Every output column is aliased — an unaliased expression comes back named after the + /// function, which makes the query undebuggable in psql, the one tool anyone reaches for. + [Fact] + public void EveryOutputColumn_IsAliased() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + var payload = PgPlanCaptureReadinessCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + /* The first SELECT carries the aliases; the UNION ALL branches inherit them by position, which is + how PostgreSQL names a union's columns. */ + foreach (var name in payload) + { + Assert.Contains($"AS {name}", sql, StringComparison.Ordinal); + } + } + + /// + /// The shared_preload_libraries test is computed EXACTLY ONCE, and every facet that judges an + /// auto_explain.* setting reads that one value. + /// + /// This is the assertion that would have caught the shipped bug. A GUC named + /// auto_explain.* can exist with no auto_explain behind it — PostgreSQL accepts any qualified + /// custom variable, so a parameter group or a plain SET defines one on a server that never + /// loaded the module. capture_threshold inferred "loaded" from the GUC merely existing, and on a + /// live Aurora target reported SATISFIED with the words "auto_explain is loaded and capturing at this + /// threshold" while library_loaded in the same result set said false. Reproduced on PostgreSQL 17 + /// with one SET. + /// + /// The root cause was one fact derived in five places, and the fifth forgot. Pinning the count at + /// one is what makes a sixth copy impossible rather than merely discouraged — the same shape as #2599, + /// where two collectors derived "is this extension installed" from different sources and disagreed. + /// + [Fact] + public void PreloadDetection_IsComputedOnce_AndGatesEveryAutoExplainFacet() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + var preloadReads = Regex.Matches(sql, @"current_setting\(\s*'shared_preload_libraries'").Count; + + Assert.True( + preloadReads == 1, + $"shared_preload_libraries is read {preloadReads} times. It must be read once and shared, or two " + + "facets will eventually disagree about whether auto_explain is loaded."); + + /* Every auto_explain.* GUC is read inside that same single row, so no facet can reach one without + also having the loaded flag beside it. */ + foreach (Match read in Regex.Matches(sql, @"current_setting\(\s*'auto_explain\.[a-z_]+'")) + { + var before = sql.Substring(0, read.Index); + Assert.True( + before.LastIndexOf("UNION ALL", StringComparison.Ordinal) < 0, + $"{read.Value} is read inside a facet branch rather than the shared row, so that facet can " + + "judge it without consulting whether the library is loaded."); + } + } + + /// + /// A setting that exists while the library does not is called out as a placeholder rather than reported + /// as working configuration. The wording matters more than usual here: the reader's next action is + /// different for "not set" (set it) than for "set but inert" (load the library first, then this starts + /// mattering), and the old text sent them to the wrong one. + /// + [Fact] + public void AnUnloadedLibrary_ReportsItsSettingsAsPlaceholders() + { + var sql = PgPlanCaptureReadinessCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("PLACEHOLDER", sql, StringComparison.Ordinal); + + /* The claim that must never be reachable without the loaded flag. */ + var claim = "auto_explain is loaded and capturing at this threshold"; + Assert.Contains(claim, sql, StringComparison.Ordinal); + + /* Scoped to the facet's is_satisfied EXPRESSION - between the facet literal and the `observed` + column that follows it - rather than to the whole branch. Searching the branch for the word + "loaded" passes on the buggy version too, because its prose says "library not loaded": the + assertion has to look where the decision is made, not where the word appears. */ + var branchStart = sql.IndexOf("'capture_threshold'::text", StringComparison.Ordinal); + Assert.True(branchStart >= 0, "the capture_threshold facet is missing"); + + var observedAt = sql.IndexOf("coalesce(p.threshold", branchStart, StringComparison.Ordinal); + Assert.True(observedAt > branchStart, "capture_threshold's observed column moved; rescope this test"); + + var satisfiedExpression = sql[branchStart..observedAt]; + + Assert.True( + satisfiedExpression.Contains("p.loaded", StringComparison.Ordinal), + "capture_threshold decides is_satisfied without consulting the shared loaded flag, so a " + + "placeholder auto_explain.* GUC on a server with no auto_explain reports as working capture."); + } +} diff --git a/Lite.Tests/PgPredicateStatsCollectorDefinitionTests.cs b/Lite.Tests/PgPredicateStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..d04abd269 --- /dev/null +++ b/Lite.Tests/PgPredicateStatsCollectorDefinitionTests.cs @@ -0,0 +1,122 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the predicate collector on the scope error that would corrupt it and the sampling fact that would +/// make its numbers lie. +/// +public class PgPredicateStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext() + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + PostgresVersionNum = 170000, + }, + }; + + private static string Sql => PgPredicateStatsCollector.Instance.BuildQuery(MakeContext()).Text; + + [Fact] + public void Identity_IsTheTableAndEngineTheStoreExpects() + { + Assert.Equal("pg_predicate_stats", PgPredicateStatsCollector.Instance.Name); + Assert.Equal("pg_predicate_stats", PgPredicateStatsCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgPredicateStatsCollector.Instance.TargetEngine); + } + + /// + /// The assertion that matters most. pg_qualstats() returns cluster-wide rows keyed by + /// dbid, but lrelid/lattnum are OIDs meaningful only inside their own database and + /// pg_class/pg_attribute are per-database catalogs. Without the scope filter the join + /// either silently drops other databases' rows or resolves them against whatever local object shares + /// the OID and reports a confident wrong column name. Measured: two cross-database rows present, neither + /// colliding locally that day — so the failure would have been silent loss, and different OID luck would + /// have produced the wrong name instead. + /// + [Fact] + public void ItScopesToTheConnectedDatabase_AndRunsPerDatabase() + { + Assert.Contains("q.dbid = (SELECT oid FROM pg_catalog.pg_database WHERE datname = current_database())", + Sql, StringComparison.Ordinal); + + Assert.True(PgPredicateStatsCollector.Instance.RunsPerDatabase( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + + /* Per-database, so the rows must be attributable — the #2599 invariant. */ + Assert.Contains("database_name", PgPredicateStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + } + + /// + /// The sample rate is stored, because the counts are a sample and the default is not 1. + /// pg_qualstats.sample_rate defaults to 1/max_connections — 0.01 on the rig, where the + /// function returned ZERO rows on a server that had just run the queries it was meant to record. + /// + [Fact] + public void TheSampleRate_IsStoredWithTheCounts() + { + Assert.Contains("pg_qualstats.sample_rate", Sql, StringComparison.Ordinal); + Assert.Contains("sample_rate", PgPredicateStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + + /* Missing-ok, like every GUC read here: an absent setting degrades one column, not the collection. */ + foreach (Match call in Regex.Matches(Sql, @"current_setting\(([^)]*)\)")) + { + Assert.Contains(", true", call.Groups[1].Value, StringComparison.Ordinal); + } + } + + /// + /// Ranked by rows filtered — the index-candidate signal. Ranking by sample count would order the grid + /// by how often the SAMPLER happened to fire, which is a property of the extension's configuration + /// rather than of the workload. + /// + [Fact] + public void ItRanksByRowsFiltered_NotBySampleCount() + { + Assert.Matches(new Regex(@"ORDER BY\s+sum\(q\.nbfiltered\) DESC"), Sql); + } + + /// + /// An extension function lives where the extension was created, not in pg_catalog. + /// + [Fact] + public void TheExtensionFunction_IsQualifiedWhereItActuallyLives() + { + Assert.Contains("public.pg_qualstats()", Sql, StringComparison.Ordinal); + Assert.DoesNotContain("pg_catalog.pg_qualstats", Sql, StringComparison.Ordinal); + } + + /// + /// The operator is stored as its SYMBOL. An OID would need a second lookup to mean anything, in a + /// catalog that is per-database — and which operator it is decides whether an index would help at all. + /// + [Fact] + public void TheOperatorIsStoredAsASymbol() + { + Assert.Contains("operator", PgPredicateStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + Assert.Contains("o.oprname", Sql, StringComparison.Ordinal); + } +} diff --git a/Lite.Tests/PgReplicationStatsCollectorDefinitionTests.cs b/Lite.Tests/PgReplicationStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..54443338f --- /dev/null +++ b/Lite.Tests/PgReplicationStatsCollectorDefinitionTests.cs @@ -0,0 +1,173 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2544, the replication slice. The assertions are about the two things measurement showed matter: that +/// BOTH the byte distance and the time lag are captured, and that every distance is measured from the +/// primary's current WAL position rather than from what was sent. +/// +public class PgReplicationStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static string Sql() + => PgReplicationStatsCollector.Instance.BuildQuery(new CollectorContext + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + }, + ExcludedDatabases = Array.Empty(), + }).Text; + + /// + /// Both measures, because they disagree in the case that matters. Measured against a standby holding + /// pg_wal_replay_pause(): 33.7 MB behind while the time lag read 2.8 seconds. Storing the lag + /// alone would make a stalled replica look survivable. + /// + [Fact] + public void BothTheByteDistanceAndTheTimeLag_AreCollected() + { + var names = PgReplicationStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("replay_bytes_behind", names); + Assert.Contains("replay_lag_ms", names); + } + + /// + /// Four distances, not one. In the stalled run sent, write and flush were all ZERO + /// while replay was 33.7 MB behind — so the fault was purely apply, which no single column shows. + /// + [Fact] + public void AllFourStages_AreMeasuredSeparately() + { + var names = PgReplicationStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + foreach (var stage in new[] { "sent_bytes_behind", "write_bytes_behind", "flush_bytes_behind", "replay_bytes_behind" }) + { + Assert.Contains(stage, names); + } + } + + /// + /// Distance is measured from the primary's CURRENT WAL position, never from sent_lsn. Using what + /// was sent as the baseline hides a sender that has itself fallen behind — the distance would read zero + /// while the standby was arbitrarily far from the truth. + /// + [Fact] + public void DistanceIsMeasuredFromCurrentWalLsn_NotFromSentLsn() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + /* Every pg_wal_lsn_diff must take pg_current_wal_lsn() as its FIRST argument. */ + var diffs = Regex.Matches(sql, @"pg_wal_lsn_diff\(\s*([^,]+),").Cast().ToArray(); + + Assert.NotEmpty(diffs); + Assert.All(diffs, m => Assert.Contains("pg_current_wal_lsn()", m.Groups[1].Value, StringComparison.Ordinal)); + } + + /// + /// The lag columns are PostgreSQL intervals and must be converted at the source. Storing an + /// interval would force every consumer to parse it. + /// + [Fact] + public void LagIntervals_AreConvertedToMilliseconds() + => Assert.Equal(3, Regex.Matches(Sql(), @"EXTRACT\(EPOCH FROM r\.\w+_lag\) \* 1000").Count); + + /// + /// sync_state is collected because a SYNC standby falling behind blocks commits on the primary. + /// Identical lag, entirely different severity from an async one's. + /// + [Fact] + public void SyncStateIsCollected_BecauseItChangesTheSeverity() + => Assert.Contains("sync_state", PgReplicationStatsCollector.Instance.PayloadColumns.Select(c => c.Name)); + + /// + /// Applies to every target INCLUDING standbys: a cascading replica's downstream is as worth watching as + /// a primary's, and recovery state changes on failover without a dispatch gate noticing. A standby with + /// no downstream returns zero rows, which is correct rather than an error. + /// + [Fact] + public void AppliesTo_EveryPostgresTarget_IncludingStandbys() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + Assert.True(PgReplicationStatsCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + } + } + + /// + /// backend_start is converted with AT TIME ZONE 'UTC' and never a bare cast — the cast + /// renders in the SESSION's TimeZone, so a store session east of UTC would record a stamp hours from the + /// one the server meant. + /// + [Fact] + public void BackendStart_IsConvertedToUtc_NotBareCast() + { + var sql = Sql(); + + Assert.Contains("AT TIME ZONE 'UTC'", sql, StringComparison.Ordinal); + Assert.DoesNotMatch(new Regex(@"backend_start\s*::\s*timestamp"), sql); + } + + /// + /// Catalog reads are schema-qualified: pg_catalog is searched implicitly but not necessarily + /// FIRST, so an unqualified read can resolve to an object a user created in a schema earlier in the + /// monitoring login's search_path. + /// + [Fact] + public void EveryCatalogRead_IsSchemaQualified() + { + var sql = Regex.Replace(Sql(), @"/\*.*?\*/", " ", RegexOptions.Singleline); + + foreach (var name in new[] { "pg_stat_replication", "pg_wal_lsn_diff", "pg_current_wal_lsn" }) + { + foreach (Match match in Regex.Matches(sql, $@"(\S*)\b{Regex.Escape(name)}\b")) + { + /* ENDS WITH, not equals. These calls NEST — pg_catalog.pg_wal_lsn_diff(pg_catalog. + pg_current_wal_lsn(), ...) — so the greedy \S* captures the enclosing call as well as the + qualification. Asserting equality fails on correctly-qualified SQL, which is what the + first draft of this test did. */ + Assert.EndsWith("pg_catalog.", match.Groups[1].Value, StringComparison.Ordinal); + } + } + } + + /// One SELECT alias per payload column, in order — a mismatch is a silently shifted binary + /// COPY, which writes every value into the wrong column rather than failing. + [Fact] + public void SelectAliases_MatchThePayloadOrder() + { + var expected = PgReplicationStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + var selected = Sql() + .Split('\n') + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(expected, selected); + } +} diff --git a/Lite.Tests/PgSessionStatesCollectorDefinitionTests.cs b/Lite.Tests/PgSessionStatesCollectorDefinitionTests.cs new file mode 100644 index 000000000..48a3a8fa1 --- /dev/null +++ b/Lite.Tests/PgSessionStatesCollectorDefinitionTests.cs @@ -0,0 +1,616 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using System.Threading; +using System.Threading.Tasks; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the session-states collector (#2540): that no raw statement text can reach the store, that the +/// horizon sentinel survives the reader instead of being floored into a claim, that the version floor is +/// handled by substitution rather than by gating the collector off, and that the query text cannot re-enter +/// through the command-tag column. +/// +public class PgSessionStatesCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext(int major = 16, ICollectorDeltaCalculator? deltas = null) + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 23, 12, 0, 0, DateTimeKind.Utc), + Deltas = deltas ?? s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + }; + + /// + /// The table name cannot be changed later without a migration, and pg_catalog is searched before + /// search_path — so a store table named after a catalog object breaks CREATE INDEX with 42809 and makes + /// unqualified reads resolve to the MONITORING store's own copy. + /// + [Fact] + public void Identity_Pinned_AndTheTableDoesNotShadowACatalogObject() + { + Assert.Equal("pg_session_states", PgSessionStatesCollector.Instance.Name); + Assert.Equal("pg_session_states", PgSessionStatesCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgSessionStatesCollector.Instance.TargetEngine); + + Assert.NotEqual("pg_stat_activity", PgSessionStatesCollector.Instance.TargetTable); + Assert.NotEqual("pg_stat_get_activity", PgSessionStatesCollector.Instance.TargetTable); + } + + /// + /// Runs on ANY PostgreSQL target including standbys, deliberately unlike pg_autovacuum_stats. + /// A standby holds its own transactions exactly like a primary, and with hot_standby_feedback on + /// its xmin propagates to the PRIMARY — which is one of the four causes pg_xmin_horizon attributes. A + /// reflex IsInRecovery gate would blind this collector to the place that cause is created. + /// PostgreSQL 13 is the product's documented floor and is included: query_id is the only column + /// this reads that 13 lacks, and that is handled by substitution below rather than by refusing the + /// target. + /// + [Theory] + [InlineData(13, false, false)] + [InlineData(13, false, true)] + [InlineData(16, true, false)] + [InlineData(17, true, true)] + [InlineData(18, false, true)] + public void AppliesToEveryPostgresTarget_IncludingStandbysAndTheVersionFloor( + int major, bool isAurora, bool inRecovery) + { + var target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + IsAurora = isAurora, + IsInRecovery = inRecovery, + }; + + Assert.True(PgSessionStatesCollector.Instance.AppliesTo(target)); + Assert.True(CollectorCatalog.AppliesTo(PgSessionStatesCollector.Instance, target)); + } + + /// + /// The engine half of the gate. A PostgreSQL definition dispatched at a SQL Server target would send + /// this dialect at T-SQL every cycle — the #2213 class of defect — so the composed gate has to say no + /// even though AppliesTo on its own says yes. + /// + [Fact] + public void TheComposedGateRefusesASqlServerTarget() + { + var sqlServer = new CollectorTargetInfo { Engine = CollectorTargetEngine.SqlServer, SqlMajorVersion = 16 }; + + Assert.False(CollectorCatalog.AppliesTo(PgSessionStatesCollector.Instance, sqlServer)); + } + + /// + /// query_id arrived in PostgreSQL 14. Below that the column does not exist and naming it is a parse + /// error that would take the whole collection down every cycle, so it is substituted with a TYPED null — + /// the row shape has to stay constant across a mixed-version fleet. + /// Confirmed by reading pg_attribute for the view on a live PostgreSQL 13.23 instance: 21 columns, + /// including leader_pid, backend_type, backend_xid and backend_xmin, and no query_id. + /// + [Theory] + [InlineData(13)] + [InlineData(12)] + public void BelowPostgres14_QueryIdIsSubstitutedWithATypedNull(int major) + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext(major)).Text; + + Assert.Contains("NULL::bigint", sql, StringComparison.Ordinal); + Assert.DoesNotContain("a.query_id", sql, StringComparison.Ordinal); + } + + [Theory] + [InlineData(14)] + [InlineData(16)] + [InlineData(18)] + public void OnPostgres14AndAbove_QueryIdIsRead(int major) + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext(major)).Text; + + Assert.Contains("a.query_id", sql, StringComparison.Ordinal); + } + + /// + /// No column of the query is projected, on any version. pg_stat_activity.query is the statement as + /// submitted, with literal parameter values inline — verified on a live target, where a probe session's + /// row came back carrying its literal argument verbatim. + /// This is the assertion that stops the column being added back by someone reasoning from + /// pg_blocking, which does store it. That collector fires on an exceptional condition where the text IS + /// the finding; this one fires on a duration floor an ordinary application crosses, so the same column + /// would mean routinely accumulating user data to answer a question that does not need it. + /// + [Fact] + public void NoRawQueryTextIsEverProjectedOrStored() + { + foreach (var major in new[] { 13, 14, 16, 17, 18 }) + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext(major)).Text; + + /* The reference is allowed in the redaction test and in the command-tag whitelist, where the + text is COMPARED but never emitted. What must not appear is a projection of it. + + A REGEX with a trailing word boundary, not DoesNotContain("AS query"): the plain + substring also matches "AS query_id", which is the column this collector deliberately + DOES emit, so the loose form failed against correct code. It failed in the safe + direction - a guard that cries wolf costs a CI round, where one that stays quiet costs + a leak - but a guard that cannot tell the thing it forbids from the thing it requires + is not yet a guard. \b after "query" needs a non-word character next, and "_" is a word + character, so "AS query_id" no longer matches while "AS query," still does. */ + Assert.DoesNotMatch(new Regex(@"\bAS\s+query\b"), sql); + Assert.DoesNotContain("a.query AS", sql, StringComparison.Ordinal); + Assert.DoesNotContain("left(a.query", sql, StringComparison.OrdinalIgnoreCase); + Assert.DoesNotContain("substring(a.query", sql, StringComparison.OrdinalIgnoreCase); + Assert.DoesNotContain("substr(a.query", sql, StringComparison.OrdinalIgnoreCase); + } + + Assert.DoesNotContain( + PgSessionStatesCollector.Instance.PayloadColumns, + c => c.Name.Contains("query_text", StringComparison.Ordinal) || c.Name == "query"); + } + + /// + /// The command tag is a CLOSED whitelist, not a substring, and that is a data-protection decision rather + /// than a formatting one: plenty of ORMs prepend a SQL comment block, so the first token of a real + /// statement can be a comment opener followed by anything the application chose to put in it, and a + /// leading-N-characters rule would carry literals straight out of a WHERE clause. + /// + [Fact] + public void TheCommandTagIsAWhitelistAndFallsBackToAConstant() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("split_part(btrim(a.query), ' ', 1)", sql, StringComparison.Ordinal); + Assert.Contains("'(other)'", sql, StringComparison.Ordinal); + Assert.Contains("'(redacted)'", sql, StringComparison.Ordinal); + Assert.Contains("'(idle)'", sql, StringComparison.Ordinal); + + /* The whitelist itself, spot-checked at both ends so a truncation of the list fails. */ + Assert.Contains("'SELECT'", sql, StringComparison.Ordinal); + Assert.Contains("'UPDATE'", sql, StringComparison.Ordinal); + Assert.Contains("'COMMIT'", sql, StringComparison.Ordinal); + Assert.Contains("'TABLE'", sql, StringComparison.Ordinal); + } + + /// + /// PostgreSQL block comments NEST, so a literal comment opener written inside a comment in this query + /// opens a nested one that never closes and the whole collection fails to parse — every cycle, on every + /// target. The first draft did exactly that and it was caught only by running the shipped string against + /// a live instance; nothing about reading the C# reveals it. + /// + [Fact] + public void TheQueryHasBalancedBlockComments() + { + foreach (var major in new[] { 13, 16 }) + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext(major)).Text; + + var depth = 0; + for (var i = 0; i < sql.Length - 1; i++) + { + if (sql[i] == '/' && sql[i + 1] == '*') + { + depth++; + i++; + } + else if (sql[i] == '*' && sql[i + 1] == '/') + { + depth--; + i++; + Assert.True(depth >= 0, $"unbalanced comment close on PostgreSQL {major}"); + } + } + + Assert.Equal(0, depth); + } + } + + /// + /// The collector's own backend is excluded, or it is a PERMANENT row: Darling's read sits in + /// pg_stat_activity with an open transaction and a backend_xmin like anything else, so the + /// zero-rows-when-healthy state becomes unreachable and every denominator is padded by one. + /// Parallel workers are excluded too, and the guard is deliberately not a bare NULL check — + /// leader_pid is documented NULL for a plain leader but is set to the process's OWN pid for a leader + /// participating in its parallel group on newer majors, so excluding on NULL alone would drop those + /// leaders entirely. + /// + [Fact] + public void TheCollectorsOwnBackendAndParallelWorkersAreExcluded() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("a.pid <> pg_backend_pid()", sql, StringComparison.Ordinal); + Assert.Contains("a.leader_pid IS NULL OR a.leader_pid = a.pid", sql, StringComparison.Ordinal); + } + + /// + /// The redaction flag is derived from the insufficient-privilege literal and NOT from a NULL state. + /// Measured under full privilege on a live target: checkpointer, walwriter and the autovacuum + /// launcher all report a NULL state and a NULL query, so testing state alone would flag a perfectly + /// healthy instance as unprivileged. The literal is unambiguous. + /// + [Fact] + public void RedactionIsDetectedFromThePrivilegeLiteralNotFromANullState() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("''", sql, StringComparison.Ordinal); + Assert.Contains("AS state_is_redacted", sql, StringComparison.Ordinal); + Assert.DoesNotContain("a.state IS NULL AS state_is_redacted", sql, StringComparison.Ordinal); + } + + /// + /// The horizon is read from BOTH backend_xmin and backend_xid, and from the GREATER of the two ages. + /// Neither column alone sees both cases. Measured on live PostgreSQL 16.15: a READ COMMITTED + /// transaction that has written holds backend_xid with backend_xmin NULL, while a REPEATABLE READ + /// transaction holds backend_xmin with backend_xid NULL. Reading only one makes the collector blind to + /// half of what pins the horizon. + /// age(), not arithmetic on the raw xid: modular wrap makes naive subtraction wrong exactly at the + /// boundary where it matters most. + /// + [Fact] + public void TheHorizonAgeReadsBothColumnsAndTakesTheGreater() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("age(a.backend_xmin)", sql, StringComparison.Ordinal); + Assert.Contains("age(a.backend_xid)", sql, StringComparison.Ordinal); + Assert.Contains("GREATEST(", sql, StringComparison.Ordinal); + Assert.Contains("THEN -1::bigint", sql, StringComparison.Ordinal); + } + + /// + /// Every duration is computed SERVER-side in milliseconds, so there is no timestamp column in this table + /// at all. timestamptz::text renders in the SESSION TimeZone and is byte-identical to UTC on a UTC + /// server, which is why that mistake survives every probe; shipping a duration removes the class rather + /// than guarding against it. + /// + [Fact] + public void DurationsAreServerComputed_AndNoTimestampIsShipped() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("clock_timestamp()", sql, StringComparison.Ordinal); + Assert.DoesNotContain("::text AS state_change", sql, StringComparison.Ordinal); + Assert.DoesNotContain("AS xact_start", sql, StringComparison.Ordinal); + + Assert.DoesNotContain( + PgSessionStatesCollector.Instance.PayloadColumns, + c => c.Type == CollectorColumnType.Timestamp); + } + + /// + /// The row cap bounds a pathological instance, and the pre-limit count travels with the rows so a + /// truncated capture is self-evident rather than silently under-reported. + /// + [Fact] + public void TheCaptureIsBoundedAndCarriesItsOwnPreLimitCount() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("LIMIT 100", sql, StringComparison.Ordinal); + Assert.Contains("count(*) OVER ()", sql, StringComparison.Ordinal); + Assert.Contains("AS reportable_sessions", sql, StringComparison.Ordinal); + } + + /// + /// The oldest holder is force-included regardless of the duration floor. That is the entire causal claim + /// this collector makes, and losing it to a floor would leave the reader exactly where pg_xmin_horizon + /// already left them — knowing a session holds the horizon and not which. + /// The winner is computed over the FULL activity set, not the reportable subset: a young + /// transaction can be the oldest holder, and crowning a filtered survivor would name the wrong session + /// precisely when the real one fell below the floor. + /// + [Fact] + public void TheOldestHolderIsIncludedRegardlessOfTheDurationFloor() + { + var sql = PgSessionStatesCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("max(horizon_age) FILTER (WHERE horizon_age >= 0)", sql, StringComparison.Ordinal); + Assert.Contains("WHERE r.is_horizon_holder", sql, StringComparison.Ordinal); + Assert.Contains("FROM activity", sql, StringComparison.Ordinal); + } + + /// + /// PayloadColumns order IS the wire format: WritePayload is positional, so a column added without a + /// matching .Value() shifts everything after it and stores data that is silently WRONG rather than + /// failing. Both the order and the count are pinned. + /// + [Fact] + public void PayloadColumns_AreInOrder_AndMatchTheRowArity() + { + var expected = new[] + { + "backend_id", "pid", "database_name", "username", "application_name", "client_addr", + "backend_type", "state", "wait_event_type", "wait_event", "command_tag", "query_id", + "state_duration_ms", "xact_duration_ms", "query_duration_ms", "backend_duration_ms", + "xmin_age", "xid_age", "horizon_age", "is_idle_in_transaction", "is_horizon_holder", + "state_is_redacted", "total_sessions", "active_sessions", "idle_in_transaction_sessions", + "reportable_sessions", + }; + + Assert.Equal(expected, PgSessionStatesCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray()); + + var writer = new RecordingCollectorRowWriter(); + PgSessionStatesCollector.Instance.WritePayload(SampleRow(), writer, MakeContext()); + + Assert.Equal(PgSessionStatesCollector.Instance.PayloadColumns.Count, writer.Values.Count); + } + + /// + /// query_id is BigInt and NULLABLE, with three separate meanings the read has to keep apart: PostgreSQL + /// 13 has no such column, 14+ reports NULL when compute_query_id is off, and a redacted row reports NULL + /// along with everything else privileged. A sentinel would collapse all three into a number, and 0 is a + /// legal query_id. + /// + [Fact] + public void QueryIdIsNullableRatherThanSentinelled() + { + var column = Assert.Single( + PgSessionStatesCollector.Instance.PayloadColumns.Where(c => c.Name == "query_id")); + Assert.Equal(CollectorColumnType.BigInt, column.Type); + + var writer = new RecordingCollectorRowWriter(); + PgSessionStatesCollector.Instance.WritePayload(SampleRow() with { QueryId = null }, writer, MakeContext()); + + Assert.Null(writer.Values[11]); + } + + /// + /// Every field mapped to its own ordinal, with deliberately distinct values so a transposed pair fails + /// rather than passing on two equal numbers. + /// + [Fact] + public async Task ReadsAFullyPopulatedRow_WithEveryFieldOnItsOwnOrdinal() + { + var reader = new FakeCollectorDataReader( + new object[] + { + 17_874_796_750_069_283L, 69_283, // backend_id, pid + "appdb", "app_user", "checkout-worker", "10.0.0.7", + "client backend", "idle in transaction", "Client", "ClientRead", + "UPDATE", -5_564_491_789_055_112_251L, // command_tag, query_id + 584_357L, 584_358L, 584_359L, 600_000L, // state/xact/query/backend durations + -1L, 1_401L, 1_401L, // xmin_age, xid_age, horizon_age + true, true, false, // iit, holder, redacted + 9, 2, 4, 4, // totals + }); + + var rows = await PgSessionStatesCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal(17_874_796_750_069_283L, row.BackendId); + Assert.Equal(69_283, row.Pid); + Assert.Equal("appdb", row.DatabaseName); + Assert.Equal("app_user", row.Username); + Assert.Equal("checkout-worker", row.ApplicationName); + Assert.Equal("10.0.0.7", row.ClientAddr); + Assert.Equal("client backend", row.BackendType); + Assert.Equal("idle in transaction", row.State); + Assert.Equal("Client", row.WaitEventType); + Assert.Equal("ClientRead", row.WaitEvent); + Assert.Equal("UPDATE", row.CommandTag); + Assert.Equal(-5_564_491_789_055_112_251L, row.QueryId); + Assert.Equal(584_357L, row.StateDurationMs); + Assert.Equal(584_358L, row.XactDurationMs); + Assert.Equal(584_359L, row.QueryDurationMs); + Assert.Equal(600_000L, row.BackendDurationMs); + Assert.Equal(-1L, row.XminAge); + Assert.Equal(1_401L, row.XidAge); + Assert.Equal(1_401L, row.HorizonAge); + Assert.True(row.IsIdleInTransaction); + Assert.True(row.IsHorizonHolder); + Assert.False(row.StateIsRedacted); + Assert.Equal(9, row.TotalSessions); + Assert.Equal(2, row.ActiveSessions); + Assert.Equal(4, row.IdleInTransactionSessions); + Assert.Equal(4, row.ReportableSessions); + } + + /// + /// The redacted shape, exactly as PostgreSQL returns it to a role without pg_monitor — measured, not + /// imagined. The row is NOT refused: every privileged column comes back NULL while backend_xmin and + /// backend_xid stay visible, so the horizon still reads as pinned and nothing can say by what. + /// The reader must not turn those NULLs into zeros. -1 for a duration is visibly not a + /// measurement; 0 would read as "this transaction started this instant". + /// + [Fact] + public async Task ARedactedRow_KeepsItsSentinels_AndTheHorizonAgeSurvives() + { + var reader = new FakeCollectorDataReader( + new object[] + { + 17_874_378_340_069_283L, 69_283, + "appdb", "app_user", "checkout-worker", DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, + "(redacted)", DBNull.Value, + DBNull.Value, DBNull.Value, DBNull.Value, DBNull.Value, + -1L, 900L, 900L, + false, true, true, + 9, 0, 0, 2, + }); + + var row = Assert.Single( + await PgSessionStatesCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None)); + + Assert.True(row.StateIsRedacted); + Assert.Null(row.State); + Assert.Null(row.BackendType); + Assert.Null(row.QueryId); + Assert.Equal("(redacted)", row.CommandTag); + + Assert.Equal(-1L, row.StateDurationMs); + Assert.Equal(-1L, row.XactDurationMs); + Assert.Equal(-1L, row.QueryDurationMs); + Assert.Equal(-1L, row.BackendDurationMs); + + /* The cruel part of the redaction, and the reason this row is stored rather than dropped: the xid + age is NOT redacted, so the horizon is still visibly pinned. */ + Assert.Equal(900L, row.HorizonAge); + Assert.True(row.IsHorizonHolder); + } + + /// + /// A horizon age of -1 must survive the reader. This is the sentinel the entire feature turns on: an + /// idle-in-transaction session that holds neither a snapshot nor a transaction id pins NOTHING, and 0 + /// would read as "holds the newest possible xid" — the opposite finding, and the one that talks somebody + /// into killing a harmless session. + /// + [Fact] + public async Task IdleInTransactionPinningNothing_KeepsMinusOne_NotZero() + { + var reader = new FakeCollectorDataReader( + new object[] + { + 17_874_796_540_069_234L, 69_234, + "appdb", "app_user", "reporting", DBNull.Value, + "client backend", "idle in transaction", "Client", "ClientRead", + "SELECT", 7_184_301_683_933_573_861L, + 605_419L, 605_420L, 605_421L, 620_000L, + -1L, -1L, -1L, + true, false, false, + 9, 0, 4, 4, + }); + + var row = Assert.Single( + await PgSessionStatesCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None)); + + Assert.True(row.IsIdleInTransaction); + Assert.False(row.IsHorizonHolder); + Assert.Equal(-1L, row.XminAge); + Assert.Equal(-1L, row.XidAge); + Assert.Equal(-1L, row.HorizonAge); + Assert.NotEqual(0L, row.HorizonAge); + + /* Ten minutes idle inside a transaction and pinning nothing at all. Both halves of that sentence + have to survive into the store or the read cannot make the distinction. */ + Assert.True(row.StateDurationMs > 600_000L); + } + + /// + /// Zero rows is the HEALTHY state — no session had a transaction open past the floor and nothing held + /// the horizon — and must never read as a failure or be padded with a placeholder. + /// + [Fact] + public async Task AnEmptyResultIsHealthy_NotAFailure() + { + var rows = await PgSessionStatesCollector.Instance.ReadAsync( + new FakeCollectorDataReader(), MakeContext(), CancellationToken.None); + + Assert.Empty(rows); + } + + /// + /// No deltas. Every column here is a LEVEL measured at the instant of the sample; differencing two + /// samples would produce the elapsed time between them, which is a property of the schedule rather than + /// of the session. + /// + [Fact] + public async Task TakesNoDeltas() + { + var deltas = new RecordingCollectorDeltaCalculator(); + var reader = new FakeCollectorDataReader( + new object[] + { + 1L, 1, "appdb", "u", "a", DBNull.Value, "client backend", "idle in transaction", + DBNull.Value, DBNull.Value, "SELECT", DBNull.Value, + 1L, 2L, 3L, 4L, -1L, 5L, 5L, true, true, false, 1, 0, 1, 1, + }); + + var rows = await PgSessionStatesCollector.Instance.ReadAsync( + reader, MakeContext(deltas: deltas), CancellationToken.None); + + var writer = new RecordingCollectorRowWriter(); + PgSessionStatesCollector.Instance.WritePayload(rows[0], writer, MakeContext(deltas: deltas)); + + Assert.Empty(deltas.Calls); + } + + private static PgSessionStatesCollector.Row SampleRow() => new( + BackendId: 17_874_796_750_069_283, + Pid: 69_283, + DatabaseName: "appdb", + Username: "app_user", + ApplicationName: "checkout-worker", + ClientAddr: "10.0.0.7", + BackendType: "client backend", + State: "idle in transaction", + WaitEventType: "Client", + WaitEvent: "ClientRead", + CommandTag: "UPDATE", + QueryId: -5_564_491_789_055_112_251, + StateDurationMs: 584_357, + XactDurationMs: 584_358, + QueryDurationMs: 584_359, + BackendDurationMs: 600_000, + XminAge: -1, + XidAge: 1_401, + HorizonAge: 1_401, + IsIdleInTransaction: true, + IsHorizonHolder: true, + StateIsRedacted: false, + TotalSessions: 9, + ActiveSessions: 2, + IdleInTransactionSessions: 4, + ReportableSessions: 4); + + /// + /// ONE MINUTE, matching pg_blocking for the same reason rather than by copying it: both read + /// pg_stat_activity, both are SAMPLES of a view that records nothing on its own, and the cadence IS the + /// resolution. This one is the cheaper of the two — no pg_blocking_pids() call, no lock-manager + /// ShareLock, no per-database fan-out. + /// 30 days, also matching pg_blocking: the question is whether this application is parking + /// transactions more than it used to, and a month covers a release cycle, which is the unit at which + /// anyone can act on the answer. + /// + [Fact] + public void RegisteredInBothTheCatalogAndTheSchedule() + { + Assert.Contains(CollectorCatalog.All, d => d.Name == "pg_session_states"); + + var schedule = CollectorScheduleDefaults.All["pg_session_states"]; + + Assert.Equal(1, schedule.FrequencyMinutes); + Assert.Equal(30, schedule.RetentionDays); + Assert.True(schedule.DefaultEnabled); + + /* Asserted against the sibling rather than as a second literal, so the shared sampling argument + cannot drift into two unrelated numbers. */ + Assert.Equal(CollectorScheduleDefaults.All["pg_blocking"].FrequencyMinutes, schedule.FrequencyMinutes); + Assert.Equal(CollectorScheduleDefaults.All["pg_blocking"].RetentionDays, schedule.RetentionDays); + } + + /// + /// The capability vocabulary has a noun phrase for this collector, so a SQL Server target asked + /// get_pg_session_states is told what it does not collect rather than getting the generic fallback. + /// + [Fact] + public void TheCapabilityMessageNamesWhatIsNotCollected() + { + var message = CollectorEngineCapability.NotCollectedMessage( + "sql-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.SqlServer, "pg_session_states"); + + Assert.NotNull(message); + Assert.Contains("idle in transaction", message, StringComparison.Ordinal); + Assert.DoesNotContain("the data this read is served from", message, StringComparison.Ordinal); + } +} diff --git a/Lite.Tests/PgStatementStatsFlavorTests.cs b/Lite.Tests/PgStatementStatsFlavorTests.cs new file mode 100644 index 000000000..5e1c39b31 --- /dev/null +++ b/Lite.Tests/PgStatementStatsFlavorTests.cs @@ -0,0 +1,177 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2625: pg_statement_stats reads TWO sources — Aurora's extended function, or the vanilla +/// pg_stat_statements view on any other PostgreSQL — and which one is a BuildQuery decision, +/// never an applicability one. +/// +/// +/// The gate it replaced cost more than a few columns. AppliesTo means permanent incapability, and the +/// capability machinery composes a deliberately final sentence from it: "does not collect per-query-shape +/// execution statistics, and never will". Every operator of a non-Aurora PostgreSQL target got that sentence +/// for the single question a database monitor exists to answer — while pg_kernel_stats collected OS +/// CPU and pg_predicate_stats collected selectivity, both keyed by the very queryids nothing was +/// identifying. It survived because no self-hosted PostgreSQL target existed to notice. +/// +/// +/// +/// The ordinals are the load-bearing detail. Both queries select the same 27 columns in the same order — the +/// vanilla one fills Aurora's six with typed NULL literals — so ReadAsync, PayloadColumns and +/// WritePayload stay single implementations. A shorter vanilla SELECT would have meant a second reader +/// whose ordinals could drift from this one, which is exactly the failure the per-major column naming in this +/// collector already exists to prevent. +/// +/// +public class PgStatementStatsFlavorTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext(bool isAurora, int major = 17) + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 26, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + IsAurora = isAurora, + PostgresMajorVersion = major, + PostgresVersionNum = major * 10000, + }, + }; + + private static string Sql(bool isAurora, int major = 17) + => PgStatementStatsCollector.Instance.BuildQuery(MakeContext(isAurora, major)).Text; + + /// + /// The whole point. This assertion is the one that would have caught the gap, and it could not have been + /// written before a non-Aurora target existed to ask the question of. + /// + [Theory] + [InlineData(true)] + [InlineData(false)] + public void ItAppliesToEveryPostgresTarget_AuroraOrNot(bool isAurora) + => Assert.True(PgStatementStatsCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, IsAurora = isAurora })); + + [Fact] + public void AuroraReadsTheExtendedFunction() + { + var sql = Sql(isAurora: true); + + Assert.Contains("FROM aurora_stat_statements(false)", sql, StringComparison.Ordinal); + Assert.DoesNotContain("pg_stat_statements", sql, StringComparison.Ordinal); + } + + [Fact] + public void EveryOtherPostgresReadsTheVanillaView() + { + var sql = Sql(isAurora: false); + + Assert.Contains("FROM public.pg_stat_statements", sql, StringComparison.Ordinal); + Assert.DoesNotContain("aurora_stat_statements", sql, StringComparison.Ordinal); + } + + /// + /// One reader, one column list, one write order — so the two queries must agree on shape, not just on + /// source. Counted at the top level so a comma inside a cast or a comment cannot inflate it. + /// + [Fact] + public void BothFlavorsSelectTheSameColumnsInTheSameOrder() + { + var aurora = SelectAliases(Sql(isAurora: true)); + var vanilla = SelectAliases(Sql(isAurora: false)); + + Assert.Equal(aurora, vanilla); + Assert.Equal(PgStatementStatsCollector.Instance.PayloadColumns.Count - 3, aurora.Count); + } + + /// + /// NULL, not 0, and typed — an untyped NULL literal would arrive as text and Npgsql's strict type + /// checking would throw on GetInt64, which is the same class of defect as the ordinal drift above. + /// + [Theory] + [InlineData("storage_blks_read", "bigint")] + [InlineData("orcache_blks_hit", "bigint")] + [InlineData("storage_blk_read_time", "double precision")] + [InlineData("orcache_blk_read_time", "double precision")] + [InlineData("total_exec_peakmem", "bigint")] + [InlineData("max_exec_peakmem", "bigint")] + public void TheAuroraOnlyColumnsAreTypedNullsOnTheVanillaPath(string alias, string type) + { + Assert.Matches(new Regex($@"NULL::{Regex.Escape(type)}\s+AS {Regex.Escape(alias)}\b"), Sql(isAurora: false)); + } + + /// + /// toplevel arrived in pg_stat_statements 1.9 (PostgreSQL 14). Before that, nested tracking + /// did not exist, so every row IS a top-level statement — true is the CORRECT value on an older + /// server, not a fallback, and the delta key stays four-part on every version. + /// + [Theory] + [InlineData(13, "true")] + [InlineData(14, "toplevel")] + [InlineData(17, "toplevel")] + public void ToplevelIsGuardedForPostgresBefore14(int major, string expected) + { + Assert.Matches(new Regex($@"{Regex.Escape(expected)}\s+AS toplevel\b"), Sql(isAurora: false, major)); + } + + /// + /// The per-major block-time naming applies to the vanilla view too — it is a pg_stat_statements + /// rename, not an Aurora one, and getting it wrong would silently shift every ordinal after it. + /// + [Theory] + [InlineData(16, "blk_read_time", "blk_write_time")] + [InlineData(17, "shared_blk_read_time", "shared_blk_write_time")] + public void TheBlockTimeColumnsFollowTheMajorVersionOnBothFlavors(int major, string read, string write) + { + foreach (var isAurora in new[] { true, false }) + { + var sql = Sql(isAurora, major); + + Assert.Matches(new Regex($@"(? + /// The vanilla query must not reach for anything Aurora-only by accident — a stray reference would fail + /// at parse time on every non-Aurora target, which is a total outage of the read rather than a missing + /// column. + /// + [Fact] + public void TheVanillaQueryTouchesNoAuroraSurface() + => Assert.DoesNotContain("aurora_", Sql(isAurora: false), StringComparison.OrdinalIgnoreCase); + + /// + /// Column aliases of the outermost SELECT, in order. Comments are stripped first so a comma inside one + /// cannot be counted — the same correction the probe-arity guard needed. + /// + private static System.Collections.Generic.List SelectAliases(string sql) + { + var body = Regex.Replace(sql, @"/\*.*?\*/", " ", RegexOptions.Singleline); + var start = body.IndexOf("SELECT", StringComparison.Ordinal) + "SELECT".Length; + var end = body.IndexOf("FROM ", start, StringComparison.Ordinal); + + return Regex.Matches(body[start..end], @"AS\s+([a-z_]+)\s*(?:,|$)", RegexOptions.IgnoreCase) + .Select(m => m.Groups[1].Value) + .ToList(); + } +} diff --git a/Lite.Tests/PgTableBloatStatsCollectorDefinitionTests.cs b/Lite.Tests/PgTableBloatStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..9ff67b4a0 --- /dev/null +++ b/Lite.Tests/PgTableBloatStatsCollectorDefinitionTests.cs @@ -0,0 +1,591 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Threading; +using System.Threading.Tasks; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the per-table bloat collector (#2542): that the headline number is stored as an ESTIMATE and never +/// as a measurement, that the trust signals which decide whether it may be published all travel with it, and +/// that the never-analyzed sentinel survives the reader instead of being floored into a claim. +/// +public class PgTableBloatStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext( + int major = 17, ICollectorDeltaCalculator? deltas = null, string? currentDatabase = "appdb") + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 20, 12, 0, 0, DateTimeKind.Utc), + Deltas = deltas ?? s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + CurrentDatabaseName = currentDatabase, + }; + + /// + /// The table name cannot be changed later without a migration, and pg_catalog is searched before + /// search_path — so a store table named after a catalog object breaks CREATE INDEX with 42809 and makes + /// unqualified reads resolve to the MONITORING store's own copy. + /// + [Fact] + public void Identity_Pinned_AndTheTableDoesNotShadowACatalogObject() + { + Assert.Equal("pg_table_bloat_stats", PgTableBloatStatsCollector.Instance.Name); + Assert.Equal("pg_table_bloat_stats", PgTableBloatStatsCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgTableBloatStatsCollector.Instance.TargetEngine); + + Assert.NotEqual("pg_stats", PgTableBloatStatsCollector.Instance.TargetTable); + Assert.NotEqual("pg_stat_user_tables", PgTableBloatStatsCollector.Instance.TargetTable); + Assert.NotEqual("pg_class", PgTableBloatStatsCollector.Instance.TargetTable); + } + + /// + /// WRITERS ONLY, on every major and both Aurora and stock — the gate reads IsInRecovery and + /// nothing else. + /// + /// The bloat ARITHMETIC would still be right on a replica, because it reads replicated catalog + /// rows. What would be wrong is mods_since_analyze and last_analyzed, which read as + /// "statistics are perfectly fresh" on a replica that has never analyzed anything — silently zeroing the + /// one signal that detects the estimator's 81-percentage-point failure mode. + /// + [Theory] + [InlineData(13, false, false, true)] + [InlineData(15, false, false, true)] + [InlineData(16, true, false, true)] + [InlineData(17, true, false, true)] + [InlineData(18, false, false, true)] + [InlineData(13, false, true, false)] + [InlineData(16, true, true, false)] + [InlineData(17, true, true, false)] + [InlineData(18, false, true, false)] + public void AppliesToWritersOnly_OnEveryMajorAndBothAuroraAndStock( + int major, bool isAurora, bool inRecovery, bool expected) + { + Assert.Equal(expected, PgTableBloatStatsCollector.Instance.AppliesTo(new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + IsAurora = isAurora, + IsInRecovery = inRecovery, + })); + } + + /// + /// The composed gate still refuses a SQL Server target. The collector's own AppliesTo only asks about + /// recovery, so the ENGINE half is the only thing keeping this PostgreSQL query text off a SQL Server + /// connection. + /// + [Fact] + public void TheEngineHalfOfTheDispatchGateStillRefusesASqlServerTarget() + { + Assert.False(CollectorCatalog.AppliesTo(PgTableBloatStatsCollector.Instance, new CollectorTargetInfo())); + Assert.True(CollectorCatalog.AppliesTo( + PgTableBloatStatsCollector.Instance, + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql })); + } + + /// + /// pg_stats and pg_stat_user_tables are scoped to the connected database, so this is + /// necessarily a fan-out — sharing pg_autovacuum_stats's hourly cadence so the CAUSE and the DAMAGE line + /// up sample for sample. + /// + [Fact] + public void RunsPerDatabase() + { + Assert.True(PgTableBloatStatsCollector.Instance.RunsPerDatabase(MakeContext().Target)); + } + + /// + /// NO version branch at all, and that is a MEASURED result rather than an omission: the whole query was + /// executed against live PostgreSQL 13, 14, 15, 16, 17 and 18 and returned the same shape and the same + /// numbers for the same fixture. Every catalog column it reads was confirmed present on all six by + /// listing the live catalogs, so there is nothing to gate on and a gate would be a source of drift. + /// + [Fact] + public void TheQueryIsOneConstantWithNoVersionBranch() + { + var majors = new[] { 13, 14, 15, 16, 17, 18 } + .Select(m => PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext(m)).Text) + .Distinct(StringComparer.Ordinal) + .ToArray(); + + Assert.Single(majors); + + /* Same REFERENCE across two majors, not merely equal text: that is what proves the query is one + stored constant rather than a string rebuilt per call, so there is no interpolation site where a + version branch could later be introduced unnoticed. */ + Assert.Same( + PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext(13)).Text, + PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext(18)).Text); + } + + /// + /// estimate_unavailable is LOAD-BEARING, not advisory: TRUE means the estimate has no basis, and + /// the read suppresses the number rather than captioning it. + /// + /// All three conditions that set it must be present. The middle one is the important one in + /// production: pg_stats is filtered by has_column_privilege(...'select') and + /// pg_monitor does NOT confer SELECT on user tables, so a correctly-provisioned monitoring login + /// sees ZERO pg_stats rows — and MEASURED against exactly such a role on a live target, the estimator + /// did not fail: it returned a confident 88.59% for a table whose true bloat is 0.50%. + /// + [Theory] + [InlineData("bool_or(att.atttypid = 'pg_catalog.name'::regtype)")] + [InlineData("count(sts.attname) <> count(att.attname)")] + [InlineData("MAX(tbl.reltuples) < 0")] + public void CarriesEveryConditionThatSetsEstimateUnavailable(string condition) + { + Assert.Contains(condition, PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + /// + /// The trust signals travel with the estimate. mods_since_analyze is the one that catches the + /// estimator's worst failure: MEASURED on two byte-identical 8,998-page tables with an identical true + /// bloat of 10.93%, the one whose column-width statistics predated a widening UPDATE estimated 92.64% + /// and the freshly-analyzed one estimated 11.01% — 81 percentage points apart, with nothing in the + /// arithmetic to show it. Width statistics can only go stale THROUGH modifications, which is what makes + /// this a sound proxy rather than a coincidence. + /// + [Theory] + [InlineData("mods_since_analyze")] + [InlineData("last_analyzed")] + [InlineData("estimate_unavailable")] + [InlineData("alignment_bytes")] + [InlineData("pgstattuple_available")] + [InlineData("estimated_tuple_bytes")] + [InlineData("estimated_heap_pages")] + [InlineData("fillfactor")] + public void CarriesEveryTrustSignalTheEstimateNeeds(string column) + { + Assert.Contains(column, PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + /// + /// MAXALIGN cannot be read from a GUC and is detected from version(). The alternation carries + /// aarch64 and arm64 explicitly because Graviton RDS and Aurora instances are aarch64 — + /// their version() string does also contain "64-bit", but relying on that one token to hold across every + /// vendor's build is a bet with no upside. The value actually used is STORED, so a platform where the + /// detection is wrong shows up in the data rather than skewing every estimate on it silently. + /// + [Fact] + public void DetectsMaxAlignIncludingOnArm() + { + var sql = PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains("version() ~ '64-bit|x86_64|ppc64|ia64|amd64|aarch64|arm64'", sql, StringComparison.Ordinal); + Assert.Contains("AS alignment_bytes", sql, StringComparison.Ordinal); + } + + /// + /// The MEASURED columns, which are true whatever the statistics or the grants look like — and the reason + /// a suppressed estimate degrades the answer rather than removing it. + /// + [Theory] + [InlineData("pg_relation_size(tbl.oid)")] + [InlineData("pg_relation_size(tbl.reltoastrelid)")] + [InlineData("pg_indexes_size(tbl.oid)")] + [InlineData("MAX(st.n_dead_tup)")] + public void SelectsTheMeasuredColumnsThatSurviveAPermissionsGap(string fragment) + { + Assert.Contains(fragment, PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + /// + /// The size floor is the fan-out cost control. It is stated in BYTES from pg_relation_size rather + /// than in relpages, because relpages is only refreshed by VACUUM or ANALYZE and would let a table + /// that has grown since its last maintenance fall through the filter it most needs to pass. + /// + [Fact] + public void FiltersOnTheHeapSizeFloor_MeasuredNotFromTheCatalogPageCount() + { + Assert.Contains( + "WHERE s.heap_bytes >= 1048576", + PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text, + StringComparison.Ordinal); + } + + /// + /// timestamptz::text renders in the SESSION's TimeZone and the store contract is naive UTC, so the + /// conversion is explicit. GREATEST ignoring NULLs is wanted here rather than a hazard: whichever of the + /// manual and automatic analyze actually happened is the later non-NULL one, and a table that has only + /// ever been autoanalyzed must not report "never analyzed". + /// + [Fact] + public void ConvertsTheAnalyzeTimestampToUtcExplicitly() + { + var sql = PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text; + + Assert.Contains( + "GREATEST(MAX(st.last_analyze), MAX(st.last_autoanalyze)) AT TIME ZONE 'UTC'", + sql, + StringComparison.Ordinal); + + Assert.DoesNotContain("last_analyze::timestamp", sql, StringComparison.Ordinal); + Assert.DoesNotContain("last_autoanalyze::timestamp", sql, StringComparison.Ordinal); + } + + /// + /// Every catalog reference is schema-qualified. pg_catalog is searched implicitly and FIRST, so an + /// unqualified reference is a hostage to whatever a future search_path or a user object named after a + /// catalog view does to it. + /// + [Theory] + [InlineData("pg_catalog.pg_attribute")] + [InlineData("pg_catalog.pg_class")] + [InlineData("pg_catalog.pg_namespace")] + [InlineData("pg_catalog.pg_stat_user_tables")] + [InlineData("pg_catalog.pg_stats")] + [InlineData("pg_catalog.pg_extension")] + public void SchemaQualifiesEveryCatalogReference(string reference) + { + Assert.Contains(reference, PgTableBloatStatsCollector.Instance.BuildQuery(MakeContext()).Text, StringComparison.Ordinal); + } + + /// + /// The estimate is named _estimate in the STORE, not only in the read's prose, so the qualifier + /// cannot be lost between the store and a screen. A bloat number that is quietly wrong will be used to + /// justify a VACUUM FULL on a production table. + /// + [Fact] + public void TheEstimateIsNamedAsAnEstimateInThePayload() + { + var names = PgTableBloatStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("bloat_bytes_estimate", names); + Assert.Contains("bloat_pct_estimate", names); + + /* And no unqualified name that would read as a measurement. */ + Assert.DoesNotContain("bloat_bytes", names); + Assert.DoesNotContain("bloat_pct", names); + } + + [Fact] + public void PayloadColumns_OrderAndKeyTypes_Pinned() + { + var columns = PgTableBloatStatsCollector.Instance.PayloadColumns; + + Assert.Equal(19, columns.Count); + Assert.Equal( + new[] + { + "database_name", "schema_name", "table_name", + "heap_bytes", "heap_pages", "toast_bytes", "index_bytes", + "live_tuples", "dead_tuples", "mods_since_analyze", "last_analyzed", + "estimated_tuple_bytes", "estimated_heap_pages", "fillfactor", + "bloat_bytes_estimate", "bloat_pct_estimate", + "estimate_unavailable", "alignment_bytes", "pgstattuple_available", + }, + columns.Select(c => c.Name).ToArray()); + + Assert.Equal(CollectorColumnType.Varchar, columns[0].Type); // database_name + Assert.Equal(CollectorColumnType.BigInt, columns[3].Type); // heap_bytes + Assert.Equal(CollectorColumnType.Timestamp, columns[10].Type); // last_analyzed + Assert.Equal(CollectorColumnType.Double, columns[11].Type); // estimated_tuple_bytes + Assert.Equal(CollectorColumnType.Integer, columns[13].Type); // fillfactor + Assert.Equal(CollectorColumnType.BigInt, columns[14].Type); // bloat_bytes_estimate + Assert.Equal(CollectorColumnType.Boolean, columns[16].Type); // estimate_unavailable + Assert.Equal(CollectorColumnType.Integer, columns[17].Type); // alignment_bytes + Assert.Equal(CollectorColumnType.Boolean, columns[18].Type); // pgstattuple_available + + /* The percentage is numeric(5,2), which is what makes "94.28" storable and "94.283" not. */ + var pct = columns[15]; + Assert.Equal(CollectorColumnType.Decimal, pct.Type); + Assert.Equal(5, pct.Precision); + Assert.Equal(2, pct.Scale); + } + + /// + /// Every field mapped to its own ordinal, with deliberately distinct values so a transposed pair fails + /// rather than passing on two equal numbers. + /// + [Fact] + public async Task ReadsAFullyPopulatedRow_WithEveryFieldOnItsOwnOrdinal() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "churny", + 56_516_608L, 6_899L, // heap_bytes, heap_pages + 24_000_000L, 6_758_400L, // toast_bytes, index_bytes + 100_000L, 500L, 1_200L, // live_tuples, dead_tuples, mods_since_analyze + new DateTime(2026, 8, 19, 3, 15, 0, DateTimeKind.Unspecified), + 436.0d, 1_715L, 70, // estimated_tuple_bytes, estimated_heap_pages, fillfactor + 42_467_328L, 75.14m, // bloat_bytes_estimate, bloat_pct_estimate + false, 8, true, // estimate_unavailable, alignment_bytes, pgstattuple_available + }); + + var rows = await PgTableBloatStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal("public", row.SchemaName); + Assert.Equal("churny", row.TableName); + Assert.Equal(56_516_608L, row.HeapBytes); + Assert.Equal(6_899L, row.HeapPages); + Assert.Equal(24_000_000L, row.ToastBytes); + Assert.Equal(6_758_400L, row.IndexBytes); + Assert.Equal(100_000L, row.LiveTuples); + Assert.Equal(500L, row.DeadTuples); + Assert.Equal(1_200L, row.ModsSinceAnalyze); + Assert.Equal(new DateTime(2026, 8, 19, 3, 15, 0), row.LastAnalyzed); + Assert.Equal(436.0d, row.EstimatedTupleBytes); + Assert.Equal(1_715L, row.EstimatedHeapPages); + Assert.Equal(70, row.FillFactor); + Assert.Equal(42_467_328L, row.BloatBytesEstimate); + Assert.Equal(75.14m, row.BloatPctEstimate); + Assert.False(row.EstimateUnavailable); + Assert.Equal(8, row.AlignmentBytes); + Assert.True(row.PgstattupleAvailable); + } + + /// + /// The never-analyzed sentinel. PostgreSQL 14 and above report reltuples = -1 for a table that has + /// never been analyzed — a deliberate "unknown" distinct from "empty" — and the reader must preserve that + /// rather than folding it to 0, which would CLAIM the table holds no rows. It is exactly what + /// estimate_unavailable keys on, and the state in which the estimator reported 92.68% against a + /// true 48.86%. + /// + [Fact] + public async Task PreservesTheNeverAnalyzedSentinelRatherThanClaimingTheTableIsEmpty() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "unanalyzed", + 22_863_872L, 2_791L, + 0L, 0L, + DBNull.Value, // live_tuples -> -1, NOT 0 + 0L, 0L, + DBNull.Value, // last_analyzed -> null: never analyzed at all + 28.0d, + DBNull.Value, // estimated_heap_pages -> -1: could not be computed + 100, + 21_176_320L, 92.62m, + true, 8, false, + }); + + var rows = await PgTableBloatStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + + /* -1, not 0. 0 would be the claim "this table is empty", which is a different and false statement. */ + Assert.Equal(-1L, row.LiveTuples); + + /* -1 again, and for the same reason: 0 would claim the table should occupy no pages. */ + Assert.Equal(-1L, row.EstimatedHeapPages); + + /* NULL means never analyzed by either route, which is a STRONGER statement than a stale timestamp + and must not be rendered as "unknown". */ + Assert.Null(row.LastAnalyzed); + + Assert.True(row.EstimateUnavailable); + Assert.False(row.PgstattupleAvailable); + } + + /// + /// A row from the pg_monitor-only permissions state: the estimate is a large confident number and the + /// flag beside it is the only thing that says it has no basis. The reader must carry the flag through + /// untouched — a defaulted-false flag here would publish the number. + /// + [Fact] + public async Task CarriesTheUnavailableFlagThroughUnchanged() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "gadget", + 49_233_920L, 6_010L, + 0L, 26_853_376L, + 200_000L, 0L, 0L, + new DateTime(2026, 8, 19, 3, 15, 0, DateTimeKind.Unspecified), + 28.0d, 686L, 100, + 43_614_208L, 88.59m, + true, 8, true, + }); + + var rows = await PgTableBloatStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.True(row.EstimateUnavailable); + + /* The number is still STORED - suppression is the read's job, not the collector's. Storing it keeps + the row diagnosable: a reader can see what the arithmetic produced and why it was rejected. */ + Assert.Equal(88.59m, row.BloatPctEstimate); + Assert.Equal(43_614_208L, row.BloatBytesEstimate); + } + + /// + /// Absent counters fall back to 0 and absent booleans to false, but fillfactor falls back to 100 + /// — the PostgreSQL default. A 0 there would make the estimate divide by nothing. + /// + [Fact] + public async Task FillFactorFallsBackToThePostgresDefault() + { + var reader = new FakeCollectorDataReader( + new object[] + { + "public", "narrow", + 3_629_056L, 443L, + 0L, 2_260_992L, + 100_000L, 0L, 0L, + DBNull.Value, + 36.0d, 441L, + DBNull.Value, // fillfactor -> 100 + 16_384L, 0.45m, + DBNull.Value, DBNull.Value, DBNull.Value, + }); + + var rows = await PgTableBloatStatsCollector.Instance.ReadAsync(reader, MakeContext(), CancellationToken.None); + + var row = Assert.Single(rows); + Assert.Equal(100, row.FillFactor); + Assert.False(row.EstimateUnavailable); + Assert.Equal(0, row.AlignmentBytes); + Assert.False(row.PgstattupleAvailable); + } + + [Fact] + public async Task ReturnsNoRowsWhenNoTableClearsTheFloor() + { + var rows = await PgTableBloatStatsCollector.Instance.ReadAsync( + new FakeCollectorDataReader(), MakeContext(), CancellationToken.None); + + Assert.Empty(rows); + } + + /// + /// No stored deltas. Every column here is a LEVEL — how big the table is now, how much of it the + /// estimate thinks is waste now. The interesting reading over time is the trend across stored samples, + /// which the read computes from the raw levels; differencing at collection time would throw away the + /// absolute number, which is the one somebody acts on. + /// + [Fact] + public void TakesNoDeltas() + { + var deltas = new RecordingCollectorDeltaCalculator(); + + PgTableBloatStatsCollector.Instance.WritePayload( + SampleRow(), + new RecordingCollectorRowWriter(), + MakeContext(deltas: deltas)); + + Assert.Empty(deltas.Calls); + } + + /// + /// Every payload column is written, in order. WritePayload is positional, so a column added without a + /// matching Value() shifts everything after it and stores data that is silently wrong — which on this + /// collector would mean an estimate landing in the column a reader trusts unconditionally. + /// + [Fact] + public void WritesEveryPayloadColumnInOrder() + { + var writer = new RecordingCollectorRowWriter(); + + PgTableBloatStatsCollector.Instance.WritePayload(SampleRow(), writer, MakeContext()); + + Assert.Equal(PgTableBloatStatsCollector.Instance.PayloadColumns.Count, writer.Values.Count); + + /* The connection's database, NOT a value parsed from the result set. */ + Assert.Equal("appdb", writer.Values[0]); + + Assert.Equal("public", writer.Values[1]); + Assert.Equal("churny", writer.Values[2]); + Assert.Equal(56_516_608L, writer.Values[3]); // heap_bytes + Assert.Equal(6_899L, writer.Values[4]); // heap_pages + Assert.Equal(24_000_000L, writer.Values[5]); // toast_bytes + Assert.Equal(100_000L, writer.Values[7]); // live_tuples + Assert.Equal(1_200L, writer.Values[9]); // mods_since_analyze + Assert.Equal(new DateTime(2026, 8, 19, 3, 15, 0), writer.Values[10]); + Assert.Equal(436.0d, writer.Values[11]); // estimated_tuple_bytes + Assert.Equal(70, writer.Values[13]); // fillfactor + Assert.Equal(42_467_328L, writer.Values[14]); // bloat_bytes_estimate + Assert.Equal(75.14m, writer.Values[15]); // bloat_pct_estimate + Assert.Equal(false, writer.Values[16]); // estimate_unavailable + Assert.Equal(8, writer.Values[17]); // alignment_bytes + Assert.Equal(true, writer.Values[18]); // pgstattuple_available + } + + private static PgTableBloatStatsCollector.Row SampleRow() => new( + SchemaName: "public", + TableName: "churny", + HeapBytes: 56_516_608, + HeapPages: 6_899, + ToastBytes: 24_000_000, + IndexBytes: 6_758_400, + LiveTuples: 100_000, + DeadTuples: 500, + ModsSinceAnalyze: 1_200, + LastAnalyzed: new DateTime(2026, 8, 19, 3, 15, 0), + EstimatedTupleBytes: 436.0, + EstimatedHeapPages: 1_715, + FillFactor: 70, + BloatBytesEstimate: 42_467_328, + BloatPctEstimate: 75.14m, + EstimateUnavailable: false, + AlignmentBytes: 8, + PgstattupleAvailable: true); + + /// + /// HOURLY, matching pg_autovacuum_stats deliberately rather than by copying: this collector measures the + /// DAMAGE whose CAUSE that one measures, and correlating the two requires a common grain. It is also the + /// second per-database fan-out, so sharing the cadence means one connection-budget decision instead of + /// two. + /// + /// 90 days because the useful reading of bloat is a TREND — is this table's waste growing, holding, + /// or being reclaimed — and a spot percentage on its own is what gets someone to run VACUUM FULL on a + /// Tuesday. + /// + [Fact] + public void RegisteredInBothTheCatalogAndTheSchedule() + { + Assert.Contains(CollectorCatalog.All, d => d.Name == "pg_table_bloat_stats"); + + var schedule = CollectorScheduleDefaults.All["pg_table_bloat_stats"]; + + Assert.Equal(60, schedule.FrequencyMinutes); + Assert.Equal(90, schedule.RetentionDays); + Assert.True(schedule.DefaultEnabled); + + /* The cause and the damage share a grain, which is the whole argument for the cadence. Asserted + against the sibling rather than as a second literal, so the two cannot drift apart silently. */ + Assert.Equal( + CollectorScheduleDefaults.All["pg_autovacuum_stats"].FrequencyMinutes, + schedule.FrequencyMinutes); + } + + /// + /// The capability vocabulary has a noun phrase for this collector, so a SQL Server target asked + /// get_pg_table_bloat is told what it does not collect rather than getting the generic fallback. + /// + [Fact] + public void TheCapabilityMessageNamesWhatIsNotCollected() + { + var message = CollectorEngineCapability.NotCollectedMessage( + "sql-01", CollectorEngineCapability.UnknownEngineEdition, MonitoredEngineKind.SqlServer, "pg_table_bloat_stats"); + + Assert.NotNull(message); + Assert.Contains("bloat estimate", message, StringComparison.Ordinal); + Assert.Contains("dead-tuple", message, StringComparison.Ordinal); + Assert.DoesNotContain("the data this read is served from", message, StringComparison.Ordinal); + } +} diff --git a/Lite.Tests/PgWaitExclusionParityTests.cs b/Lite.Tests/PgWaitExclusionParityTests.cs new file mode 100644 index 000000000..a242bac1e --- /dev/null +++ b/Lite.Tests/PgWaitExclusionParityTests.cs @@ -0,0 +1,117 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2630: the two PostgreSQL wait collectors must agree about what counts as a wait. +/// +/// +/// They read different sources for the same measurement — pg_wait_stats takes Aurora's own +/// instrumentation, pg_wait_sampling takes a sampling profiler — and #2625 now tells a +/// stock-PostgreSQL operator, in as many words, to read the second one INSTEAD of the first. Two +/// exclusion lists means one server answers "what is it waiting on" differently depending on its flavor, +/// which is worse than either answer alone. +/// +/// +/// +/// They did disagree. The sampler excluded Activity and nothing else, and on the first target with +/// real client connections ClientRead was 2,717,290 of 2,717,989 samples — 100.0%, with every +/// real event rounded to zero. The unfiltered profile on that target held 17,864,575 samples of +/// Activity, Client and Timeout against 2,150 samples of everything else. +/// +/// +/// +/// Activity was excluded from the start because it is visibly absurd — background sleep loops +/// accumulating a second per second of uptime. Client and Timeout need a CLIENT to be idle +/// before they dominate, which no container and no CI run provides. That is the whole reason this went +/// unnoticed, and the reason the guard is a parity assertion rather than a list: whichever list is edited +/// next, the other has to follow. +/// +/// +public class PgWaitExclusionParityTests +{ + [Fact] + public void TheAuroraCollectorExcludesTheThreeTypesThatAreNotWork() + { + Assert.Equal( + new[] { "Activity", "Client", "Timeout" }, + PgWaitStatsCollector.IgnoredWaitTypes.OrderBy(t => t, StringComparer.Ordinal).ToArray()); + } + + /// + /// The sampler splices the SAME set into its SQL. Asserted against the set rather than against three + /// string literals, so adding a fourth type to the shared definition carries the sampler with it. + /// + [Fact] + public void TheSamplerExcludesEveryTypeTheAuroraCollectorDoes() + { + var sql = PgWaitSamplingCollector.Instance.BuildQuery(Context()).Text; + + foreach (var type in PgWaitStatsCollector.IgnoredWaitTypes) + { + Assert.Contains($"'{type}'", sql, StringComparison.Ordinal); + } + } + + /// + /// The one type that must NOT be filtered, and the reason the predicate coalesces before it compares. + /// + /// A backend on CPU arrives with a NULL event_type. Comparing NULL against a list yields + /// NULL, which a WHERE discards — so a filter written the obvious way would silently drop this + /// collector's distinctive signal, the one that lets its share column answer "waiting or working?" + /// before it answers "waiting on what?". Measured after the fix: CPU 1,914 samples, IO 204, LWLock 9, + /// IPC 2 — the CPU row survived and leads. + /// + [Fact] + public void TheCpuRowSurvivesTheFilter_BecauseTheTypeIsCoalescedBeforeItIsCompared() + { + var sql = PgWaitSamplingCollector.Instance.BuildQuery(Context()).Text; + + Assert.Contains("coalesce(p.event_type, 'CPU') NOT IN", sql, StringComparison.Ordinal); + Assert.DoesNotContain("'CPU'", string.Join(",", PgWaitStatsCollector.IgnoredWaitTypes), StringComparison.Ordinal); + } + + /// + /// And the filter is applied at COLLECTION, not at read. These are cumulative counters: leaving the + /// excluded types in would spend store on millions of samples of a server doing nothing, and every + /// reader would have to remember to filter them again. + /// + [Fact] + public void TheExclusionIsInTheCollectorsQuery_NotLeftToTheReader() + { + var sql = PgWaitSamplingCollector.Instance.BuildQuery(Context()).Text; + + var whereIndex = sql.IndexOf("WHERE", StringComparison.Ordinal); + var groupIndex = sql.IndexOf("GROUP BY", StringComparison.Ordinal); + + Assert.True(whereIndex >= 0 && groupIndex > whereIndex, "The collector query has no WHERE ahead of its GROUP BY."); + Assert.Contains("Client", sql[whereIndex..groupIndex], StringComparison.Ordinal); + } + + private static CollectorContext Context() + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 26, 12, 0, 0, DateTimeKind.Utc), + Deltas = new RecordingCollectorDeltaCalculator(), + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + PostgresVersionNum = 170000, + }, + }; +} diff --git a/Lite.Tests/PgWaitSamplingCollectorDefinitionTests.cs b/Lite.Tests/PgWaitSamplingCollectorDefinitionTests.cs new file mode 100644 index 000000000..eae72e4ea --- /dev/null +++ b/Lite.Tests/PgWaitSamplingCollectorDefinitionTests.cs @@ -0,0 +1,153 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// Pins the wait-sampling collector on the three things measurement decided, each of which produces +/// correct-looking output when it is wrong. +/// +public class PgWaitSamplingCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext() + => new() + { + ServerId = 42, + ServerName = "test-server", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = 17, + PostgresVersionNum = 170000, + }, + }; + + private static string Sql => PgWaitSamplingCollector.Instance.BuildQuery(MakeContext()).Text; + + [Fact] + public void Identity_IsTheTableAndEngineTheStoreExpects() + { + Assert.Equal("pg_wait_sampling", PgWaitSamplingCollector.Instance.Name); + Assert.Equal("pg_wait_sampling", PgWaitSamplingCollector.Instance.TargetTable); + Assert.Equal(CollectorTargetEngine.PostgreSql, PgWaitSamplingCollector.Instance.TargetEngine); + } + + /// + /// Activity must never be stored. Measured on an idle PostgreSQL 17, the entire top of the raw + /// profile was background processes waiting for work — AutovacuumMain, LogicalLauncherMain, + /// WalWriterMain, CheckpointerMain — and they accumulate forever precisely BECAUSE the + /// server is quiet. Rank that and every healthy server reports autovacuum's idle loop as its top wait. + /// + /// #2630 made that three types rather than one: Client and Timeout are the same kind + /// of not-work and need a live CLIENT to be idle before they dominate, which is why an idle PostgreSQL 17 + /// did not surface them. On the first target profiled with real connections, ClientRead was 100.0% + /// of the profile. The set now comes from so the two + /// wait collectors cannot disagree about what counts as a wait — + /// owns that half. + /// + [Fact] + public void IdleBackgroundWaits_AreExcludedAtTheSource() + { + Assert.Contains("'Activity'", Sql, StringComparison.Ordinal); + Assert.Contains("'Client'", Sql, StringComparison.Ordinal); + Assert.Contains("'Timeout'", Sql, StringComparison.Ordinal); + } + + /// + /// The filter must survive a NULL event type, which means a backend that was NOT waiting — real signal, + /// deliberately kept and labelled CPU. + /// + /// It was IS DISTINCT FROM for exactly this reason: NULL <> 'Activity' is NULL + /// rather than true, so a plain inequality would silently discard every on-CPU sample. #2630 widened the + /// filter to a three-type list and the same trap reappears one step along — NULL against a list is also + /// NULL — so the type is coalesced to 'CPU' BEFORE it is compared, and CPU is not in + /// the excluded set. + /// + [Fact] + public void TheActivityFilter_IsNullSafe() + { + Assert.DoesNotMatch(new Regex(@"event_type\s*<>\s*'Activity'"), Sql); + Assert.DoesNotMatch(new Regex(@"(? + /// The profile period travels with the counts. A sample tally is uninterpretable without the period it + /// was gathered at, and storing a derived millisecond figure would bake today's period into history. + /// + [Fact] + public void TheProfilePeriod_IsCollectedAlongsideTheCounts() + { + Assert.Contains("pg_wait_sampling.profile_period", Sql, StringComparison.Ordinal); + Assert.Contains("profile_period_ms", Sql, StringComparison.Ordinal); + + var columns = PgWaitSamplingCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + Assert.Contains("profile_period_ms", columns); + Assert.Contains("sample_count", columns); + + /* The name must not promise a duration it is not. */ + Assert.DoesNotContain("wait_ms", columns); + Assert.DoesNotContain("wait_time_ms", columns); + } + + /// + /// Read with the MISSING_OK form, like every other GUC this codebase reads: a renamed or absent setting + /// must degrade one column rather than fail the whole collection. + /// + [Fact] + public void TheGucRead_ToleratesTheSettingBeingAbsent() + { + foreach (Match call in Regex.Matches(Sql, @"current_setting\(([^)]*)\)")) + { + Assert.Contains(", true", call.Groups[1].Value, StringComparison.Ordinal); + } + } + + /// + /// Cluster-wide, so no database_name — the profile spans every backend and version 1.1 exposes no + /// database column. Inventing that attribution is precisely the scope error #2599 removed from three + /// other collectors. + /// + [Fact] + public void ItDoesNotClaimPerDatabaseAttribution() + { + Assert.False(PgWaitSamplingCollector.Instance.RunsPerDatabase( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = 17 })); + + Assert.DoesNotContain( + "database_name", + PgWaitSamplingCollector.Instance.PayloadColumns.Select(c => c.Name)); + } + + /// + /// queryid is a 64-bit hash — measured at nineteen digits on the live rig — so the column has to + /// be a bigint, and unattributed waits are kept rather than filtered so the stored profile still agrees + /// with the server's own totals. + /// + [Fact] + public void QueryId_IsBigIntAndUnattributedWaitsSurvive() + { + var queryId = PgWaitSamplingCollector.Instance.PayloadColumns.Single(c => c.Name == "query_id"); + Assert.Equal(CollectorColumnType.BigInt, queryId.Type); + + Assert.DoesNotMatch(new Regex(@"queryid\s*(<>|!=)\s*0"), Sql); + } +} diff --git a/Lite.Tests/PgWriteStatsCollectorDefinitionTests.cs b/Lite.Tests/PgWriteStatsCollectorDefinitionTests.cs new file mode 100644 index 000000000..57d80a3c7 --- /dev/null +++ b/Lite.Tests/PgWriteStatsCollectorDefinitionTests.cs @@ -0,0 +1,261 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Linq; +using System.Text.RegularExpressions; +using Lite.Tests.Helpers; +using PerformanceMonitor.Collectors; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2544: the write-side collector. Almost every assertion here is about the VERSION SURFACE, because that +/// is where this collector can actually fail — and it fails in the one way a text-only test cannot normally +/// catch: selecting a column that a major removed raises 42703 at run time on that major only, and takes the +/// whole collection with it. The fleet this was written for is an even split across the 16→17 break, so +/// "works on the version I tested" is not a property worth having. +/// +public class PgWriteStatsCollectorDefinitionTests +{ + private static readonly RecordingCollectorDeltaCalculator s_deltas = new(); + + private static CollectorContext MakeContext(int major) + => new() + { + ServerId = 42, + ServerName = "pg-target", + CollectionTime = new DateTime(2026, 8, 25, 12, 0, 0, DateTimeKind.Utc), + Deltas = s_deltas, + Target = new CollectorTargetInfo + { + Engine = CollectorTargetEngine.PostgreSql, + PostgresMajorVersion = major, + }, + ExcludedDatabases = Array.Empty(), + }; + + private static string Sql(int major) + => PgWriteStatsCollector.Instance.BuildQuery(MakeContext(major)).Text; + + /// + /// The columns 17 REMOVED from pg_stat_bgwriter must not be referenced on 17 or later. Five of + /// them were renamed into pg_stat_checkpointer; buffers_backend and + /// buffers_backend_fsync went to pg_stat_io instead and have no successor here at all. + /// + [Theory] + [InlineData(17)] + [InlineData(18)] + [InlineData(19)] + public void PreSeventeenBgwriterColumns_AreNotSelectedOnSeventeenOrLater(int major) + { + var sql = Sql(major); + + foreach (var gone in new[] + { + "checkpoints_timed", "checkpoints_req", "checkpoint_write_time", "checkpoint_sync_time", + "buffers_checkpoint", "buffers_backend", "buffers_backend_fsync", + }) + { + /* Word-boundary matched. A plain Contains would fire on the OUTPUT aliases, which deliberately + keep some of these names (buffers_backend is a stored column on every major, it is simply + NULL from 17) — the assertion is about what is READ FROM THE VIEW, not what is emitted. */ + Assert.DoesNotMatch(new Regex($@"pg_stat_bgwriter\)[^\n]*\b{gone}\b"), sql); + Assert.DoesNotMatch(new Regex($@"\b{gone}\s+FROM\s+pg_catalog\.pg_stat_bgwriter"), sql); + } + } + + /// + /// And the converse: below 17 those columns are exactly where the values must come from, because + /// pg_stat_checkpointer does not exist there at all. A collector that referenced it on 16 would + /// fail with 42P01 on half this fleet. + /// + [Theory] + [InlineData(14)] + [InlineData(15)] + [InlineData(16)] + public void BelowSeventeen_ReadsTheOldColumns_AndNeverTouchesPgStatCheckpointer(int major) + { + var sql = Sql(major); + + Assert.DoesNotContain("pg_stat_checkpointer", sql, StringComparison.Ordinal); + Assert.Contains("checkpoints_timed", sql, StringComparison.Ordinal); + Assert.Contains("checkpoints_req", sql, StringComparison.Ordinal); + Assert.Contains("buffers_checkpoint", sql, StringComparison.Ordinal); + Assert.Contains("buffers_backend", sql, StringComparison.Ordinal); + } + + /// + /// The 18 hazard, and the specific one #2544 was reopened to describe: pg_stat_wal LOST + /// wal_write, wal_sync, wal_write_time and wal_sync_time. A collector written + /// and tested against 17 selects all four and raises 42703 on an 18 target every cycle. + /// + [Theory] + [InlineData(18)] + [InlineData(19)] + public void EighteenAndLater_DoesNotSelectTheRemovedWalColumns(int major) + { + var sql = Sql(major); + + foreach (var gone in new[] { "wal_write", "wal_sync", "wal_write_time", "wal_sync_time" }) + { + /* Qualified with the view's alias, so this fires on a READ of the column and not on the output + alias of the same name, which is retained on every major and simply carries NULL here. */ + Assert.DoesNotContain($"w.{gone}", sql, StringComparison.Ordinal); + } + } + + [Theory] + [InlineData(14)] + [InlineData(16)] + [InlineData(17)] + public void BelowEighteen_StillReadsTheWalTimingColumns(int major) + { + var sql = Sql(major); + + Assert.Contains("w.wal_write_time", sql, StringComparison.Ordinal); + Assert.Contains("w.wal_sync_time", sql, StringComparison.Ordinal); + } + + /// + /// Columns that arrived at 18 must not be read below it — the mirror of the removal case, and the one + /// that would fail on the 26 clusters of this fleet running 17. + /// + [Theory] + [InlineData(14)] + [InlineData(16)] + [InlineData(17)] + public void BelowEighteen_DoesNotSelectEighteenOnlyColumns(int major) + { + var sql = Sql(major); + + /* Asserted as "not READ FROM the view", not "the string is absent". Both names survive as OUTPUT + ALIASES on every major - the stored shape is the union and simply carries NULL here - so a plain + absence check fails against a perfectly correct query. */ + Assert.DoesNotMatch(new Regex(@"num_done\s+FROM\s+pg_catalog\.pg_stat_checkpointer"), sql); + Assert.DoesNotMatch(new Regex(@"slru_written\s+FROM\s+pg_catalog\.pg_stat_checkpointer"), sql); + } + + /// + /// Every version gate is a >= floor, so a future major that keeps 18's shape does not fall off + /// the end into a column that no longer exists. Asserted by behaviour on a version that does not exist + /// yet rather than by reading the source for a comparison operator. + /// + [Fact] + public void AFutureMajor_BehavesLikeTheNewestKnownOne() + => Assert.Equal(Sql(18), Sql(99)); + + /// + /// The STORED shape is identical on every major — that is what lets one store hold a 16 and an 18 target + /// and mean the same thing in both rows. Only the SQL varies. + /// + [Fact] + public void ThePayloadShape_DoesNotVaryByVersion() + { + var columns = PgWriteStatsCollector.Instance.PayloadColumns; + + Assert.Equal(26, columns.Count); + Assert.Equal(columns.Count, columns.Select(c => c.Name).Distinct(StringComparer.Ordinal).Count()); + + /* One row per SELECT column, in order — a mismatch here is a silently shifted binary COPY, which + writes every value into the wrong column rather than failing. + + FROM and CROSS JOIN lines are excluded before matching: the two TABLE aliases (AS b, AS w) are + spelled with the same keyword as the column aliases, and counting them shifted the whole list by + two while looking like an ordering bug in the collector. */ + foreach (var major in new[] { 14, 16, 17, 18 }) + { + var selected = Sql(major) + .Split('\n') + .Where(line => !line.TrimStart().StartsWith("FROM", StringComparison.Ordinal) + && !line.TrimStart().StartsWith("CROSS JOIN", StringComparison.Ordinal)) + .Select(line => Regex.Match(line, @"\bAS\s+([a-z_]+),?\s*$")) + .Where(m => m.Success) + .Select(m => m.Groups[1].Value) + .ToArray(); + + Assert.Equal(columns.Select(c => c.Name).ToArray(), selected); + } + } + + /// + /// wal_bytes is numeric upstream, not bigint, precisely because cumulative WAL + /// volume may exceed 2^63. Storing it as a bigint would work for years and then wrap. + /// + [Fact] + public void WalBytes_IsDecimal_NotBigInt() + { + var column = PgWriteStatsCollector.Instance.PayloadColumns.Single(c => c.Name == "wal_bytes"); + + Assert.Equal(CollectorColumnType.Decimal, column.Type); + } + + /// + /// All three stats_reset stamps are stored, because pg_stat_reset_shared takes a target and + /// the three families reset independently. A read that differenced across a reset it could not see would + /// report a negative interval as an enormous positive one. + /// + [Fact] + public void AllThreeStatsResetStamps_AreStored() + { + var names = PgWriteStatsCollector.Instance.PayloadColumns.Select(c => c.Name).ToArray(); + + Assert.Contains("checkpointer_stats_reset", names); + Assert.Contains("bgwriter_stats_reset", names); + Assert.Contains("wal_stats_reset", names); + } + + /// + /// pg_stat_wal is the binding floor at 14. Below that the collector must not dispatch at all + /// rather than fail per cycle. + /// + [Theory] + [InlineData(13, false)] + [InlineData(14, true)] + [InlineData(17, true)] + public void AppliesTo_FloorsAtFourteen(int major, bool expected) + => Assert.Equal(expected, PgWriteStatsCollector.Instance.AppliesTo( + new CollectorTargetInfo { Engine = CollectorTargetEngine.PostgreSql, PostgresMajorVersion = major })); + + /// + /// Timestamps come back as naive UTC via AT TIME ZONE 'UTC', never a bare ::timestamp cast + /// — the cast renders in the SESSION's TimeZone, so a store session east of UTC would record a stamp + /// hours away from the one the server meant. Same rule pg_stat_io follows. + /// + [Fact] + public void StatsResetStamps_AreConvertedToUtc_NotBareCast() + { + var sql = Sql(17); + + Assert.Equal(3, Regex.Matches(sql, @"AT TIME ZONE 'UTC'").Count); + Assert.DoesNotMatch(new Regex(@"stats_reset\s*::\s*timestamp"), sql); + } + + /// + /// Catalog reads are schema-qualified. pg_catalog is searched implicitly but not necessarily + /// FIRST, so an unqualified read can resolve to an object a user created in a schema earlier in the + /// monitoring login's search_path. + /// + [Theory] + [InlineData(16)] + [InlineData(17)] + [InlineData(18)] + public void EveryCatalogRead_IsSchemaQualified(int major) + { + var sql = Sql(major); + + foreach (var view in new[] { "pg_stat_bgwriter", "pg_stat_wal", "pg_stat_checkpointer" }) + { + foreach (Match match in Regex.Matches(sql, $@"(\S*){Regex.Escape(view)}")) + { + Assert.Equal("pg_catalog.", match.Groups[1].Value); + } + } + } +} diff --git a/Lite.Tests/QueryHeatmapToolTests.cs b/Lite.Tests/QueryHeatmapToolTests.cs new file mode 100644 index 000000000..64a22a4ff --- /dev/null +++ b/Lite.Tests/QueryHeatmapToolTests.cs @@ -0,0 +1,342 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_query_heatmap (#2484), the twin of Darling's. +/// +/// Two things are pinned here that no reader-level test would see. The three empty branches — the third +/// of which is unique to this read: collection ran, the window has captures, and every one recorded zero +/// executions, which is an IDLE server rather than a broken one. And the BUCKETING, which is the whole reason +/// the read exists in this shape: the desktop viewer hardcodes 5-minute bins, so an agent and a desktop +/// looking at the same server over the same window have to land on the same grid. +/// +public sealed class QueryHeatmapToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "HeatmapSrv"; + private const string Db = "AppDb"; + + /* Lite DERIVES its server id from the storage name; a hardcoded one would seed rows the tool looks + straight past and pass the never-collected assertion for the wrong reason. */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public QueryHeatmapToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-heatmap-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task ThreeKindsOfNothing_AreThreeDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + var baseNow = Truncate(DateTime.UtcNow); + + /* 1. nothing ever collected. */ + var never = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24)); + Assert.Equal("unavailable", never.GetProperty("status").GetString()); + var neverText = never.GetProperty("message").GetString()!; + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + Assert.Contains("PERIODIC table rather than an edge table", neverText, StringComparison.Ordinal); + + /* + 2. rows exist, but not in the window. Seeded FIRST, and the ordering is load-bearing: once a row + exists 30 minutes ago no window a caller can legally ask for excludes it, so this branch is + unreachable afterwards. The states are walked by moving hours_back over one growing set of rows. + */ + await SeedAsync(baseNow.AddHours(-40), "0xOLD", deltaExec: 4, deltaElapsed: 200_000); + + var noWindow = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 2)); + Assert.Equal("empty", noWindow.GetProperty("status").GetString()); + var noWindowText = noWindow.GetProperty("message").GetString()!; + Assert.Contains("Widen hours_back", noWindowText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", noWindowText, StringComparison.Ordinal); + + /* + 3. collected in the window, and every capture recorded ZERO executions. The branch only this read + has: collection is healthy and the server is idle, so telling the caller to widen the window would + point them at the wrong problem. + */ + await SeedAsync(baseNow.AddMinutes(-30), "0xIDLE", deltaExec: 0, deltaElapsed: 999_000); + + var idle = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24)); + Assert.Equal("empty", idle.GetProperty("status").GetString()); + var idleText = idle.GetProperty("message").GetString()!; + Assert.Contains("zero execution delta", idleText, StringComparison.Ordinal); + Assert.DoesNotContain("Widen hours_back", idleText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", idleText, StringComparison.Ordinal); + + /* Three sentences, not one sentence three times. */ + var messages = new[] { neverText, noWindowText, idleText }; + Assert.Equal(3, messages.Distinct(StringComparer.Ordinal).Count()); + } + + /// + /// The grid itself, on the viewer's own 5-minute bins. + /// The seed is floored to the hour, and that is not tidiness: time_bucket aligns bins to the origin, + /// so an unfloored "three hours ago" lands at an arbitrary minute and the three rows below straddle a + /// 5-minute boundary roughly three times in five. It would fail on a clock, not on a defect. + /// + [Fact] + public async Task TheGrid_BinsAtFiveMinutes_BucketsByMagnitude_AndNamesTheHottestQuery() + { + var service = new LocalDataService(_duckDb); + var t1 = FloorToHour(Truncate(DateTime.UtcNow).AddHours(-3)); + + /* Two queries at ~50 ms/exec share bucket 2 of one bin; 0.5 ms/exec lands in bucket 0 of the SAME + bin; a fourth 35 minutes later gets its own bin. The zero-execution row contributes to nothing. */ + await SeedAsync(t1, "0xHOT", deltaExec: 5, deltaElapsed: 250_000); + await SeedAsync(t1.AddMinutes(1), "0xCOLD", deltaExec: 1, deltaElapsed: 50_000); + await SeedAsync(t1.AddMinutes(2), "0xLOW", deltaExec: 2, deltaElapsed: 1_000); + await SeedAsync(t1.AddMinutes(35), "0xNEXT", deltaExec: 3, deltaElapsed: 90_000); + await SeedAsync(t1.AddMinutes(3), "0xZERO", deltaExec: 0, deltaElapsed: 999_000); + + var grid = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24)); + + Assert.Equal(ServerName, grid.GetProperty("server").GetString()); + Assert.Equal("duration", grid.GetProperty("metric").GetString()); + Assert.Equal(5, grid.GetProperty("bucket_minutes").GetInt32()); + Assert.True(grid.GetProperty("bucket_minutes_matches_desktop_viewer").GetBoolean()); + Assert.Equal(2, grid.GetProperty("time_bin_count").GetInt32()); + Assert.Equal(3, grid.GetProperty("cell_count").GetInt32()); + Assert.False(grid.GetProperty("truncated").GetBoolean()); + Assert.Equal(7, grid.GetProperty("magnitude_buckets").GetArrayLength()); + + var cells = grid.GetProperty("cells").EnumerateArray().ToArray(); + + /* Chronological, whatever order the cap fetched them in. */ + var times = cells.Select(c => DateTime.Parse(c.GetProperty("time_bucket").GetString()!, + System.Globalization.CultureInfo.InvariantCulture, + System.Globalization.DateTimeStyles.RoundtripKind)).ToArray(); + Assert.Equal(times.OrderBy(t => t).ToArray(), times); + + /* + The reported window has to be the window the read USED. It spans exactly hours_back and brackets + every bin that came back. Taken after the read returns it drifts by the read's own duration — + on the one read whose entire output is a time axis, so a window that disagrees with the bins + under it is worse here than almost anywhere (review catch on the first cut of this file). + */ + var windowStart = ParseUtc(grid.GetProperty("window_start").GetString()!); + var windowEnd = ParseUtc(grid.GetProperty("window_end").GetString()!); + Assert.Equal(24.0, (windowEnd - windowStart).TotalHours, 3); + Assert.True(windowStart <= times[0], "window_start is after the first bin the read returned"); + Assert.True(windowEnd >= times[^1], "window_end is before the last bin the read returned"); + + var bucket0 = cells.Single(c => c.GetProperty("bucket_index").GetInt32() == 0); + Assert.Equal(1, bucket0.GetProperty("query_count").GetInt64()); + Assert.Equal("0xLOW", bucket0.GetProperty("top_query_hash").GetString()); + Assert.Equal("0-1ms", bucket0.GetProperty("bucket_label").GetString()); + + /* Two DISTINCT queries in the cell, and the cell's top query is the most-EXECUTED rather than the + slowest — 0xHOT ran five times, 0xCOLD once, and they are the same speed. */ + var bucket2 = cells.First(c => c.GetProperty("bucket_index").GetInt32() == 2); + Assert.Equal(2, bucket2.GetProperty("query_count").GetInt64()); + Assert.Equal("0xHOT", bucket2.GetProperty("top_query_hash").GetString()); + Assert.Equal("10-100ms", bucket2.GetProperty("bucket_label").GetString()); + + /* A coarser bin collapses the two columns into one — the parameter does something. */ + var coarse = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, null, null, 60)); + Assert.Equal(60, coarse.GetProperty("bucket_minutes").GetInt32()); + Assert.False(coarse.GetProperty("bucket_minutes_matches_desktop_viewer").GetBoolean()); + Assert.Equal(1, coarse.GetProperty("time_bin_count").GetInt32()); + Assert.Equal(2, coarse.GetProperty("cell_count").GetInt32()); + + /* Another metric is a different grid, not the same one relabelled. */ + var execCount = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, "execution_count")); + Assert.Equal("execution_count", execCount.GetProperty("metric").GetString()); + Assert.Equal("0-1", execCount.GetProperty("magnitude_buckets")[0].GetProperty("label").GetString()); + Assert.Contains("a total, not a per-execution average", execCount.GetProperty("metric_unit").GetString()!, StringComparison.Ordinal); + } + + /// + /// Truncation keeps the RECENT end of the window and hands back no partial column. + /// A cap of 2 against a 3-cell grid whose oldest bin holds 2 of those cells: the newest bin's one + /// cell fits, the oldest bin does not fit whole, so it is dropped rather than shown with a hole. A column + /// missing its low buckets reads as "nothing fast ran then" rather than "we stopped looking". + /// + [Fact] + public async Task ACappedGrid_KeepsTheRecentEnd_AndNoPartialColumn() + { + var service = new LocalDataService(_duckDb); + var t1 = FloorToHour(Truncate(DateTime.UtcNow).AddHours(-3)); + + await SeedAsync(t1, "0xHOT", deltaExec: 5, deltaElapsed: 250_000); + await SeedAsync(t1.AddMinutes(2), "0xLOW", deltaExec: 2, deltaElapsed: 1_000); + await SeedAsync(t1.AddMinutes(35), "0xNEXT", deltaExec: 3, deltaElapsed: 90_000); + + var capped = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, null, null, 5, 2)); + + Assert.True(capped.GetProperty("truncated").GetBoolean()); + Assert.Equal(1, capped.GetProperty("time_bin_count").GetInt32()); + Assert.Equal(1, capped.GetProperty("cell_count").GetInt32()); + Assert.Equal("0xNEXT", capped.GetProperty("cells")[0].GetProperty("top_query_hash").GetString()); + + /* first/last say which slice came back, so the missing part of the window is visible rather than + inferred. */ + Assert.Equal( + capped.GetProperty("first_time_bin").GetString(), + capped.GetProperty("last_time_bin").GetString()); + + /* An uncapped call over the same rows is the whole grid, so the cap above really was the cause. */ + var whole = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24)); + Assert.False(whole.GetProperty("truncated").GetBoolean()); + Assert.Equal(3, whole.GetProperty("cell_count").GetInt32()); + } + + /// + /// The anchor moves the window, and the resolved instant reaches the QUERY. + /// #2495's own failure mode is a tool that takes as_of, validates it, refuses a bad one + /// correctly, and then queries NOW — the validation succeeding is what makes the caller believe the + /// window moved, and nothing in the result says otherwise. So this proves it by CONTENT: an incident 30 + /// hours old is outside every default window on the surface, and only the anchored call can see it. + /// + [Fact] + public async Task TheAnchor_MovesTheWindow_AndTheDefaultAnchorCannotSeeAPastIncident() + { + var service = new LocalDataService(_duckDb); + var incident = FloorToHour(Truncate(DateTime.UtcNow).AddHours(-30)); + await SeedAsync(incident, "0xINCIDENT", deltaExec: 5, deltaElapsed: 250_000); + + var anchor = DateTime.SpecifyKind(incident.AddMinutes(30), DateTimeKind.Utc).ToString("o"); + var anchored = Root(await McpQueryTools.GetQueryHeatmap( + service, _serverManager, ServerName, 1, null, null, 5, 500, anchor)); + + Assert.Equal(1, anchored.GetProperty("cell_count").GetInt32()); + Assert.Equal("0xINCIDENT", anchored.GetProperty("cells")[0].GetProperty("top_query_hash").GetString()); + + /* The reported window is the ANCHORED one. Taken from a fresh clock it would say "now" over bins + that are 30 hours old. */ + Assert.Equal(anchor, anchored.GetProperty("window_end").GetString()); + Assert.Equal( + ParseUtc(anchored.GetProperty("cells")[0].GetProperty("time_bucket").GetString()!), + incident); + + /* The same LENGTH of window at the default anchor cannot reach it — so it is the anchor doing the + work, not hours_back. */ + var unanchored = Root(await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 1)); + Assert.Equal("empty", unanchored.GetProperty("status").GetString()); + Assert.Contains("Widen hours_back", unanchored.GetProperty("message").GetString()!, StringComparison.Ordinal); + + /* An anchor we cannot use is refused, never silently treated as now. */ + var badAnchor = await McpQueryTools.GetQueryHeatmap( + service, _serverManager, ServerName, 1, null, null, 5, 500, "last tuesday"); + Assert.Contains("Invalid as_of", badAnchor, StringComparison.Ordinal); + } + + [Fact] + public async Task OutOfRangeKnobs_AreRefused_NotSilentlyClamped() + { + var service = new LocalDataService(_duckDb); + await SeedAsync(Truncate(DateTime.UtcNow).AddMinutes(-30), "0xANY", deltaExec: 3, deltaElapsed: 60_000); + + var tooManyCells = await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, null, null, 5, 5000); + Assert.Contains("exceeds maximum of", tooManyCells, StringComparison.Ordinal); + Assert.Contains("1000", tooManyCells, StringComparison.Ordinal); + + var zeroBucket = await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, null, null, 0); + Assert.Contains("Must be between 1 and 1440", zeroBucket, StringComparison.Ordinal); + + var hugeBucket = await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, null, null, 1441); + Assert.Contains("Must be between 1 and 1440", hugeBucket, StringComparison.Ordinal); + + /* An unknown metric is REFUSED rather than turned into duration: a caller who asked for CPU and + silently got elapsed time would read the wrong grid with nothing to tell them so. */ + var badMetric = await McpQueryTools.GetQueryHeatmap(service, _serverManager, ServerName, 24, "reads"); + Assert.Contains("Invalid metric 'reads'", badMetric, StringComparison.Ordinal); + Assert.Contains("logical_reads", badMetric, StringComparison.Ordinal); + } + + private static JsonElement Root(string json) => JsonDocument.Parse(json).RootElement; + + private static DateTime ParseUtc(string value) => DateTime.Parse( + value, System.Globalization.CultureInfo.InvariantCulture, System.Globalization.DateTimeStyles.RoundtripKind); + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + /// Floors to the top of the hour, which is also a 5-minute boundary on the epoch grid the read + /// bins against, so the seeded rows land where the test says they do whatever time CI runs at. + private static DateTime FloorToHour(DateTime value) => + new(value.Ticks - (value.Ticks % TimeSpan.TicksPerHour), value.Kind); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedAsync(DateTime collectionTime, string queryHash, long deltaExec, long deltaElapsed) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO query_stats + (collection_id, collection_time, server_id, server_name, database_name, + query_hash, sql_handle, last_execution_time, delta_execution_count, + delta_worker_time, delta_elapsed_time, query_text) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12)"; + var naive = DateTime.SpecifyKind(collectionTime, DateTimeKind.Unspecified); + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = Db }); + cmd.Parameters.Add(new DuckDBParameter { Value = queryHash }); + cmd.Parameters.Add(new DuckDBParameter { Value = "0xSQLH" }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = deltaExec }); + cmd.Parameters.Add(new DuckDBParameter { Value = 0L }); + cmd.Parameters.Add(new DuckDBParameter { Value = deltaElapsed }); + cmd.Parameters.Add(new DuckDBParameter { Value = "SELECT * FROM Orders" }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/QueryStoreRegressionsToolTests.cs b/Lite.Tests/QueryStoreRegressionsToolTests.cs new file mode 100644 index 000000000..ae67ce7a0 --- /dev/null +++ b/Lite.Tests/QueryStoreRegressionsToolTests.cs @@ -0,0 +1,226 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_query_store_regressions (#2484), the twin of Darling's. +/// +/// Two things are pinned here that no reader-level test would see. The four empty branches - only one +/// of which is good news - and the DEDUP, which is correctness rather than performance on this read: the +/// baseline arm is unbounded while the recent arm is a short window, so the two sides being compared have +/// systematically different re-collection density per interval, and losing the dedup moves the averages and +/// the 25% gate for reasons that have nothing to do with the query. +/// +public sealed class QueryStoreRegressionsToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "RegressionsSrv"; + private const string Db = "AppDb"; + + /* Lite DERIVES its server id from the storage name; a hardcoded one would seed rows the tool looks + straight past and pass the never-collected assertion for the wrong reason. */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public QueryStoreRegressionsToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-regressions-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task FourKindsOfNothing_AreFourDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + var baseNow = Truncate(DateTime.UtcNow); + + /* 1. nothing ever collected. */ + var never = Root(await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 24)); + Assert.Equal("unavailable", never.GetProperty("status").GetString()); + var neverText = never.GetProperty("message").GetString()!; + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + Assert.Contains("Query Store may be OFF", neverText, StringComparison.Ordinal); + + /* + 2. a baseline exists but nothing landed in the window. Seeded FIRST, and the ordering is + load-bearing: once a row exists 30 minutes ago no window a caller can legally ask for excludes + it, so this branch is unreachable after the recent row is planted. The four states are walked + by moving hours_back over one fixed set of rows rather than by deleting rows between + assertions. + */ + await SeedAsync(baseNow.AddHours(-40), executions: 100, avgDurationUs: 1000, avgCpuUs: 1000, intervalId: 2); + + var noRecent = Root(await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 2)); + Assert.Equal("empty", noRecent.GetProperty("status").GetString()); + var noRecentText = noRecent.GetProperty("message").GetString()!; + Assert.Contains("Widen hours_back", noRecentText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", noRecentText, StringComparison.Ordinal); + + /* 3. both sides collected, nothing regressed: the ONE good-news answer. */ + await SeedAsync(baseNow.AddMinutes(-30), executions: 100, avgDurationUs: 1000, avgCpuUs: 1000, intervalId: 1); + + var clear = Root(await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 24)); + Assert.Equal("empty", clear.GetProperty("status").GetString()); + var clearText = clear.GetProperty("message").GetString()!; + Assert.Contains("all-clear", clearText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", clearText, StringComparison.Ordinal); + Assert.DoesNotContain("Widen", clearText, StringComparison.Ordinal); + + /* + 4. recent rows but NO baseline - the branch this read exists to get right, reached from the + SAME rows by widening the window until every one of them falls inside it. There is then nothing + to compare against, and answering "no regressions" would be a confident wrong answer rather + than a missing one. + */ + var noBaseline = Root(await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 48)); + Assert.Equal("unavailable", noBaseline.GetProperty("status").GetString()); + var noBaselineText = noBaseline.GetProperty("message").GetString()!; + Assert.Contains("no baseline", noBaselineText, StringComparison.Ordinal); + Assert.Contains("NOT a clean bill of health", noBaselineText, StringComparison.Ordinal); + + /* Widening makes the window bigger and the baseline SHORTER - the wrong direction. */ + Assert.Contains("Shorten hours_back", noBaselineText, StringComparison.Ordinal); + Assert.DoesNotContain("Widen", noBaselineText, StringComparison.Ordinal); + + /* Four sentences, not one sentence four times. */ + var messages = new[] { neverText, noBaselineText, noRecentText, clearText }; + Assert.Equal(4, messages.Distinct(StringComparer.Ordinal).Count()); + } + + [Fact] + public async Task ARealRegression_IsRanked_AndSurvivesTheDedup() + { + var service = new LocalDataService(_duckDb); + var baseNow = Truncate(DateTime.UtcNow); + + /* + Query 7 doubles its CPU and quadruples its duration between the baseline and the window. The + baseline interval is ALSO re-collected with a higher cumulative count, which is what the dedup + has to survive: un-deduped, the baseline exec_count would be 140 (50 + 90) rather than 90 and + the baseline averages would be an avg-of-avgs over two snapshots of one interval. + */ + await SeedAsync(baseNow.AddHours(-40), 50, 1000, 1000, intervalId: 3, queryId: 7); + await SeedAsync(baseNow.AddHours(-39), 90, 1000, 1000, intervalId: 3, queryId: 7); + await SeedAsync(baseNow.AddMinutes(-30), 200, 4000, 4000, intervalId: 4, queryId: 7); + + var hit = Root(await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 24)); + Assert.Equal(ServerName, hit.GetProperty("server").GetString()); + Assert.Equal(1, hit.GetProperty("regression_count").GetInt32()); + Assert.False(hit.GetProperty("truncated").GetBoolean()); + + var row = hit.GetProperty("regressions")[0]; + Assert.Equal(7, row.GetProperty("query_id").GetInt64()); + Assert.Equal(Db, row.GetProperty("database_name").GetString()); + + /* 1 ms -> 4 ms is +300%, the CRITICAL band (> 100%). */ + Assert.Equal("CRITICAL", row.GetProperty("severity").GetString()); + Assert.Equal(1.0, row.GetProperty("baseline_duration_ms").GetDouble(), 6); + Assert.Equal(4.0, row.GetProperty("recent_duration_ms").GetDouble(), 6); + Assert.Equal(300.0, row.GetProperty("duration_regression_percent").GetDouble(), 6); + + /* The ranking key: 3 ms per execution across the 200 recent executions. */ + Assert.Equal(600.0, row.GetProperty("additional_duration_ms").GetDouble(), 6); + + /* 90, not 140. This is the dedup, and it is the assertion the whole read leans on. */ + Assert.Equal(90, row.GetProperty("baseline_exec_count").GetInt64()); + Assert.Equal(200, row.GetProperty("recent_exec_count").GetInt64()); + } + + [Fact] + public async Task AnOutOfRangeCap_IsRefused_NotSilentlyClamped() + { + var service = new LocalDataService(_duckDb); + await SeedAsync(Truncate(DateTime.UtcNow).AddMinutes(-30), 100, 1000, 1000, intervalId: 1); + + var tooBig = await McpQueryTools.GetQueryStoreRegressions(service, _serverManager, ServerName, 24, null, 5000); + Assert.Contains("exceeds maximum of", tooBig, StringComparison.Ordinal); + Assert.Contains("1000", tooBig, StringComparison.Ordinal); + } + + private static JsonElement Root(string json) => JsonDocument.Parse(json).RootElement; + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedAsync( + DateTime collectionTime, long executions, long avgDurationUs, long avgCpuUs, long intervalId, long queryId = 1) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO query_store_stats + (collection_id, collection_time, server_id, server_name, database_name, query_id, plan_id, + execution_type_desc, execution_count, avg_duration_us, avg_cpu_time_us, avg_logical_io_reads, + runtime_stats_interval_id, query_text, last_execution_time) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12, $13, $14, $15)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = collectionTime }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = Db }); + cmd.Parameters.Add(new DuckDBParameter { Value = queryId }); + cmd.Parameters.Add(new DuckDBParameter { Value = 9L }); + cmd.Parameters.Add(new DuckDBParameter { Value = "Regular" }); + cmd.Parameters.Add(new DuckDBParameter { Value = executions }); + cmd.Parameters.Add(new DuckDBParameter { Value = avgDurationUs }); + cmd.Parameters.Add(new DuckDBParameter { Value = avgCpuUs }); + cmd.Parameters.Add(new DuckDBParameter { Value = 100L }); + cmd.Parameters.Add(new DuckDBParameter { Value = intervalId }); + cmd.Parameters.Add(new DuckDBParameter { Value = "SELECT * FROM dbo.Widgets" }); + cmd.Parameters.Add(new DuckDBParameter { Value = collectionTime }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/QueryStoreSliceRepairTests.cs b/Lite.Tests/QueryStoreSliceRepairTests.cs index d5a8cc512..dd6f5adac 100644 --- a/Lite.Tests/QueryStoreSliceRepairTests.cs +++ b/Lite.Tests/QueryStoreSliceRepairTests.cs @@ -344,12 +344,65 @@ public void EveryMutatingPhase_SitsInsideAWriteLockScope() /* And the REWRITE stays under the read lock: writing a temp sibling nothing can see yet must not take the UI's exclusivity for the whole of a large archive rewrite. */ - Assert.Contains("using (_duckDb.AcquireReadLock())", source, StringComparison.Ordinal); - var readLock = source.IndexOf("using (_duckDb.AcquireReadLock())", StringComparison.Ordinal); + Assert.Contains("using (_duckDb.AcquireReadLock(cancellationToken))", source, StringComparison.Ordinal); + var readLock = source.IndexOf("using (_duckDb.AcquireReadLock(cancellationToken))", StringComparison.Ordinal); var copyToTemp = source.IndexOf("COMPRESSION ZSTD)\";", StringComparison.Ordinal); Assert.True(copyToTemp > readLock, "the rewrite-to-temp belongs under the read lock, not the write lock"); } + /// + /// Source guard: no lock acquisition in this service is silent about cancellation (#2465). + /// + /// The sibling test above pins WHICH lock each phase takes. This one pins whether the phase can be + /// ABANDONED while it waits for it, which is a separate property and the one CA2016 was pointing at: every + /// method here holds a token, and every one of them takes the lock BEFORE it opens its connection, so a + /// no-arg AcquireReadLock() leaves the token stopped at the door — abandonable everywhere except + /// where the pass is actually stuck, which is the exact defect #2454 closed for the analysis pass. + /// + /// A source guard rather than a behavioral one because what is being pinned is a DECISION, not a + /// behavior: three sites forward the token and two decline it, and on the shipped caller + /// (MainWindow fires the repair un-awaited with no token) both spellings run identically today. + /// Nothing observable separates them, and that is precisely why they need pinning — a later edit that + /// "tidied" the two declines into forwards, or blanket-suppressed the rule, would cost nothing at runtime + /// and quietly convert five stated decisions back into an unknown. + /// + /// The declines are pinned WITH their reason, because a bare CancellationToken.None is the + /// oversight the analyzer complained about with an extra token typed in. + /// + [Fact] + public void EveryLockAcquisition_SaysWhetherItCanBeAbandoned() + { + var source = File.ReadAllText(SourcePath("Lite", "Services", "QueryStoreSliceRepairService.cs")); + + /* Not one silent acquisition left. This is the assertion that goes red on the pre-#2465 file. */ + Assert.Empty(Regex.Matches(source, @"_duckDb\.AcquireReadLock\(\)")); + + /* Three FORWARD it: the marker read, the survey, and the archive rewrite-to-temp. Each abandons into + a state the next launch reproduces for free, so there is nothing to protect by waiting. */ + Assert.Equal(3, Regex.Matches(source, @"_duckDb\.AcquireReadLock\(cancellationToken\)").Count); + + /* Two DECLINE it, and both are the marker write — the only thing here that records work already done + and irreversible, where abandoning costs the next launch the whole survey to learn nothing. */ + var declined = Regex.Matches(source, @"_duckDb\.AcquireReadLock\(CancellationToken\.None\)"); + Assert.Equal(2, declined.Count); + + foreach (Match site in declined) + { + var reason = source[Math.Max(0, site.Index - 1500)..site.Index]; + Assert.Contains("#2465", reason, StringComparison.Ordinal); + Assert.Contains("marker", reason, StringComparison.Ordinal); + } + + /* And each decline is WHOLE. A lock that will not be abandoned in front of a write that will be is + worse than either choice made consistently, so the open and both statements decline too. */ + Assert.Equal(2, Regex.Matches(source, @"await connection\.OpenAsync\(CancellationToken\.None\);").Count); + Assert.Equal(2, Regex.Matches(source, @"await MarkRepairedAsync\(connection, [^;]*CancellationToken\.None\);").Count); + + /* The write locks stay out of this: AcquireWriteLock has no token-taking overload, so there is no + decision to state at those two. Their timeout question is #2463's, not this test's. */ + Assert.Equal(2, Regex.Matches(source, @"_duckDb\.AcquireWriteLock\(\)").Count); + } + /* ─────────────────────────── helpers ─────────────────────────── */ /// Walks up from the test binary to the repo root so the pin works from any run directory. diff --git a/Lite.Tests/RemainingEmptyReadsToolTests.cs b/Lite.Tests/RemainingEmptyReadsToolTests.cs new file mode 100644 index 000000000..f4ab5491e --- /dev/null +++ b/Lite.Tests/RemainingEmptyReadsToolTests.cs @@ -0,0 +1,275 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitor.Notifications; +using PerformanceMonitor.Ui; +using PerformanceMonitorLite.Analysis; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// The last four reads on #2485's list, each of which needs a different shape of answer — which is why they +/// are here together rather than given one boilerplate envelope. +/// +/// +/// get_wait_types is windowed, so a LIMIT 1 probe against the same source separates a quiet +/// window from a server nothing was ever stored for. +/// get_memory_clerks is NOT windowed — it returns every clerk at MAX(collection_time) — so zero +/// rows is logically the same statement as zero rows in the table and a probe would agree with the read by +/// construction. What it owes the caller instead is that an empty clerk list is never a quiet period. +/// get_mute_rules is config, not collection: both empties are true negatives, but "no rule was +/// ever written" and "five rules that all lapsed" are different states and the second is a mute somebody +/// intended. +/// compare_analysis needs no probe at all — its own fact counts already hold the answer, and +/// the defect was that all-zero counters read as "nothing changed" rather than "nothing to compare". +/// +/// +/// Every message here is Darling's word for word; McpMissMessageParityPinTests holds that down +/// against both source trees. +/// +public sealed class RemainingEmptyReadsToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "RemainingEmptySrv"; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private DuckDBConnection? _seedConn; + private long _nextId = 850000; + + public RemainingEmptyReadsToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-remaining-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + + /* Derived, not stored -- seeding under a hardcoded id would write rows the tool looks past. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task WaitTypes_NeverCollected_IsNotAWindowToWiden() + { + var service = new LocalDataService(_duckDb); + + var root = JsonDocument.Parse( + await McpWaitTools.GetWaitTypes(service, _serverManager, ServerName, 4)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("EVER", text, StringComparison.Ordinal); + Assert.Contains("not an empty window", text, StringComparison.Ordinal); + + /* The second-cycle nuance is what makes this branch self-clearing on a freshly added server. */ + Assert.Contains("SECOND collection cycle", text, StringComparison.Ordinal); + Assert.DoesNotContain("widen hours_back", text, StringComparison.Ordinal); + } + + [Fact] + public async Task WaitTypes_CollectedButOutsideTheWindow_IsAWindowToWiden() + { + var service = new LocalDataService(_duckDb); + await SeedWaitAsync(DateTime.UtcNow.AddHours(-48), "CXPACKET", 5000L); + + var root = JsonDocument.Parse( + await McpWaitTools.GetWaitTypes(service, _serverManager, ServerName, 1)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("widen hours_back", text, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + } + + [Fact] + public async Task WaitTypes_InsideTheWindow_StillReturnsTheList() + { + var service = new LocalDataService(_duckDb); + await SeedWaitAsync(DateTime.UtcNow.AddMinutes(-10), "PAGEIOLATCH_SH", 900L); + + var root = JsonDocument.Parse( + await McpWaitTools.GetWaitTypes(service, _serverManager, ServerName, 4)).RootElement; + + Assert.False(root.TryGetProperty("status", out _)); + Assert.Equal(1, root.GetProperty("wait_types").GetArrayLength()); + } + + /// + /// One branch, on purpose. What is pinned is what the sentence REFUSES to imply: an empty clerk list is + /// not a quiet period and not a window that wants widening, because a live SQL Server always has clerks. + /// + [Fact] + public async Task MemoryClerks_EmptySnapshot_IsNeverDescribedAsAQuietPeriod() + { + var service = new LocalDataService(_duckDb); + + var root = JsonDocument.Parse( + await McpMemoryTools.GetMemoryClerks(service, _serverManager, ServerName)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("never a quiet period", text, StringComparison.Ordinal); + Assert.Contains("LATEST snapshot", text, StringComparison.Ordinal); + + /* No window, so no window to widen -- offering that would be advice that cannot work. */ + Assert.DoesNotContain("widen hours_back", text, StringComparison.Ordinal); + } + + [Fact] + public async Task MuteRules_NoneConfigured_SaysNothingIsSuppressedAnywhere() + { + var service = new MuteRuleService(new InMemoryMuteRuleStore(), new AppLoggerAdapter()); + + var root = JsonDocument.Parse(await McpAlertTools.GetMuteRules(service)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("No mute rules are configured", text, StringComparison.Ordinal); + Assert.Equal(0, root.GetProperty("hints").GetProperty("configured_count").GetInt32()); + } + + /// + /// The state the bare array could never express: rules exist, somebody meant them, and every one has + /// lapsed. Nothing is being suppressed either way — which is why both branches are "empty" — but a + /// caller auditing why a server looks quiet needs to know the rules are there. + /// + [Fact] + public async Task MuteRules_AllLapsed_IsADifferentAnswerFromNoneConfigured() + { + var store = new InMemoryMuteRuleStore(); + var service = new MuteRuleService(store, new AppLoggerAdapter()); + + await service.AddRuleAsync(new MuteRule + { + Id = "lapsed-1", + Enabled = true, + CreatedAtUtc = DateTime.UtcNow.AddDays(-7), + ExpiresAtUtc = DateTime.UtcNow.AddDays(-1), + ServerName = ServerName, + MetricName = "High CPU", + }); + + var root = JsonDocument.Parse(await McpAlertTools.GetMuteRules(service)).RootElement; + + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("disabled or expired", text, StringComparison.Ordinal); + Assert.Contains("enabled_only=false", text, StringComparison.Ordinal); + + /* And it must NOT reach for the none-configured sentence, which would be a flat lie here. */ + Assert.DoesNotContain("No mute rules are configured", text, StringComparison.Ordinal); + + var hints = root.GetProperty("hints"); + Assert.Equal(1, hints.GetProperty("configured_count").GetInt32()); + Assert.Equal(1, hints.GetProperty("excluded_by_filter").GetInt32()); + + /* enabled_only=false is the escape hatch the message names, so it has to actually work. */ + var all = JsonDocument.Parse(await McpAlertTools.GetMuteRules(service, enabled_only: false)).RootElement; + Assert.False(all.TryGetProperty("status", out _)); + Assert.Equal(1, all.GetProperty("total_count").GetInt32()); + } + + /// + /// All-zero counters and facts: [] read as "nothing changed" when the truth is "there was nothing + /// to compare" — opposite conclusions about the same server. No probe is needed: the comparison list is + /// the union of both windows' keys, so zero entries IS both fact sets being empty. + /// + [Fact] + public async Task CompareAnalysis_WithNoFactsInEitherWindow_SaysThereWasNothingToCompare() + { + var analysis = new AnalysisService(_duckDb) { MinimumDataHours = 0 }; + + var root = JsonDocument.Parse( + await McpAnalysisTools.CompareAnalysis(analysis, _serverManager, ServerName, 4, 28)).RootElement; + + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("EITHER window", text, StringComparison.Ordinal); + Assert.Contains("NOT a report that nothing changed", text, StringComparison.Ordinal); + + /* The windows it actually looked at, so the caller can check collection covered them. */ + var hints = root.GetProperty("hints"); + Assert.False(string.IsNullOrEmpty(hints.GetProperty("baseline_start").GetString())); + Assert.False(string.IsNullOrEmpty(hints.GetProperty("comparison_end").GetString())); + } + + /// In-memory store so the mute-rule branches are decided by the test, not by whatever rules a + /// shared store happens to hold. + private sealed class InMemoryMuteRuleStore : IMuteRuleStore + { + private readonly List _rules = new(); + + public Task> LoadAllAsync() => Task.FromResult>(_rules.ToList()); + public Task InsertAsync(MuteRule rule) { _rules.Add(rule); return Task.CompletedTask; } + public Task UpdateAsync(MuteRule rule) => Task.CompletedTask; + public Task SetEnabledAsync(string ruleId, bool enabled) => Task.CompletedTask; + public Task DeleteAsync(string ruleId) { _rules.RemoveAll(r => r.Id == ruleId); return Task.CompletedTask; } + public Task DeleteExpiredAsync(IReadOnlyList expiredIds) => Task.CompletedTask; + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedWaitAsync(DateTime collectionTimeUtc, string waitType, long deltaWaitMs) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO wait_stats + (collection_id, collection_time, server_id, server_name, wait_type, + waiting_tasks_count, wait_time_ms, signal_wait_time_ms, + delta_waiting_tasks, delta_wait_time_ms, delta_signal_wait_time_ms) +VALUES ($1, $2, $3, $4, $5, 0, 0, 0, $6, $7, 0)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = waitType }); + cmd.Parameters.Add(new DuckDBParameter { Value = 10L }); + cmd.Parameters.Add(new DuckDBParameter { Value = deltaWaitMs }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/RunningJobsCollectorDefinitionTests.cs b/Lite.Tests/RunningJobsCollectorDefinitionTests.cs index 660ca4aa7..82c1d5151 100644 --- a/Lite.Tests/RunningJobsCollectorDefinitionTests.cs +++ b/Lite.Tests/RunningJobsCollectorDefinitionTests.cs @@ -41,9 +41,12 @@ public void AppliesTo_SkipsAzureSqlDbRdsAndNoMsdb_ButCollectsOnPremAndManagedIns /* AWS RDS blocks msdb.dbo.syssessions (the join needs sysadmin, which RDS never grants) — it would raise "SELECT permission was denied" every cycle, so skip it there. */ Assert.False(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAwsRds = true })); - /* No msdb access → every table this reads lives in msdb; skip (gate collapsed from - IsCollectorSupported into the shared AppliesTo). */ - Assert.False(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); + /* NOT gated on msdb access (#2559). It is a GRANT rather than an engine capability, and it was + probed once and cached for the connection's life - so running the GRANT we advise did nothing + until a restart. This now attempts and fails into PERMISSIONS, which error 916 already maps to, + and CollectorHealthClassifier bands a never-permitted collector as NO_PERMISSIONS ahead of + FAILING, so it does not read as broken. */ + Assert.True(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo { HasMsdbAccess = false })); /* Managed Instance has Agent; on-prem collects too. */ Assert.True(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo { IsAzureManagedInstance = true })); Assert.True(RunningJobsCollector.Instance.AppliesTo(new CollectorTargetInfo())); diff --git a/Lite.Tests/ServerCollectionFreshnessTests.cs b/Lite.Tests/ServerCollectionFreshnessTests.cs new file mode 100644 index 000000000..fac9de3b3 --- /dev/null +++ b/Lite.Tests/ServerCollectionFreshnessTests.cs @@ -0,0 +1,382 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Windows.Media; +using PerformanceMonitor.Common; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2452: Lite's Overview card banded five metric rows and a status word, and none of them was collection +/// FRESHNESS. A server could go quiet and stay green — the connection check still passing, no collector +/// reporting an error, and the store taking no new rows for hours — while the card showed "Online" in green, +/// a neutral border, and a "Last Collect" timestamp rendered in the plain foreground brush, where a reading +/// from four hours ago looked exactly like one from four seconds ago. +/// +/// Freshness is a THIRD axis and these pins exist mostly to keep it one. The card already +/// answers two questions: did the last connection check succeed (IsOnline) and are any collectors +/// erroring (HasCollectorErrors). "Has anything landed lately" is neither of those, and the way this +/// goes wrong is well documented next door: #2429 found the Darling viewer says the same word, "Warning", for +/// a stale collection AND for a metric breach, with nothing on the card telling them apart — the card @ehaar +/// wrote in about in #2422. So and +/// are the load-bearing tests here; the +/// band arithmetic below is the shared classifier's and is pinned in PerformanceMonitor.Common. +/// +/// The thresholds are read, never restated. Every boundary case is built from +/// rather than from the literals 2 and 15, so a test cannot pass while +/// the card has quietly grown numbers of its own — which is the drift the shared class was created to end +/// (#1562). +/// +public class ServerCollectionFreshnessTests +{ + private static readonly DateTime Now = new(2026, 8, 21, 12, 0, 0, DateTimeKind.Utc); + + /// A card in the state #2452 describes: reachable, nothing erroring, every metric calm. The + /// only variable is how long ago the newest collection_log row landed. + private static ServerSummaryItem QuietServer(TimeSpan? sinceLastCollection) + { + var card = new ServerSummaryItem + { + DisplayName = "sql-01", + ServerId = 1, + IsOnline = true, + HasCollectorErrors = false, + CpuPercent = 4, + OtherProcessCpuPercent = 1, + MemoryMb = 8192, + BlockingCount = 0, + DeadlockCount = 0, + LastCollectionTime = sinceLastCollection.HasValue ? Now - sinceLastCollection.Value : null, + }; + + card.ApplyCollectionFreshness(Now); + return card; + } + + private static string Hex(SolidColorBrush brush) => brush.Color.ToString(System.Globalization.CultureInfo.InvariantCulture); + + private const string Green = "#FF81C784"; + private const string RowAmber = "#FFFFB74D"; + private const string Red = "#FFE57373"; + private const string StatusAmber = "#FFFFD54F"; + private const string NeutralBorder = "#FF2A2D35"; + private const string UnknownGrey = "#FF888888"; + + // ── The band itself, off the shared thresholds ──────────────────────────────────────────────────── + + [Fact] + public void ACollectionInsideTheCadenceIsFresh() + { + var card = QuietServer(ServerHealthThresholds.CollectorCadence); + + Assert.Equal(ServerFreshness.Fresh, card.CollectionFreshness); + Assert.False(card.CollectionIsNotFresh); + Assert.Equal(Green, Hex(card.LastCollectionBrush)); + } + + [Fact] + public void PastTwiceTheCadenceTheRowGoesAmberAndSaysStale() + { + var card = QuietServer(ServerHealthThresholds.StaleThreshold + TimeSpan.FromSeconds(1)); + + Assert.Equal(ServerFreshness.Stale, card.CollectionFreshness); + Assert.Equal(RowAmber, Hex(card.LastCollectionBrush)); + Assert.EndsWith(" (stale)", card.LastCollectionDisplay, StringComparison.Ordinal); + } + + [Fact] + public void PastTheOfflineThresholdTheRowGoesRedAndSaysStopped() + { + var card = QuietServer(ServerHealthThresholds.OfflineThreshold + TimeSpan.FromSeconds(1)); + + Assert.Equal(ServerFreshness.Offline, card.CollectionFreshness); + Assert.Equal(Red, Hex(card.LastCollectionBrush)); + Assert.EndsWith(" (stopped)", card.LastCollectionDisplay, StringComparison.Ordinal); + } + + /// Exactly ON a threshold is still the lower band — the shared classifier compares with a strict + /// >. Pinned because a card that re-derived the comparison is exactly how the two SKUs would + /// come to disagree by one tick, and nothing else would ever show it. + [Fact] + public void TheBoundariesAreTheSharedClassifiersAndAreInclusiveOfTheLowerBand() + { + Assert.Equal(ServerFreshness.Fresh, QuietServer(ServerHealthThresholds.StaleThreshold).CollectionFreshness); + Assert.Equal(ServerFreshness.Stale, QuietServer(ServerHealthThresholds.OfflineThreshold).CollectionFreshness); + } + + /// A server that has never been collected is queued, not dead. Amber, never the red the + /// stopped band gets — the viewer learned this from a 24-server field report that went chasing a + /// phantom scheduler bug (). + [Fact] + public void NeverCollectedIsAmberRatherThanRed() + { + var card = QuietServer(null); + + Assert.Equal(ServerFreshness.NeverCollected, card.CollectionFreshness); + Assert.Equal(RowAmber, Hex(card.LastCollectionBrush)); + Assert.Equal("Never", card.LastCollectionDisplay); + } + + /// An item nobody classified renders EXACTLY what dev's row rendered: the bare stamp, the + /// card's unknown grey, and no tooltip. Unreachable in the shipped app — every ServerSummaryItem is + /// built by GetServerSummaryAsync, which stamps the band — but a fixture or a new data path can produce + /// it, and a named "not classified" beats falling through onto a band that would be a guess. + [Fact] + public void AnUnclassifiedCardClaimsNothing() + { + var card = new ServerSummaryItem { LastCollectionTime = Now.AddHours(-4) }; + + Assert.Null(card.CollectionFreshness); + Assert.False(card.CollectionIsNotFresh); + Assert.Equal(UnknownGrey, Hex(card.LastCollectionBrush)); + Assert.Null(card.CollectionFreshnessTooltip); + Assert.DoesNotContain("(", card.LastCollectionDisplay, StringComparison.Ordinal); + } + + // ── The defect, and the conflation the fix must not create ──────────────────────────────────────── + + /// + /// The #2452 card itself. Everything the card used to answer still says the server is fine, because + /// those answers are still TRUE — the connection check really did succeed and no collector really is + /// erroring. What changes is that the card no longer looks calm: the border escalates and the row that + /// holds the evidence is red and says so. + /// + [Fact] + public void AServerThatWentQuietNoLongerLooksCalm() + { + var quiet = QuietServer(TimeSpan.FromHours(4)); + + /* Unchanged, and deliberately so: these two answer different questions and both are still yes. */ + Assert.Equal("Online", quiet.StatusDisplay); + Assert.Equal(Green, Hex(quiet.StatusBrush)); + + /* Changed: the card now shows the third axis. */ + Assert.Equal(StatusAmber, Hex(quiet.CardBorderBrush)); + Assert.Equal(Red, Hex(quiet.LastCollectionBrush)); + Assert.Contains("stopped", quiet.LastCollectionDisplay, StringComparison.Ordinal); + Assert.Contains("Collection has stopped", quiet.CollectionFreshnessTooltip!, StringComparison.Ordinal); + + /* And #2451's card tooltip picks it up through the gate pattern, which is what option 1 in the + issue meant by "let the tooltip name it, the way it names CPU and Blocking today". */ + Assert.Contains("Needs attention: Last collect ", quiet.StatusTooltip, StringComparison.Ordinal); + } + + /// + /// The invariant #2451 established for the metric rows, extended to this one: the tooltip names the Last + /// Collect row exactly when that row is not green. Asserted against the SHIPPED brush rather than against + /// a copy of the thresholds, so the clause and the colour cannot drift apart — a tooltip that named a + /// green row, or stayed silent about a red one, is the single thing this property exists to prevent. + /// + [Fact] + public void TheTooltipNamesTheCollectionRowExactlyWhenItIsNotGreen() + { + foreach (var age in new TimeSpan?[] + { + TimeSpan.FromSeconds(20), + ServerHealthThresholds.StaleThreshold + TimeSpan.FromSeconds(1), + TimeSpan.FromHours(4), + null, + }) + { + var card = QuietServer(age); + + Assert.Equal( + Hex(card.LastCollectionBrush) != Green, + card.StatusTooltip.Contains("Last collect ", StringComparison.Ordinal)); + } + } + + /// + /// Collection goes FIRST in the reason, ahead of the metrics. It is the row that says whether the other + /// four can be believed: a card whose collection stopped four hours ago is showing four-hour-old CPU, + /// and meeting "CPU 4%" before learning that is the wrong order to be told the two facts in. + /// + [Fact] + public void TheReasonLeadsWithCollectionWhenBothAxesAreInTrouble() + { + var card = QuietServer(TimeSpan.FromHours(4)); + card.CpuPercent = 96; + card.OtherProcessCpuPercent = 0; + card.DeadlockCount = 2; + + Assert.StartsWith("Last collect ", card.StatusReason, StringComparison.Ordinal); + Assert.Contains("CPU ", card.StatusReason, StringComparison.Ordinal); + Assert.Contains("Deadlocks ", card.StatusReason, StringComparison.Ordinal); + } + + /// + /// The same card while collection is current is the control: neutral border, green row. Without this the + /// test above would pass on a card that had simply been painted amber for everyone. + /// + [Fact] + public void ACalmCardWithCurrentCollectionStaysNeutral() + { + var healthy = QuietServer(TimeSpan.FromSeconds(20)); + + Assert.Equal(NeutralBorder, Hex(healthy.CardBorderBrush)); + Assert.Equal(Green, Hex(healthy.LastCollectionBrush)); + } + + /// + /// The anti-conflation pin, and the reason option 2 in #2452 was turned down. Lite's status word is a + /// CONNECTION word — IsOnline comes from a live check that succeeds whether or not anything is + /// being collected, and "Warning" there already means collectors are erroring. If freshness were folded + /// into it, one amber word would mean two unrelated failures, which is precisely the viewer defect + /// #2429 found and #2422 reported. So a stale card's word and colour are byte-identical to a fresh + /// card's, and every difference is on the row that owns the axis. + /// + [Fact] + public void TheStatusWordIsNotToldAboutFreshness() + { + var fresh = QuietServer(TimeSpan.FromSeconds(20)); + var stale = QuietServer(ServerHealthThresholds.StaleThreshold + TimeSpan.FromSeconds(1)); + var stopped = QuietServer(TimeSpan.FromHours(4)); + + Assert.Equal(fresh.StatusDisplay, stale.StatusDisplay); + Assert.Equal(fresh.StatusDisplay, stopped.StatusDisplay); + Assert.Equal(Hex(fresh.StatusBrush), Hex(stale.StatusBrush)); + Assert.Equal(Hex(fresh.StatusBrush), Hex(stopped.StatusBrush)); + } + + /// + /// Both directions of the split. A stale collection must not be describable as a metric problem, and a + /// metric breach must not be describable as a collection problem — the card carries both at once often + /// enough that only an adversarial pair proves the two vocabularies stayed apart. + /// + [Fact] + public void EachAxisIsNamedOnTheCardWithoutBorrowingTheOthersWords() + { + var stale = QuietServer(ServerHealthThresholds.StaleThreshold + TimeSpan.FromSeconds(1)); + var tooltip = stale.CollectionFreshnessTooltip!; + + Assert.Contains("Collection is stale", tooltip, StringComparison.Ordinal); + Assert.Contains("not about the server's metrics", tooltip, StringComparison.Ordinal); + /* The one word that means "collectors are erroring" on this card. Freshness may never claim it. */ + Assert.DoesNotContain("Warning", tooltip, StringComparison.Ordinal); + + /* A busy server whose collection is perfectly current: the border escalates for the METRIC, and the + collection row keeps saying collection is fine. */ + var busy = QuietServer(TimeSpan.FromSeconds(20)); + busy.CpuPercent = 96; + busy.OtherProcessCpuPercent = 0; + + Assert.Equal(RowAmber, Hex(busy.CardBorderBrush)); + Assert.Equal(Green, Hex(busy.LastCollectionBrush)); + Assert.Contains("Collection is current", busy.CollectionFreshnessTooltip!, StringComparison.Ordinal); + Assert.DoesNotContain("Last collect ", busy.StatusTooltip, StringComparison.Ordinal); + } + + /// The tooltip quotes the shared thresholds instead of restating them, so the sentence cannot + /// come to disagree with the band it is explaining. + [Fact] + public void TheTooltipQuotesTheSharedThresholds() + { + var stale = QuietServer(ServerHealthThresholds.StaleThreshold + TimeSpan.FromSeconds(1)); + var stopped = QuietServer(ServerHealthThresholds.OfflineThreshold + TimeSpan.FromSeconds(1)); + + Assert.Contains( + $"{ServerHealthThresholds.StaleThreshold.TotalMinutes:0.#} minute", + stale.CollectionFreshnessTooltip!, + StringComparison.Ordinal); + Assert.Contains( + $"{ServerHealthThresholds.OfflineThreshold.TotalMinutes:0.#} minute", + stopped.CollectionFreshnessTooltip!, + StringComparison.Ordinal); + } + + /// The band is a pure function of (last collection, now): the same card classified against two + /// clocks lands in two bands. That is what makes the stamp testable without a store, and what keeps the + /// clock out of the card's property getters. + [Fact] + public void TheBandIsPureOverTheClockItIsGiven() + { + var card = new ServerSummaryItem { LastCollectionTime = Now }; + + card.ApplyCollectionFreshness(Now); + Assert.Equal(ServerFreshness.Fresh, card.CollectionFreshness); + + card.ApplyCollectionFreshness(Now + ServerHealthThresholds.OfflineThreshold + TimeSpan.FromMinutes(1)); + Assert.Equal(ServerFreshness.Offline, card.CollectionFreshness); + } + + // ── Wiring: the half that lives in XAML, where no assertion about a C# object can reach it ───────── + + /// + /// The Last Collect row has to actually BE bound to the band, and removing either binding compiles + /// perfectly clean — the properties would simply go unread and the row would render exactly as it did + /// on dev. Background="Transparent" is checked with them because a TextBlock with a null + /// Background hit-tests on its rendered glyphs alone, so without it the tooltip appears over the letters + /// and nowhere in the space around them (#2429 found this on the viewer's status line). + /// + [Fact] + public void TheCardsLastCollectRowIsBoundToTheBand() + { + var xaml = ParitySource.ReadFile("Lite/MainWindow.xaml"); + var row = RowFour(xaml); + + Assert.Contains("{Binding LastCollectionBrush}", row, StringComparison.Ordinal); + Assert.Contains("{Binding CollectionFreshnessTooltip}", row, StringComparison.Ordinal); + Assert.DoesNotContain("Foreground=\"{DynamicResource ForegroundBrush}\"", row, StringComparison.Ordinal); + Assert.Equal(2, CountOccurrences(row, "Background=\"Transparent\"")); + Assert.Equal(2, CountOccurrences(row, "ToolTip=\"{Binding CollectionFreshnessTooltip}\"")); + } + + /// + /// The band is stamped at the ONE place a ServerSummaryItem is built, which is what stops the Overview + /// and the two MCP reads from getting different answers about the same server. It is pinned in the + /// source because deleting that single line compiles perfectly clean and silently returns every card to + /// dev's behaviour — no band, so a grey row and a neutral border. A property that nothing sets is the + /// failure mode this whole issue is about, one level up. + /// + [Fact] + public void EveryCardIsStampedWithItsBandWhereItIsBuilt() + { + var source = ParitySource.ReadFile("Lite/Services/LocalDataService.Overview.cs"); + + Assert.Equal(1, CountOccurrences(source, "new ServerSummaryItem")); + Assert.Equal(1, CountOccurrences(source, "ApplyCollectionFreshness(DateTime.UtcNow)")); + + /* And the stamp is after the object is built: banding an object that has already been handed back + would band nothing, and the ordering is the half a count cannot see. */ + Assert.True( + source.IndexOf("new ServerSummaryItem", StringComparison.Ordinal) + < source.IndexOf("ApplyCollectionFreshness(DateTime.UtcNow)", StringComparison.Ordinal), + "The freshness stamp no longer follows the summary it is meant to band."); + } + + /// The two Grid.Row="4" TextBlocks of the Overview card template — the "Last Collect:" label + /// and its value. Sliced by the marker rather than by line numbers so an edit above the row does not + /// silently move the window this test reads. + private static string RowFour(string xaml) + { + const string marker = "Text=\"Last Collect:\""; + var start = xaml.IndexOf(marker, StringComparison.Ordinal); + Assert.True(start >= 0, "The Overview card no longer has a Last Collect row."); + Assert.Equal(start, xaml.LastIndexOf(marker, StringComparison.Ordinal)); + + var lineStart = xaml.LastIndexOf('<', start); + var end = xaml.IndexOf("", start, StringComparison.Ordinal); + Assert.True(end > lineStart, "The Last Collect row is no longer inside the card's metric Grid."); + return xaml[lineStart..end]; + } + + private static int CountOccurrences(string haystack, string needle) + { + var count = 0; + for (var i = haystack.IndexOf(needle, StringComparison.Ordinal); i >= 0; + i = haystack.IndexOf(needle, i + needle.Length, StringComparison.Ordinal)) + { + count++; + } + + return count; + } +} diff --git a/Lite.Tests/SettingsFileGuardTests.cs b/Lite.Tests/SettingsFileGuardTests.cs new file mode 100644 index 000000000..a5485cbf1 --- /dev/null +++ b/Lite.Tests/SettingsFileGuardTests.cs @@ -0,0 +1,405 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Text.RegularExpressions; +using PerformanceMonitor.Common; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2425. One trailing comma in a hand-edited settings.json used to reset all eighty-eight Lite settings +/// at once, say nothing about it anywhere, and then leave the file exposed to the next whole-document +/// rewrite. Two properties are pinned here, and the second is the one that turns an annoyance into data +/// loss. +/// +/// Absent is not unreadable. The old loaders had a single bare catch, so a first run +/// with no file and a corrupt file took the same silent path to defaults. That is why the silence looked +/// reasonable in the code and was wrong in practice — half the traffic through it really was fine. The +/// split has to hold in BOTH directions: a missing file must stay completely silent, because a first-run +/// warning is pure noise, and a present-but-broken file must never be. +/// +/// Nothing overwrites what it could not read. Every Save in Lite rewrites the whole document, +/// including saves nobody thinks of as saves, so a writer that starts from a fresh object after a failed +/// parse destroys the only record of the user's real configuration. The copy has to exist BEFORE the write, +/// and the original has to survive making it. +/// +public sealed class SettingsFileGuardTests +{ + /// A settings.json a user would recognize: real keys, one syntax error. + private const string TrailingComma = @"{ + ""alerts_enabled"": true, + ""alert_cpu_threshold"": 91, +}"; + + private static string NewTempDir(string tag) + { + var dir = Path.Combine(Path.GetTempPath(), $"pmlite_{tag}_{Guid.NewGuid():N}"); + Directory.CreateDirectory(dir); + return dir; + } + + private static string WriteSettings(string dir, string content) + { + var path = Path.Combine(dir, "settings.json"); + File.WriteAllText(path, content); + return path; + } + + /// + /// The legitimate first run. No file, no problem, nothing to say — and specifically no Problem string, + /// because anything non-null there becomes a log line and a dialog for a user who has done nothing + /// wrong. + /// + [Fact] + public void Read_IsSilentlyAbsent_WhenThereIsNoFile() + { + var dir = NewTempDir("absent"); + try + { + var read = SettingsFileGuard.Read(Path.Combine(dir, "settings.json")); + + Assert.Equal(SettingsFileState.Absent, read.State); + Assert.Null(read.Problem); + Assert.Null(read.Root); + } + finally + { + Directory.Delete(dir, true); + } + } + + [Fact] + public void Read_IsReadable_ForAnOrdinarySettingsFile() + { + var dir = NewTempDir("ok"); + try + { + var path = WriteSettings(dir, @"{""alerts_enabled"":true,""alert_cpu_threshold"":91}"); + + var read = SettingsFileGuard.Read(path); + + Assert.Equal(SettingsFileState.Readable, read.State); + Assert.Null(read.Problem); + Assert.NotNull(read.Root); + Assert.True(read.Root!["alerts_enabled"]!.GetValue()); + Assert.NotNull(read.Text); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The headline case, and the reason the diagnostic carries a position rather than the word "failed": + /// "settings.json is broken" sends someone looking through a file they already believe is correct, + /// while "line 4" is a minute's work. The System.Text.Json " Path: $ | LineNumber: ..." tail is cut + /// because the same facts are already in the sentence, in the form a person reads them. + /// + [Fact] + public void Read_ReportsTheLineAndPosition_ForATrailingComma() + { + var dir = NewTempDir("comma"); + try + { + var path = WriteSettings(dir, TrailingComma); + + var read = SettingsFileGuard.Read(path); + + Assert.Equal(SettingsFileState.Unreadable, read.State); + Assert.NotNull(read.Problem); + Assert.Contains("line 4", read.Problem!, StringComparison.Ordinal); + Assert.Contains("position", read.Problem!, StringComparison.Ordinal); + Assert.DoesNotContain(" Path: ", read.Problem!, StringComparison.Ordinal); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// A file holding the JSON literal null parses fine and is still unusable. It gets its own case because + /// it is the shape that slipped past the writers' old JsonNode.Parse(json) ?? new JsonObject() + /// read: no exception, no warning, and the very next Save replaced the document with a fresh one + /// holding a single key. + /// + [Fact] + public void Read_IsUnreadable_ForARootThatIsNotAnObject() + { + var dir = NewTempDir("notobject"); + try + { + Assert.Equal(SettingsFileState.Unreadable, SettingsFileGuard.Read(WriteSettings(dir, "null")).State); + + var array = SettingsFileGuard.Read(WriteSettings(dir, "[1,2,3]")); + Assert.Equal(SettingsFileState.Unreadable, array.State); + Assert.Contains("array", array.Problem!, StringComparison.Ordinal); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// An empty file has nothing to preserve but plenty to explain — settings that were there yesterday are + /// gone today, and a write interrupted by a full disk or a crash is the likeliest reason. Reported, not + /// quietly folded into "absent". + /// + [Fact] + public void Read_IsUnreadable_ForAnEmptyFile() + { + var dir = NewTempDir("empty"); + try + { + var read = SettingsFileGuard.Read(WriteSettings(dir, " \n")); + + Assert.Equal(SettingsFileState.Unreadable, read.State); + Assert.NotNull(read.Problem); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The path that has always worked keeps working: a readable file is MERGED into, so keys the writer + /// never mentions — hand-edited ones with no UI, most of all — survive the save. + /// + [Fact] + public void RootForWrite_MergesIntoTheExistingDocument_WhenReadable() + { + var dir = NewTempDir("merge"); + try + { + var path = WriteSettings(dir, @"{""check_for_updates_on_startup"":false,""alert_cpu_threshold"":91}"); + + var forWrite = SettingsFileGuard.RootForWrite(path, DateTime.Now); + + Assert.Null(forWrite.Problem); + Assert.Null(forWrite.QuarantinedTo); + Assert.False(forWrite.Root["check_for_updates_on_startup"]!.GetValue()); + Assert.Equal(91, forWrite.Root["alert_cpu_threshold"]!.GetValue()); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The first run again, from the writer's side: no file means a fresh document and still no diagnostic. + /// A quarantine copy here would be litter, and a warning would be a lie. + /// + [Fact] + public void RootForWrite_StartsFreshAndSilent_WhenTheFileIsAbsent() + { + var dir = NewTempDir("firstwrite"); + try + { + var forWrite = SettingsFileGuard.RootForWrite(Path.Combine(dir, "settings.json"), DateTime.Now); + + Assert.Null(forWrite.Problem); + Assert.Null(forWrite.QuarantinedTo); + Assert.Empty(forWrite.Root); + Assert.Empty(Directory.GetFiles(dir)); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The data-loss case. An unreadable file is copied aside BEFORE the caller is handed a document to + /// write, the copy holds the original bytes verbatim, and the original is still where it was — a move + /// would leave a hole if the write that follows also failed. + /// + [Fact] + public void RootForWrite_CopiesTheUnreadableFileAside_BeforeHandingBackAFreshDocument() + { + var dir = NewTempDir("quarantine"); + try + { + var path = WriteSettings(dir, TrailingComma); + + var forWrite = SettingsFileGuard.RootForWrite(path, new DateTime(2026, 8, 21, 14, 5, 2, DateTimeKind.Local)); + + Assert.NotNull(forWrite.Problem); + Assert.NotNull(forWrite.QuarantinedTo); + Assert.Equal(path + ".unreadable-20260821-140502", forWrite.QuarantinedTo); + Assert.Equal(TrailingComma, File.ReadAllText(forWrite.QuarantinedTo!)); + + /* Copied, not moved: the caller's write can still fail. */ + Assert.Equal(TrailingComma, File.ReadAllText(path)); + + /* And what the caller writes must not carry anything from the file nobody could read. */ + Assert.Empty(forWrite.Root); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// Two unreadable files quarantined inside the same second must not collide, because the collision + /// would destroy exactly the bytes the first copy was made to keep. Same timestamp, twice, deliberately. + /// + [Fact] + public void Quarantine_DoesNotOverwriteAnEarlierCopyFromTheSameSecond() + { + var dir = NewTempDir("collide"); + try + { + var path = WriteSettings(dir, TrailingComma); + var stamp = new DateTime(2026, 8, 21, 14, 5, 2, DateTimeKind.Local); + + var first = SettingsFileGuard.Quarantine(path, stamp); + File.WriteAllText(path, "{ still not json"); + var second = SettingsFileGuard.Quarantine(path, stamp); + + Assert.NotNull(first); + Assert.NotNull(second); + Assert.NotEqual(first, second); + Assert.Equal(TrailingComma, File.ReadAllText(first!)); + Assert.Equal("{ still not json", File.ReadAllText(second!)); + } + finally + { + Directory.Delete(dir, true); + } + } +} + +/// +/// The category guard behind #2425 and #2433, rather than the instance. The defect was never "WriteSetting +/// is wrong" — it was that separate methods each rolled their own read of settings.json in front of their +/// own whole-document rewrite, so the safety of a Save depended on which method you happened to be in, and +/// so did whether anyone could find out afterwards that it had failed. +/// +/// #2425 answered the first half by making every one of those reads the quarantining one. #2433 +/// answered the second by removing the reads and the writes: the Settings window opens the document once, +/// hands it to all ten of its writers, and writes it once. So the invariant this guard pins is now +/// stronger and simpler than counting reads against writes per file — settings.json has exactly ONE +/// writer in the whole of Lite, and that writer takes the quarantining read. +/// +/// Source-parsing because the invariant is a wiring one. A behavioral test can prove the guard works +/// and still not notice a caller that never asks it. +/// +public sealed class SettingsWriterQuarantineWiringTests +{ + [Fact] + public void SettingsJson_HasExactlyOneWriterInAllOfLite() + { + var writers = new List(); + var described = new List(); + var total = 0; + + foreach (var file in LiteSourceFiles()) + { + var rewrites = Regex.Matches(WithoutComments(File.ReadAllText(file)), @"File\.WriteAllText\(\s*settingsPath").Count; + if (rewrites == 0) + { + continue; + } + + total += rewrites; + writers.Add(file); + described.Add($"{Path.GetFileName(file)} ({rewrites})"); + } + + Assert.True(total > 0, + "No settings.json rewrite found anywhere in Lite — this guard's anchor moved and it is testing nothing."); + Assert.True(total == 1, + $"{total} whole-document rewrite(s) of settings.json across {writers.Count} file(s): " + + string.Join(", ", described) + ". A second writer brings back both defects at once — a rewrite " + + "that reads the file itself can replace an unparseable settings.json without copying it aside " + + "(#2425), and a save split across several writes has no single honest answer to report (#2433)."); + Assert.EndsWith("App.xaml.cs", writers[0], StringComparison.Ordinal); + } + + [Fact] + public void TheOneWriter_TakesTheQuarantiningRead() + { + var source = File.ReadAllText(FindRepoFile(Path.Combine("Lite", "App.xaml.cs"))); + + Assert.Matches(new Regex(@"[=(,]\s*SettingsRootForWrite\(\)"), source); + } + + /// + /// The exact shape that made a non-object root a silent total overwrite: Parse returns null for the JSON + /// literal null, the null-coalesce reads that as "no file", and the save replaces the document. Banned + /// across all of Lite so it cannot be reintroduced by copy-paste from a sibling writer. + /// + [Fact] + public void NoWriter_FallsBackToAFreshDocumentOnAFailedParse() + { + var banned = new Regex(@"JsonNode\.Parse\([^;]*\)\s*\?\?\s*new JsonObject\(\)"); + + foreach (var file in LiteSourceFiles()) + { + Assert.DoesNotMatch(banned, WithoutComments(File.ReadAllText(file))); + } + } + + /* Both scans run on code with the comments removed, and SettingsFileGuard is why: the one place that + explains WHY `JsonNode.Parse(json) ?? new JsonObject()` is banned has to quote it to explain it, and + a scanner that cannot tell a comment from code reads the explanation as the offence. Same trap as + #2418's key extractor, one file over. */ + private static string WithoutComments(string source) + { + var withoutBlocks = Regex.Replace(source, @"/\*.*?\*/", "", RegexOptions.Singleline); + return Regex.Replace(withoutBlocks, @"//[^\r\n]*", ""); + } + + private static IEnumerable LiteSourceFiles() => + Directory.EnumerateFiles(FindRepoDirectory("Lite"), "*.cs", SearchOption.AllDirectories) + .Where(f => !f.Contains($"{Path.DirectorySeparatorChar}obj{Path.DirectorySeparatorChar}", StringComparison.Ordinal) + && !f.Contains($"{Path.DirectorySeparatorChar}bin{Path.DirectorySeparatorChar}", StringComparison.Ordinal)) + .OrderBy(f => f, StringComparer.Ordinal); + + private static string FindRepoFile(string relativePath) + { + var dir = AppContext.BaseDirectory; + for (var i = 0; i < 8 && dir is not null; i++) + { + var candidate = Path.Combine(dir, relativePath); + if (File.Exists(candidate)) + { + return candidate; + } + dir = Path.GetDirectoryName(dir); + } + + throw new FileNotFoundException($"Could not locate {relativePath} walking up from {AppContext.BaseDirectory}"); + } + + private static string FindRepoDirectory(string relativePath) + { + var dir = AppContext.BaseDirectory; + for (var i = 0; i < 8 && dir is not null; i++) + { + var candidate = Path.Combine(dir, relativePath); + if (Directory.Exists(candidate) && File.Exists(Path.Combine(dir, "PerformanceMonitor.sln"))) + { + return candidate; + } + dir = Path.GetDirectoryName(dir); + } + + throw new DirectoryNotFoundException($"Could not locate {relativePath} walking up from {AppContext.BaseDirectory}"); + } +} diff --git a/Lite.Tests/SettingsSampleTests.cs b/Lite.Tests/SettingsSampleTests.cs new file mode 100644 index 000000000..dbd6d920b --- /dev/null +++ b/Lite.Tests/SettingsSampleTests.cs @@ -0,0 +1,322 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Text.Json; +using System.Text.RegularExpressions; +using PerformanceMonitor.Common; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Keeps Lite\config\settings.sample.json honest against the code that actually reads +/// settings.json, in both directions: every key a loader reads is documented, and every key the +/// sample documents is really read. +/// +/// Why this exists. The file this replaces (#2418) shipped nowhere, seeded nothing and +/// was never read, so nothing could tell it had gone stale — and it had: four of its eight keys were +/// read by no code anywhere in the repo, including "theme", which the loader has always spelled +/// color_theme. A stale reference is worse than none, because it is what someone finds when they +/// go looking for what settings.json can hold, and it teaches wrong keys with nothing to signal that it +/// is wrong. The only thing that makes a reference file worth keeping is a check that fails when it +/// drifts, so this is that check. +/// +/// Derived from the shipped source, not a copy. The key list is regexed out of the real +/// App.xaml.cs and Mcp\McpSettings.cs, copied beside the test binary by the csproj. A +/// hand-maintained list here would be a third thing to keep in sync and would rot the same way the +/// sample did. +/// +/// Scoping. Each TryGetProperty is attributed to the most recent *.json +/// string literal above it, which is the shape every loader in these files has: open the file, then +/// read keys out of it. That means a loader added later for a DIFFERENT json file is excluded +/// automatically instead of needing an exemption — which matters, because App.xaml.cs already reads +/// servers.json, ignored_wait_types.json and collection_schedule.json elsewhere. +/// +public sealed class SettingsSampleTests +{ + /// + /// The two source files that read settings.json. Both are copied to Fixtures\ by the csproj. + /// If a THIRD loader appears, add it here — a reader this list does not know about is a key the + /// sample can omit forever without failing anything. + /// + private static readonly string[] ReaderSources = { "App.xaml.cs", "McpSettings.cs" }; + + /// + /// Keys deliberately documented in the sample that no loader reads. EMPTY, and it should stay that + /// way: a key with no reader is the exact defect this test exists to catch. The seam is here so that + /// adding one has to be a deliberate, commented act rather than a silent omission — the old file's + /// dead "theme" is discussed in a sample COMMENT, which carries the warning without claiming + /// to be a live key. + /// + private static readonly HashSet SampleOnlyKeys = new(StringComparer.Ordinal); + + /// + /// Keys a loader reads that the sample deliberately leaves undocumented. EMPTY. If a key is ever + /// too dangerous to publish, exempt it here WITH the reason — do not quietly drop it from the + /// sample, because an undocumented knob is how #2418 started. + /// + private static readonly HashSet LoaderOnlyKeys = new(StringComparer.Ordinal); + + [Fact] + public void Sample_DocumentsEveryKeyTheLoadersRead() + { + var read = ReadKeysFromLoaders(); + var documented = SampleKeys(); + + var missing = UndocumentedKeys(read, documented); + + Assert.True( + missing.Count == 0, + "settings.json keys read by Lite but absent from config\\settings.sample.json: " + + string.Join(", ", missing.Select(k => $"{k} (in {read[k]})")) + + ". Document them in the sample, or exempt them in LoaderOnlyKeys with a reason."); + } + + [Fact] + public void Sample_DocumentsNoKeyTheLoadersIgnore() + { + var read = ReadKeysFromLoaders(); + var documented = SampleKeys(); + + var dead = documented + .Where(k => !read.ContainsKey(k) && !SampleOnlyKeys.Contains(k)) + .OrderBy(k => k, StringComparer.Ordinal) + .ToList(); + + Assert.True( + dead.Count == 0, + "config\\settings.sample.json documents keys nothing reads: " + string.Join(", ", dead) + + ". Either the loader lost them or the sample invented them; a reader who copies one " + + "gets a setting that silently does nothing."); + } + + [Fact] + public void Sample_ParsesAsCommentedJson_WithNoDuplicateKeys() + { + using var doc = ParseSample(); + + var seen = new HashSet(StringComparer.Ordinal); + var duplicates = doc.RootElement + .EnumerateObject() + .Where(p => !seen.Add(p.Name)) + .Select(p => p.Name) + .ToList(); + + Assert.True( + duplicates.Count == 0, + "config\\settings.sample.json repeats keys: " + string.Join(", ", duplicates) + + ". A duplicate silently documents two different defaults for one setting."); + } + + /// + /// Guards the guard: the regex above is the whole test, so a restructured loader that stops matching + /// would make BOTH symmetry checks pass on an empty set. Lite reads dozens of keys and always will; + /// the floor is deliberately far below today's count so it pins "the extraction still works" rather + /// than becoming a second thing to bump on every new setting. + /// + [Fact] + public void KeyExtraction_StillFindsTheLoaders() + { + var read = ReadKeysFromLoaders(); + + Assert.True( + read.Count >= 50, + $"Only {read.Count} settings.json keys were found across {string.Join(", ", ReaderSources)}. " + + "The loaders were restructured and the extraction no longer sees them, which would make " + + "the symmetry tests pass vacuously."); + + /* Anchors: one no-UI key (the reason the sample is kept at all) and one from each source file, + so a fixture that silently stopped being copied fails here rather than passing on the other. */ + Assert.Contains("analysis_timeout_seconds", read.Keys); + Assert.Contains("alerts_enabled", read.Keys); + Assert.Contains("mcp_port", read.Keys); + } + + /// + /// Alternation, evaluated left to right in one pass, so the "which file is open" state and the key hits + /// stay in source order. + /// The TryGetProperty half keys on a METHOD NAME, which is the constraint every refactor of + /// the loaders runs into — see TheReadHelpersName_IsWhatThisExtractionKeysOn for why the #2444 + /// read helper is spelled the way it is. + /// + private static readonly Regex Scanner = new( + "\"(?[A-Za-z0-9_.\\-]+\\.json)\"|TryGetProperty\\(\"(?[^\"]+)\"", + RegexOptions.CultureInvariant); + + /// + /// The self-test #2444 owed this guard. The loaders' reads moved onto a helper + /// (), and the reason they still work is that the helper's method carries the + /// same NAME the extraction keys on — which is a fact about a string, held together by nothing the + /// compiler checks. So it is checked here, against a source this test writes, in both shapes: the one the + /// loaders used before #2444 and the one they use now. + /// + /// And it does not stop at "the keys are found". A guard that extracts keys but no longer COMPARES + /// them is the same vacuous green as one that extracts nothing, so this runs the real + /// over a sample that is deliberately missing one key and asserts that it + /// is named. That is the property the two symmetry tests exist for, proved on a fixture rather than on + /// the tree — where it can only ever be proved by breaking the tree. + /// + [Fact] + public void KeyExtraction_SeesBothReaderShapes_AndStillCatchesAnUndocumentedKey() + { + const string source = """ + var settings = SettingsFileGuard.Read(Path.Combine(configDirectory, "settings.json")); + if (root.TryGetProperty("old_shape_key", out var a)) A = a.GetBoolean(); + if (read.TryGetProperty("new_shape_key", out var b)) B = b.Bool(B); + if (read.TryGetProperty("undocumented_key", out var c)) C = c.Int(C); + """; + + var found = new Dictionary(StringComparer.Ordinal); + ExtractKeys(source, "synthetic.cs", found); + + Assert.Contains("old_shape_key", found.Keys); + Assert.Contains("new_shape_key", found.Keys); + + /* The guard still bites: one key is left out of the "sample", and it is the one named. */ + var documented = new HashSet(StringComparer.Ordinal) { "old_shape_key", "new_shape_key" }; + Assert.Equal(new[] { "undocumented_key" }, UndocumentedKeys(found, documented)); + + /* ...and stays silent when the sample really does document everything. */ + documented.Add("undocumented_key"); + Assert.Empty(UndocumentedKeys(found, documented)); + + /* The trap itself, so the paragraph above is a measurement rather than a warning. This is the shape + #2428 tried — a read helper that takes the key as an ordinary argument — and every key written that + way is invisible here. That is what makes the helper's NAME load-bearing rather than incidental. */ + const string hidden = """ + var settings = SettingsFileGuard.Read(Path.Combine(configDirectory, "settings.json")); + ReadInt(root, "invisible_key", ref X); + """; + + var missed = new Dictionary(StringComparer.Ordinal); + ExtractKeys(hidden, "synthetic.cs", missed); + Assert.Empty(missed); + } + + /// + /// The other half of the scoping contract, pinned for the same reason: a read attributed to the WRONG + /// file would make the sample look as though it were missing keys it has no business documenting. Each + /// hit belongs to the most recent *.json literal above it, which is how App.xaml.cs's + /// servers.json and collection_schedule.json loaders stay out of this set without needing an exemption. + /// + [Fact] + public void KeyExtraction_AttributesEachReadToTheFileOpenedAboveIt() + { + const string source = """ + var servers = Path.Combine(dir, "servers.json"); + if (root.TryGetProperty("not_a_settings_key", out var a)) A = a.GetBoolean(); + var settings = SettingsFileGuard.Read(Path.Combine(dir, "settings.json")); + if (read.TryGetProperty("a_settings_key", out var b)) B = b.Bool(B); + """; + + var found = new Dictionary(StringComparer.Ordinal); + ExtractKeys(source, "synthetic.cs", found); + + Assert.Equal(new[] { "a_settings_key" }, found.Keys.ToArray()); + } + + /// + /// Names the dependency this guard acquired in #2444 so a rename fails HERE, with the reason, rather + /// than as an unexplained collapse of the key count somewhere else. + /// + /// SettingsReader's read method is called TryGetProperty because that is the literal + /// this extraction matches. Spell it TryRead and all eighty-seven of Lite's keys vanish from the + /// extracted set, both symmetry tests pass on what is left, and settings.sample.json is free to drift + /// exactly the way #2418 was filed about. PR #2428 hit this and had to read back through + /// JsonDocument; #2444 kept the name instead, and this is the note that says so out loud. + /// + [Fact] + public void TheReadHelpersName_IsWhatThisExtractionKeysOn() + { + Assert.Contains("TryGetProperty", Scanner.ToString(), StringComparison.Ordinal); + + Assert.True( + typeof(SettingsReader).GetMethod("TryGetProperty") is not null, + "SettingsReader no longer has a public TryGetProperty. That name is what this file's extraction " + + "matches on, so renaming it removes every Lite settings key from the extracted set and lets " + + "settings.sample.json drift undetected (#2418, #2428, #2444). If it really must be renamed, " + + "teach the Scanner regex the new name in the SAME change and prove it here."); + } + + /// + /// key -> the source file it was found in, for a failure message that names where to look. + /// + private static Dictionary ReadKeysFromLoaders() + { + var keys = new Dictionary(StringComparer.Ordinal); + + foreach (var name in ReaderSources) + { + var path = Path.Combine(AppContext.BaseDirectory, "Fixtures", name); + Assert.True( + File.Exists(path), + $"{name} was not copied beside the test binary — check the csproj None/Link item."); + + ExtractKeys(File.ReadAllText(path), name, keys); + } + + return keys; + } + + /// + /// The extraction itself, over TEXT rather than over a path, so the self-tests below can run it against a + /// source they control instead of only against whatever the tree happens to contain today. Split out + /// during #2444: the loaders' read shape changed, and a guard whose only exercise is the real file cannot + /// show that it still SEES the new shape — it can only fail later, obliquely, on a count. + /// + private static void ExtractKeys(string source, string sourceName, Dictionary into) + { + var openFile = string.Empty; + foreach (Match match in Scanner.Matches(source)) + { + if (match.Groups["file"].Success) + { + openFile = match.Groups["file"].Value; + } + else if (openFile == "settings.json") + { + into.TryAdd(match.Groups["key"].Value, sourceName); + } + } + } + + /// + /// The comparison makes, so the self-test below + /// exercises the REAL check rather than a re-implementation of it that could drift away from it. + /// + private static List UndocumentedKeys(Dictionary read, HashSet documented) => + read.Keys + .Where(k => !documented.Contains(k) && !LoaderOnlyKeys.Contains(k)) + .OrderBy(k => k, StringComparer.Ordinal) + .ToList(); + + private static HashSet SampleKeys() + { + using var doc = ParseSample(); + return doc.RootElement.EnumerateObject().Select(p => p.Name).ToHashSet(StringComparer.Ordinal); + } + + private static JsonDocument ParseSample() + { + var path = Path.Combine(AppContext.BaseDirectory, "Fixtures", "settings.sample.json"); + Assert.True( + File.Exists(path), + "settings.sample.json was not copied beside the test binary — check the csproj None/Link item."); + + /* The sample is JSONC on purpose: the comments ARE the documentation. Lite's own loader parses + settings.json with default options and would reject them, which is why the sample's header + says to copy keys out of it rather than the file itself. */ + return JsonDocument.Parse( + File.ReadAllText(path), + new JsonDocumentOptions { CommentHandling = JsonCommentHandling.Skip, AllowTrailingCommas = false }); + } +} diff --git a/Lite.Tests/SettingsSaveReportTests.cs b/Lite.Tests/SettingsSaveReportTests.cs new file mode 100644 index 000000000..9d089eac2 --- /dev/null +++ b/Lite.Tests/SettingsSaveReportTests.cs @@ -0,0 +1,258 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.RegularExpressions; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2433. Lite's Settings window said "Settings saved." on a save that wrote nothing. +/// +/// Ten writers backed that one button and each rewrote the whole of settings.json behind its own +/// catch. Seven returned void and could not report a write failure by construction; the three that +/// returned something returned whether the BOXES validated, which is a different question. So the toast +/// was shown whenever nothing objected, and "nothing objected" was never about whether a byte reached +/// disk. +/// +/// The repair was to remove the question rather than answer it: one read, ten mutators on one +/// document, one write. That leaves a single ordering rule, which is what these pin — and the rule that +/// matters is that a failed write outranks a validation objection. A writer that rejected a value has +/// already raised its own dialog naming it; what the user has not been told, and can find out nowhere +/// else, is that none of it was saved. +/// +public sealed class SettingsSaveReportTests +{ + [Fact] + public void AnOrdinarySave_Reports_Saved() + { + Assert.Equal( + SettingsSaveOutcome.Saved, + SettingsSaveReport.Classify(documentWritten: true, mcpChanged: false, + alertsValid: true, mcpValid: true, webhooksValid: true)); + } + + [Fact] + public void AnMcpChange_AsksForARestart() + { + Assert.Equal( + SettingsSaveOutcome.SavedAndMcpNeedsRestart, + SettingsSaveReport.Classify(documentWritten: true, mcpChanged: true, + alertsValid: true, mcpValid: true, webhooksValid: true)); + } + + /// + /// The headline case. Everything on the page validated, so on dev this is exactly the state that + /// produced "Settings saved." over a save that wrote nothing at all. + /// + [Fact] + public void AFailedWrite_IsNeverReportedAsSaved() + { + Assert.Equal( + SettingsSaveOutcome.NothingWritten, + SettingsSaveReport.Classify(documentWritten: false, mcpChanged: false, + alertsValid: true, mcpValid: true, webhooksValid: true)); + } + + /// + /// And it stays the headline case when a writer ALSO objected. The objection has its own dialog; the + /// failed write does not, and it is the one that describes what happened to the other nine writers' + /// work. Ordering the other way round would leave the most important fact of the save unsaid. + /// + [Theory] + [InlineData(false, true, true)] + [InlineData(true, false, true)] + [InlineData(true, true, false)] + public void AFailedWrite_OutranksAValidationObjection(bool alertsValid, bool mcpValid, bool webhooksValid) + { + Assert.Equal( + SettingsSaveOutcome.NothingWritten, + SettingsSaveReport.Classify(documentWritten: false, mcpChanged: true, + alertsValid, mcpValid, webhooksValid)); + } + + /// + /// A rejected value suppresses the toast without claiming nothing was saved — because the rest of the + /// document DID reach disk. This is the arm that would be a lie under any of the three options #2433 + /// listed that kept the writes separate. + /// + [Theory] + [InlineData(false, true, true)] + [InlineData(true, false, true)] + [InlineData(true, true, false)] + public void AnObjection_SuppressesTheToast_WithoutClaimingNothingWasSaved( + bool alertsValid, bool mcpValid, bool webhooksValid) + { + Assert.Equal( + SettingsSaveOutcome.WrittenWithObjections, + SettingsSaveReport.Classify(documentWritten: true, mcpChanged: false, + alertsValid, mcpValid, webhooksValid)); + } +} + +/// +/// The category behind #2433 rather than the instance, and the companion to +/// 's one-writer pin: the Save button must ask whether the +/// write happened before it is allowed to say anything, and none of its writers may go behind its back to +/// settings.json. +/// +/// Source-parsing because the defect was a wiring one. Every individual writer was correct about the +/// thing it knew; what was missing was any path from "the write failed" to the sentence the user reads. +/// +public sealed class SettingsSaveButtonHonestyTests +{ + private static string SettingsWindowSource() => + File.ReadAllText(FindRepoFile(Path.Combine("Lite", "Windows", "SettingsWindow.xaml.cs"))); + + /// + /// The toast is reachable only through the classifier, so a future edit cannot restore the old + /// "nothing objected, therefore say it saved" shortcut without deleting this. + /// + [Fact] + public void TheSaveButton_AsksTheClassifierBeforeItSaysAnything() + { + var body = MethodBody(SettingsWindowSource(), "SaveButton_Click"); + + Assert.Contains("App.WriteSettingsDocument(", body, StringComparison.Ordinal); + Assert.Contains("SettingsSaveReport.Classify(", body, StringComparison.Ordinal); + + /* Every "Settings saved" string in the method must sit under a classifier arm; the only way to + check that cheaply is that the classify call comes first. */ + var classify = body.IndexOf("SettingsSaveReport.Classify(", StringComparison.Ordinal); + var firstToast = body.IndexOf("Settings saved", StringComparison.Ordinal); + Assert.True(classify < firstToast, + "SaveButton_Click can reach \"Settings saved.\" without going through SettingsSaveReport.Classify, " + + "which is #2433: the toast is about whether the boxes validated rather than whether anything " + + "reached disk."); + } + + /// + /// Each Save* the button calls takes the shared document. A writer that opened settings.json for itself + /// would be back to its own read, its own write and its own swallowed failure — one write is what makes + /// one answer possible. + /// + [Fact] + public void EveryWriterTheSaveButtonCalls_TakesTheSharedDocument() + { + var source = SettingsWindowSource(); + var body = MethodBody(source, "SaveButton_Click"); + + var called = Regex.Matches(body, @"\b(Save[A-Za-z]+)\(root\)") + .Select(m => m.Groups[1].Value) + .ToList(); + + Assert.True(called.Count >= 9, + $"SaveButton_Click hands the shared document to only {called.Count} writer(s) — the rest are " + + "opening settings.json for themselves again."); + + foreach (var name in called) + { + Assert.Matches(new Regex($@"\b{name}\(JsonNode root\)"), source); + } + } + + /// + /// The window's other claim about a save, and the one that used to contradict the dialog. The theme + /// selector applies its choice LIVE, and _saved is what stops the close handlers reverting that + /// preview — so setting it on the click rather than on the write left an unpersisted theme applied for + /// the rest of the run while the dialog said nothing had been saved. It has to track the write. + /// + [Fact] + public void TheLiveThemePreview_OnlySurvivesASaveThatWrote() + { + var source = SettingsWindowSource(); + var body = MethodBody(source, "SaveButton_Click"); + + Assert.Contains("_saved = written;", body, StringComparison.Ordinal); + Assert.DoesNotContain("_saved = true;", body, StringComparison.Ordinal); + + /* And the gate it feeds is still there, so this cannot quietly stop meaning anything. */ + Assert.Contains("if (!_saved)", source, StringComparison.Ordinal); + } + + /// + /// The restart is a bigger claim than the toast, so it needs the same gate. MainWindow reads + /// McpSettingsChanged after ShowDialog and, when it is set, stops and restarts the MCP + /// server — dropping every connected client — then reloads the port from settings.json on DISK. Set on + /// the click rather than the write, a failed save produced a disruptive restart back onto the OLD + /// configuration, moments after the app had said nothing was saved. + /// + [Fact] + public void TheMcpRestart_OnlyFollowsASaveThatWrote() + { + var body = MethodBody(SettingsWindowSource(), "SaveButton_Click"); + + Assert.Contains("if (mcpChanged && written) McpSettingsChanged = true;", body, StringComparison.Ordinal); + } + + /// + /// Every mutator sits inside the guarded region, not just the read and the write. + /// + /// Before the consolidation each writer carried its own try, so an exception thrown while BUILDING + /// a value — not only on the disk I/O — was caught, logged, and the remaining writers still ran. + /// Guarding only the ends would leave the mutators between them able to escape into the generic + /// "An error occurred" dispatcher dialog, which is precisely the class of silence this PR is about. + /// Nothing is written in that case either, and that is what the user should be told. + /// + [Fact] + public void EveryMutator_SitsInsideTheGuardedRegion() + { + var body = MethodBody(SettingsWindowSource(), "SaveButton_Click"); + + var openTry = body.IndexOf("try", StringComparison.Ordinal); + var firstMutator = body.IndexOf("SaveMcpSettingsAsync(root)", StringComparison.Ordinal); + var lastMutator = body.IndexOf("SaveWebhookSettings(root)", StringComparison.Ordinal); + var theCatch = body.IndexOf("catch (Exception", StringComparison.Ordinal); + + Assert.True(openTry >= 0 && firstMutator >= 0 && lastMutator >= 0 && theCatch >= 0, + "SaveButton_Click no longer has the shape this guard describes — its anchors moved and it is " + + "testing nothing."); + Assert.True(openTry < firstMutator && lastMutator < theCatch, + "A mutator sits outside SaveButton_Click's try, so an exception building a settings value " + + "escapes into App's generic unhandled-exception dialog instead of the honest \"Nothing was " + + "saved\" this method owes the user."); + } + + private static string MethodBody(string source, string methodName) + { + var signature = source.IndexOf(methodName + "(", StringComparison.Ordinal); + Assert.True(signature >= 0, $"No method named {methodName} — this guard's anchor moved and it is testing nothing."); + + var open = source.IndexOf('{', signature); + Assert.True(open >= 0, $"{methodName} has no body."); + + var depth = 0; + for (var i = open; i < source.Length; i++) + { + if (source[i] == '{') depth++; + else if (source[i] == '}' && --depth == 0) return source[open..i]; + } + + throw new InvalidOperationException($"{methodName}'s body is unbalanced."); + } + + private static string FindRepoFile(string relativePath) + { + var dir = AppContext.BaseDirectory; + for (var i = 0; i < 8 && dir is not null; i++) + { + var candidate = Path.Combine(dir, relativePath); + if (File.Exists(candidate)) + { + return candidate; + } + dir = Path.GetDirectoryName(dir); + } + + throw new FileNotFoundException($"Could not locate {relativePath} walking up from {AppContext.BaseDirectory}"); + } +} diff --git a/Lite.Tests/SettingsValueReadTests.cs b/Lite.Tests/SettingsValueReadTests.cs new file mode 100644 index 000000000..c3dc4dac9 --- /dev/null +++ b/Lite.Tests/SettingsValueReadTests.cs @@ -0,0 +1,307 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.Json; +using PerformanceMonitor.Common; +using PerformanceMonitorLite; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2444 — one badly-shaped value used to cost every setting after it. +/// +/// #2425 split the DOCUMENT fault out of App.LoadAlertSettings's single try, so a +/// trailing comma is now reported rather than silently reverting everything. What it left behind was the +/// VALUE fault: a key holding a string where an int belongs threw on its own Get* call and abandoned +/// every read after it, so which settings survived depended on where the bad key happened to sit in the +/// file. One wrong value near the top cost almost everything; the same value near the bottom cost almost +/// nothing. That ordering is an implementation detail of the loader's line order, and it was invisible. +/// +/// The position tests below are the ones that matter, and they are position tests on purpose: a +/// fixture with a single bad key and nothing after it passes against the OLD code too. Only a bad key with +/// good keys on BOTH sides can tell "this key fell back" from "this key and the rest of the file fell +/// back". +/// +/// Shares the app-alert-statics collection with the other classes that drive +/// App.LoadAlertSettings: it rewrites the whole alert block, and xUnit runs classes in parallel. +/// +[Collection("app-alert-statics")] +public class SettingsValueReadTests +{ + private static void FailOnWrite(string key, string value) => + Assert.Fail($"LoadAlertSettings wrote credential '{key}' when nothing should have been saved."); + + private static string WriteSettings(string tag, string json) + { + var dir = Path.Combine(Path.GetTempPath(), $"pmlite_{tag}_{Guid.NewGuid():N}"); + Directory.CreateDirectory(dir); + File.WriteAllText(Path.Combine(dir, "settings.json"), json); + return dir; + } + + private static SettingsReader ReaderOver(string json) => + new(JsonDocument.Parse(json).RootElement); + + /* ---- the defect, end to end through the real loader ---- */ + + /// + /// The headline. A quoted number sits between two perfectly good settings; the good ones both apply and + /// only the bad key falls back. Against dev the LAST key is never reached at all, which is the whole of + /// #2444 in one assertion. + /// + [Fact] + public void ABadValue_CostsItsOwnKeyAndNothingAfterIt() + { + var dir = WriteSettings("badvalue", + """ + { + "alerts_enabled": true, + "alert_cpu_threshold": "ninety", + "analysis_timeout_seconds": 123 + } + """); + + try + { + App.AlertsEnabled = false; + App.AlertCpuThreshold = 42; + App.AnalysisTimeoutSeconds = 99; + + App.LoadAlertSettings(dir, key => "stored:" + key, FailOnWrite); + + /* Before the bad key: always worked, still works. */ + Assert.True(App.AlertsEnabled); + + /* The bad key itself: keeps the value it had, which is the only correct answer. */ + Assert.Equal(42, App.AlertCpuThreshold); + + /* After the bad key: THIS is the regression. It was unreachable on dev. */ + Assert.Equal(123, App.AnalysisTimeoutSeconds); + } + finally + { + Directory.Delete(dir, true); + } + } + + /// + /// The same shape one level down, and a second throw site the old single try hid: an array whose + /// ELEMENT is the wrong kind. elem.GetString() throws on the number, and on dev that one element + /// cost every setting in the rest of the file. + /// + [Fact] + public void AWrongShapedArrayElement_CostsNeitherTheListNorTheRestOfTheFile() + { + var dir = WriteSettings("badelement", + """ + { + "alert_excluded_databases": ["keep_me", 5, "keep_me_too"], + "analysis_timeout_seconds": 321 + } + """); + + try + { + App.AnalysisTimeoutSeconds = 99; + + App.LoadAlertSettings(dir, key => "stored:" + key, FailOnWrite); + + Assert.Equal(new[] { "keep_me", "keep_me_too" }, App.AlertExcludedDatabases); + Assert.Equal(321, App.AnalysisTimeoutSeconds); + } + finally + { + Directory.Delete(dir, true); + } + } + + /* ---- the reader's own policy ---- */ + + /// + /// Reports the whole set, not the first. A user who hand-edited one line has one mistake; a user who + /// pasted a block out of settings.sample.json has several, and stopping at the first means they fix, + /// restart, and discover the next one — several times over. + /// + [Fact] + public void EveryBadKeyIsReported_NotOnlyTheFirst() + { + var read = ReaderOver( + """ + { "a": "no", "b": 1, "c": true, "d": [1], "e": null } + """); + + /* Every one of the five is a different wrong shape, and every one hands back the caller's value. */ + Assert.True(read.TryGetProperty("a", out var a)); + Assert.True(a.Bool(fallback: true)); + + Assert.True(read.TryGetProperty("b", out var b)); + Assert.Equal("kept", b.Text("kept")); + + Assert.True(read.TryGetProperty("c", out var c)); + Assert.Equal(5, c.Int(5)); + + Assert.True(read.TryGetProperty("d", out var d)); + Assert.Equal(2.5, d.Double(2.5, 0, 10)); + + Assert.True(read.TryGetProperty("e", out var e)); + Assert.Equal("kept", e.Text("kept")); + + Assert.Equal(new[] { "a", "b", "c", "d", "e" }, read.Problems.Select(p => p.Key).ToArray()); + + /* The message names the kind that WAS there, which is what lets someone find the line. */ + Assert.Contains("string", read.Problems[0].Problem, StringComparison.Ordinal); + Assert.Contains("true or false", read.Problems[0].Problem, StringComparison.Ordinal); + Assert.Contains("null", read.Problems[4].Problem, StringComparison.Ordinal); + } + + /// An absent key is how every default works and how an older settings.json behaves. It is never + /// a problem, and reporting it would make the dialog fire on nearly every install. + [Fact] + public void AnAbsentKeyIsNotAProblem() + { + var read = ReaderOver("""{ "present": 1 }"""); + + Assert.False(read.TryGetProperty("absent", out _)); + Assert.True(read.TryGetProperty("present", out var present)); + Assert.Equal(1, present.Int(0)); + Assert.Empty(read.Problems); + } + + /// + /// The latent bug the clamps' move made unreachable. The old inline form was + /// (int)Math.Max(0, v.GetInt64()), which floors and then NARROWS — so a hand-typed value beyond + /// int range wrapped, and a "bigger threshold" became a negative one. The reader clamps before it + /// narrows, so the worst a huge number can do is land on the bound. This is not a shape fault, so it is + /// deliberately NOT reported: the value was readable, it was just out of range. + /// + [Fact] + public void AnOutOfRangeNumberClampsToTheBound_RatherThanWrappingNegative() + { + var read = ReaderOver("""{ "huge": 5000000000, "negative": -7 }"""); + + Assert.True(read.TryGetProperty("huge", out var huge)); + Assert.Equal(int.MaxValue, huge.Int(1, 0, int.MaxValue)); + + Assert.True(read.TryGetProperty("negative", out var negative)); + Assert.Equal(0, negative.Int(1, 0, 100)); + + Assert.Empty(read.Problems); + } + + /// + /// Review found the boundary the clamp promise stopped at (#2453): the clamped readers go through + /// TryGetInt64, so a number beyond Int64 fell out of the clamp and was reported as though it were + /// not a whole number — a nonsense sentence about a number, and in a dialog. It clamps by SIGN instead, + /// which makes the promise absolute rather than "up to Int64". + /// + [Fact] + public void ANumberBeyondInt64_StillClampsToABound() + { + var read = ReaderOver("""{ "vast": 99999999999999999999, "vast_negative": -99999999999999999999 }"""); + + Assert.True(read.TryGetProperty("vast", out var vast)); + Assert.Equal(100, vast.Int(1, 0, 100)); + + Assert.True(read.TryGetProperty("vast_negative", out var negative)); + Assert.Equal(0, negative.Int(1, 0, 100)); + + Assert.Empty(read.Problems); + } + + /// + /// The regression the FIRST cut of that fix introduced, and the sharper of the two review findings. + /// TryGetInt64 returns false for a number that is not a pure integer token as well as for one that + /// is too big, so choosing the bound by SIGN turned analysis_timeout_seconds: 30.0 into 600 — a + /// value the user never asked for, never saw flagged, and could not have found. Reading the remainder as + /// a double is what tells the two apart, and 30.0 is plainly 30. + /// + [Fact] + public void AFractionalNumber_ReadsAsItsValue_RatherThanLandingOnTheMaximum() + { + var read = ReaderOver("""{ "whole": 30.0, "fraction": 5.5, "under": -0.5 }"""); + + Assert.True(read.TryGetProperty("whole", out var whole)); + Assert.Equal(30, whole.Int(99, 30, 600)); + + /* Truncated INTO the range, which is the same silent adjustment the clamp already is and is confined + to the readers whose caller declared a range for exactly that. It is not 100. */ + Assert.True(read.TryGetProperty("fraction", out var fraction)); + Assert.Equal(5, fraction.Int(99, 0, 100)); + + Assert.True(read.TryGetProperty("under", out var under)); + Assert.Equal(0, under.Int(99, 0, 100)); + + Assert.Empty(read.Problems); + } + + /// + /// The other side of the boundary. An UNCLAMPED read has no range to put an unusable number into, so it + /// must not invent one — and it names which of the two problems the value has, because "holds a JSON + /// number where a whole number belongs" is a nonsense sentence about a number. A number that IS exact is + /// still taken however it was written. + /// + [Fact] + public void AnUnclampedRead_TakesAnExactNumber_AndNamesWhyItRejectsTheRest() + { + var exact = ReaderOver("""{ "n": 30.0 }"""); + Assert.True(exact.TryGetProperty("n", out var n)); + Assert.Equal(30, n.Int(7)); + Assert.Empty(exact.Problems); + + var fractional = ReaderOver("""{ "n": 90.7 }"""); + Assert.True(fractional.TryGetProperty("n", out var f)); + Assert.Equal(7, f.Int(7)); + Assert.Contains("not a whole number", Assert.Single(fractional.Problems).Problem, StringComparison.Ordinal); + + var vast = ReaderOver("""{ "n": 99999999999999999999 }"""); + Assert.True(vast.TryGetProperty("n", out var v)); + Assert.Equal(7, v.Int(7)); + + var problem = Assert.Single(vast.Problems); + Assert.Contains("out of range", problem.Problem, StringComparison.Ordinal); + Assert.DoesNotContain("where a whole number belongs", problem.Problem, StringComparison.Ordinal); + } + + /// + /// The line between "wrong shape" and "wrong word", which decides what the startup dialog is allowed to + /// complain about. A number where a theme name belongs is this file's business; a string that is simply + /// not one of the three theme names is the caller's own vocabulary, has always been ignored quietly, and + /// widening THAT into a dialog is a separate decision from this one. + /// + [Fact] + public void OnlyTheShapeIsReported_NotAnUnrecognisedButWellShapedString() + { + var wrongShape = ReaderOver("""{ "color_theme": 5 }"""); + Assert.True(wrongShape.TryGetProperty("color_theme", out var number)); + Assert.Null(number.TextOrNull()); + Assert.Equal("color_theme", Assert.Single(wrongShape.Problems).Key); + + var wrongWord = ReaderOver("""{ "color_theme": "Chartreuse" }"""); + Assert.True(wrongWord.TryGetProperty("color_theme", out var text)); + Assert.Equal("Chartreuse", text.TextOrNull()); + Assert.Empty(wrongWord.Problems); + } + + /// A value read twice cannot be reported twice — the dialog lists keys, and a key listed + /// twice reads as two different problems with the same name. + [Fact] + public void OneKeyIsReportedOnce() + { + var read = ReaderOver("""{ "n": "not a number" }"""); + + Assert.True(read.TryGetProperty("n", out var n)); + Assert.Equal(1, n.Int(1)); + Assert.Equal(2, n.Int(2, 0, 10)); + + Assert.Equal("n", Assert.Single(read.Problems).Key); + } +} diff --git a/Lite.Tests/SharedDuckDbFixture.cs b/Lite.Tests/SharedDuckDbFixture.cs index 084665002..5ef45906c 100644 --- a/Lite.Tests/SharedDuckDbFixture.cs +++ b/Lite.Tests/SharedDuckDbFixture.cs @@ -20,8 +20,40 @@ namespace PerformanceMonitorLite.Tests; /// /// Deliberately IClassFixture (one database per class), NOT a collection fixture: a /// single shared collection would serialize these classes against each other and give -/// back most of the win. Each class owns its own database file, keeping xUnit's -/// cross-class parallelism intact. +/// back most of the win. Each class owns its own database file, so the SCHEMA-BUILD cost +/// — the ~80 DDL statements above, which is what this fixture exists to amortise — runs +/// concurrently across classes. +/// +/// What that does NOT buy (#2376). The file separation is real, but it does +/// not make the classes independent at RUNTIME: DuckDbInitializer's lock is +/// private static readonly, so it is ONE lock for the whole process no matter which +/// database file a call targets. +/// +/// Be precise about what that costs, because the obvious reading overstates it. The +/// lock is a ReaderWriterLockSlim, so readers do NOT queue behind each other — any +/// number of AcquireReadLock holders run concurrently across classes, and this suite +/// is overwhelmingly readers (~184 read call sites against ~20 write sites). The +/// serialization is writer-driven: a writer excludes everyone, and readers block only while +/// a writer holds the lock or is waiting for it. The contention is therefore bursty around +/// the write sites — archival, compaction, CHECKPOINT, the mute and alert-history stores — +/// rather than a flat tax on every database call. +/// +/// That distinction decides what a fix could even look like: a collection fixture +/// would serialize the READERS as well, so it does not address writer-driven contention and +/// would cost more than it saves. Nobody has measured how much of the suite's ~190-220s is +/// this, and the honest answer is that it may be very little. +/// +/// That lock is correct and must not be narrowed to fix this: production creates +/// several DuckDbInitializer instances over the same App.DatabasePath +/// (MainWindow, DatabaseStateOverridesWindow, and DuckDbAlertHistoryStore), and the static +/// lock is exactly what keeps those mutually exclusive. Per-instance would trade a slow +/// suite for a real data race. +/// +/// The practical consequence is worth knowing when a test here fails oddly: a +/// scheduling-pressure window can starve the 5-second write-lock acquisition in +/// LocalDataService.GetDatabaseStateDeviationsAsync, whose maintenance block is +/// best-effort and simply SKIPS on timeout. That is #2374 — a test that assumed the +/// maintenance had run, rather than waiting for it, failed a nightly. /// public sealed class SharedDuckDbFixture : IAsyncLifetime { diff --git a/Lite.Tests/SignificantWaitsToolTests.cs b/Lite.Tests/SignificantWaitsToolTests.cs new file mode 100644 index 000000000..b2f62bb66 --- /dev/null +++ b/Lite.Tests/SignificantWaitsToolTests.cs @@ -0,0 +1,192 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitor.Common; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Lite's get_health_parser_significant_waits (#2484), the twin of Darling's. The tool landed on both SKUs +/// in one change rather than on the divergence ratchet, so the promises the SKUs make to each other are +/// pinned here: the same three kinds of empty, in words that cannot be mistaken for one another. +/// +/// Written at the TOOL level, not the reader level. The reader is the easy half; the three-way empty +/// branch and the cap contract both live in the tool, and a reader-level test would see neither. +/// +public sealed class SignificantWaitsToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "SigWaitsSrv"; + + /* + Lite does not store a server id -- ServerResolver DERIVES it from the storage name, so seeded rows + have to be written under the same derived value the tool resolves to. A hardcoded id would seed + rows the tool looks straight past, and the never-captured assertion would pass for the wrong reason. + */ + private readonly int _serverId; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private DuckDBConnection? _seedConn; + private long _nextId = 1; + + public SignificantWaitsToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-sigwaits-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + private static string LoadFixture(string name) => + File.ReadAllText(Path.Combine(AppContext.BaseDirectory, "Fixtures", "SystemHealth", name)); + + [Fact] + public async Task ThreeKindsOfNothing_AreThreeDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + /* Registered, nothing ever captured: NOT an all-clear, and it must refuse to read as one. */ + var never = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, ServerName, 24, 50); + var neverRoot = JsonDocument.Parse(never).RootElement; + Assert.Equal("unavailable", neverRoot.GetProperty("status").GetString()); + var neverText = neverRoot.GetProperty("message").GetString()!; + Assert.Contains("EVER", neverText, StringComparison.Ordinal); + Assert.Contains("NOT an all-clear", neverText, StringComparison.Ordinal); + Assert.DoesNotContain("widen", neverText, StringComparison.OrdinalIgnoreCase); + + /* Captured, but outside the asked-for window: a quiet window, and widening IS the move. */ + await SeedWaitAsync(LoadFixture("wait_info.xml"), Truncate(DateTime.UtcNow.AddHours(-48))); + + var quiet = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, ServerName, 1, 50); + var quietRoot = JsonDocument.Parse(quiet).RootElement; + Assert.Equal("empty", quietRoot.GetProperty("status").GetString()); + var quietText = quietRoot.GetProperty("message").GetString()!; + Assert.Contains("widen", quietText, StringComparison.OrdinalIgnoreCase); + Assert.DoesNotContain("EVER", quietText, StringComparison.Ordinal); + + /* + Captured IN the window and gated out: the healthy answer, and the one the other two must never + be confused with. Same fixture, duration dropped under the 500 ms bar. + */ + var tooShort = LoadFixture("wait_info.xml").Replace("1500", "100", StringComparison.Ordinal); + Assert.DoesNotContain("1500", tooShort, StringComparison.Ordinal); + await SeedWaitAsync(tooShort, Truncate(DateTime.UtcNow.AddMinutes(-10))); + + var gated = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, ServerName, 4, 50); + var gatedRoot = JsonDocument.Parse(gated).RootElement; + Assert.Equal("empty", gatedRoot.GetProperty("status").GetString()); + var gatedText = gatedRoot.GetProperty("message").GetString()!; + Assert.Contains("none was significant", gatedText, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", gatedText, StringComparison.Ordinal); + Assert.DoesNotContain("widen", gatedText, StringComparison.OrdinalIgnoreCase); + } + + [Fact] + public async Task ThePayloadCarriesTheStatementThatPaidForTheWait() + { + var service = new LocalDataService(_duckDb); + + /* One significant wait and one gated-out event, both inside the window. */ + await SeedWaitAsync(LoadFixture("wait_info.xml"), Truncate(DateTime.UtcNow.AddMinutes(-9))); + await SeedWaitAsync( + LoadFixture("wait_info.xml").Replace("1500", "100", StringComparison.Ordinal), + Truncate(DateTime.UtcNow.AddMinutes(-10))); + + var hit = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, ServerName, 4, 50); + var root = JsonDocument.Parse(hit).RootElement; + Assert.Equal(ServerName, root.GetProperty("server").GetString()); + Assert.Equal(1, root.GetProperty("wait_count").GetInt32()); + + var wait = root.GetProperty("waits")[0]; + Assert.Equal("PAGEIOLATCH_SH", wait.GetProperty("wait_type").GetString()); + Assert.Equal(1500, wait.GetProperty("duration_ms").GetInt64()); + Assert.Equal(12, wait.GetProperty("signal_duration_ms").GetInt64()); + Assert.Equal(57, wait.GetProperty("session_id").GetInt32()); + + /* + The SQL text is the half get_wait_stats can never give: the instance-wide totals name a wait + type and never the statement that paid for it. Both SKUs advertise the same field name. + */ + Assert.Contains("dbo.big_table", wait.GetProperty("query_text").GetString()!, StringComparison.Ordinal); + + /* The event time comes back as the raw naive-UTC XE @timestamp, not the grid's local render. */ + Assert.StartsWith("2026-07-05T12:04:30", wait.GetProperty("event_time").GetString()!, StringComparison.Ordinal); + } + + [Fact] + public async Task AnOutOfRangeCap_IsRefused_NotSilentlyClamped() + { + var service = new LocalDataService(_duckDb); + await SeedWaitAsync(LoadFixture("wait_info.xml"), Truncate(DateTime.UtcNow.AddMinutes(-9))); + + var tooBig = await McpHealthParserTools.GetSignificantWaits(service, _serverManager, ServerName, 24, 5000); + Assert.Contains("exceeds maximum of", tooBig, StringComparison.Ordinal); + Assert.Contains("1000", tooBig, StringComparison.Ordinal); + } + + private static DateTime Truncate(DateTime value) => + DateTime.SpecifyKind(new DateTime(value.Ticks - (value.Ticks % TimeSpan.TicksPerSecond)), DateTimeKind.Unspecified); + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedWaitAsync(string eventXml, DateTime eventTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO system_health_events + (system_health_event_id, collection_time, server_id, server_name, event_time, event_type, event_xml) +VALUES ($1, $2, $3, $4, $5, $6, $7)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.UtcNow }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = eventTimeUtc }); + cmd.Parameters.Add(new DuckDBParameter { Value = SystemHealthParser.WaitInfoEvent }); + cmd.Parameters.Add(new DuckDBParameter { Value = eventXml }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/SqlServerPermissionErrorsTests.cs b/Lite.Tests/SqlServerPermissionErrorsTests.cs new file mode 100644 index 000000000..baabf6e66 --- /dev/null +++ b/Lite.Tests/SqlServerPermissionErrorsTests.cs @@ -0,0 +1,116 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using PerformanceMonitor.Collectors; +using PerformanceMonitor.Common; +using Xunit; + +namespace Lite.Tests; + +/// +/// #2512: the permission-denial number set, and the Azure sentence that rides on 262. +/// +/// Why this set is a shared predicate rather than three literals, and why that is what gets +/// tested. The numbers lived in three hand-maintained places — SqlServerTargetProvider.Classify, +/// Darling's worker catch filter, and Lite's RunCollectorAsync catch — and they had already +/// drifted apart (916 was in the first and neither of the other two). None of those three call sites can be +/// unit-tested directly, because SqlException cannot be constructed: the only way to pin them is to +/// give them one set to read and pin the set. That is what these do. +/// +/// What 262 costs when it is missing. "%ls permission denied in database '%.*ls'" is what a +/// database-scoped DMV read raises when the login does not hold the permission IN the named database. It is +/// unmistakably a permission denial, and it was classified Unclassified — which means ERROR, on +/// every collection cycle, forever. That is what the #2150 field report looked like (11x consecutive on an +/// Azure SQL Database elastic pool), and it is why was gated off the +/// entire Azure SQL Database tier rather than allowed to degrade. The gate is gone; this is what makes that +/// safe. +/// +public sealed class SqlServerPermissionErrorsTests +{ + /// + /// 262 is the number this issue adds; the other five are the pre-existing set, asserted so the + /// extraction cannot have quietly dropped one on its way into a shared method. + /// + [Theory] + [InlineData(229)] /* EXECUTE/SELECT permission denied on an object */ + [InlineData(262)] /* permission denied IN a database — the #2150/#2512 tempdb case */ + [InlineData(297)] /* the user does not have permission to perform this action */ + [InlineData(300)] /* VIEW SERVER STATE denied (a service-objective limit on Azure SQL DB) */ + [InlineData(916)] /* the principal cannot access the database under the current security context */ + [InlineData(8189)] /* sys.traces' own denial, ALTER TRACE missing (#1823) */ + public void PermissionDenials_AreClassifiedAsPermissions(int number) + => Assert.True(SqlServerPermissionErrors.IsPermissionDenied(number)); + + /// + /// The other side of the set. These must NOT degrade to PERMISSIONS: a missing object wants an install + /// or an upgrade rather than a grant, a lock-timeout yield is evidence about the monitored server, and a + /// timeout or an unrecognized number has to stay loud. A predicate that swallowed them would turn the + /// non-fatal bucket into a place failures go to be ignored. + /// + [Theory] + [InlineData(208)] /* invalid object name -> ObjectMissing, not Permissions */ + [InlineData(1222)] /* lock request timeout -> LockTimeoutYield for collectors that declare it */ + [InlineData(-2)] /* command timeout */ + [InlineData(207)] /* invalid column name (version drift) */ + [InlineData(40615)] /* Azure firewall rejection */ + [InlineData(0)] + public void NonPermissionFailures_StayLoud(int number) + => Assert.False(SqlServerPermissionErrors.IsPermissionDenied(number)); + + /// + /// 262 in TEMPDB on Azure SQL Database gets the same treatment 300 already got (#1631): the raw error + /// names tempdb and reads as a missing GRANT, and there is no grant to issue there. Empty off Azure, + /// where a 262 IS a missing grant and the fix is to issue it — the same asymmetry the 300 hint draws. + /// + [Fact] + public void AzureDmvPermissionHint_ExplainsTempDb262_OnAzureOnly() + { + var azure262 = AzureDmvPermissionHint.For(262, isAzureSqlDb: true, TempDbDenial); + + Assert.Contains("TEMPDB", azure262, System.StringComparison.Ordinal); + Assert.Contains("##MS_ServerStateReader##", azure262, System.StringComparison.Ordinal); + Assert.Contains("no grant to issue", azure262, System.StringComparison.Ordinal); + + Assert.Empty(AzureDmvPermissionHint.For(262, isAzureSqlDb: false, TempDbDenial)); + + /* The 300 arm is untouched by the switch that replaced its if-guard, and 229 still says nothing. */ + Assert.Contains("SERVICE OBJECTIVE", AzureDmvPermissionHint.For(300, isAzureSqlDb: true), System.StringComparison.Ordinal); + Assert.Empty(AzureDmvPermissionHint.For(229, isAzureSqlDb: true)); + } + + /// + /// The review catch on #2512, and the reason 262 reads the message where 300 does not. 300 is + /// server-scoped, so its number settles it. 262 names a DATABASE, and the advice inverts on which one: + /// in tempdb there is no grant to issue, in a user database there is and the raw error already names + /// it. Keying purely off the number appended tempdb guidance — "reach tempdb's space DMVs + /// through server-level state access" — to a denial in someone's user database, which is + /// worse than appending nothing, because it sends them after a role membership that would not have + /// helped. No collector raises that today, but the per-database loop collectors run against arbitrary + /// user databases on Azure SQL DB and are one permission change away from it. + /// A null message stays silent rather than guessing, so a call site that forgets to pass it + /// loses a helpful sentence instead of gaining a wrong one. + /// + [Theory] + [InlineData("VIEW DATABASE PERFORMANCE STATE permission denied in database 'AdventureWorks'.")] + [InlineData("VIEW DATABASE STATE permission denied in database 'reporting_tempdb_stage'.")] + [InlineData("")] + [InlineData(null)] + public void AzureDmvPermissionHint_SaysNothingAbout262_WhenTheDenialIsNotTempDb(string? message) + => Assert.Empty(AzureDmvPermissionHint.For(262, isAzureSqlDb: true, message)); + + /// The name is matched QUOTED, and the collation of the server decides its casing. + [Theory] + [InlineData("permission denied in database 'tempdb'.")] + [InlineData("permission denied in database 'TempDB'.")] + [InlineData("permission denied in database 'TEMPDB'.")] + public void AzureDmvPermissionHint_MatchesTempDb_WhateverTheCasing(string message) + => Assert.NotEmpty(AzureDmvPermissionHint.For(262, isAzureSqlDb: true, message)); + + private const string TempDbDenial = + "VIEW DATABASE PERFORMANCE STATE permission denied in database 'tempdb'."; +} diff --git a/Lite.Tests/StatusBarSizeReadLockTests.cs b/Lite.Tests/StatusBarSizeReadLockTests.cs new file mode 100644 index 000000000..eb03d418a --- /dev/null +++ b/Lite.Tests/StatusBarSizeReadLockTests.cs @@ -0,0 +1,88 @@ +using System; +using System.Diagnostics; +using System.IO; +using System.Threading; +using System.Threading.Tasks; +using PerformanceMonitorLite.Database; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// Pins the status bar's used-size read against the two ways it can be wrong (#2594). +/// +/// It must not open a connection without the lock. This was the one PERIODIC connection site +/// in that did — driven by the dashboard's 30-second +/// DispatcherTimer, so a live handle sat on the database file at arbitrary moments, including +/// moments the archival path may be deleting and recreating it. +/// +/// And it must not block the dashboard to get it. The obvious fix — take the read lock like +/// every other read — runs on the dispatcher thread, so a size figure nobody is reading would freeze the +/// window behind a long archival. So the contract is specifically a BOUNDED attempt: acquire if free, +/// give up quickly otherwise, and let the caller render the file size alone. Asserting only "it takes a +/// lock" would pass a fix that hangs the UI, which is why this test measures the time. +/// +public class StatusBarSizeReadLockTests +{ + /// + /// Generous against the 100 ms budget rather than tight: this asserts the read GAVE UP rather than + /// waited, and a CI machine under load can overshoot a 100 ms timeout considerably without the + /// behaviour being wrong. The distinction being drawn is against a wait for the full hold below, + /// which is an order of magnitude larger. + /// + private static readonly TimeSpan GaveUpCeiling = TimeSpan.FromSeconds(2); + + private static readonly TimeSpan WriteLockHold = TimeSpan.FromSeconds(10); + + [Fact] + public async Task GetUsedDataSizeMb_WhenTheWriteLockIsHeld_GivesUpInsteadOfBlocking() + { + var tempDir = Path.Combine(Path.GetTempPath(), "pmlite-statusbar-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(tempDir); + + try + { + var initializer = new DuckDbInitializer(Path.Combine(tempDir, "test.duckdb")); + + var lockHeld = new ManualResetEventSlim(false); + var release = new ManualResetEventSlim(false); + + /* The write lock is thread-affine, so it has to be taken and released on its own thread. */ + var holder = Task.Run(() => + { + using var writeLock = initializer.AcquireWriteLock(); + lockHeld.Set(); + release.Wait(WriteLockHold); + }); + + Assert.True(lockHeld.Wait(TimeSpan.FromSeconds(5)), "the write lock was never acquired"); + + var stopwatch = Stopwatch.StartNew(); + var used = initializer.GetUsedDataSizeMb(); + stopwatch.Stop(); + + release.Set(); + await holder.WaitAsync(TimeSpan.FromSeconds(15)); + + /* Null, not a number: the caller renders the file size alone in this state, which is the + degraded answer this is choosing on purpose. */ + Assert.Null(used); + + Assert.True( + stopwatch.Elapsed < GaveUpCeiling, + $"the status-bar size read waited {stopwatch.Elapsed.TotalMilliseconds:F0} ms for the write " + + "lock. It runs on the dispatcher thread and must give up rather than block the window."); + } + finally + { + try + { + Directory.Delete(tempDir, recursive: true); + } + catch (IOException) + { + /* A leftover temp directory is not worth failing a passing test over. */ + } + } + } +} diff --git a/Lite.Tests/SweepPressureClassifierTests.cs b/Lite.Tests/SweepPressureClassifierTests.cs index 379b94f18..736f5f792 100644 --- a/Lite.Tests/SweepPressureClassifierTests.cs +++ b/Lite.Tests/SweepPressureClassifierTests.cs @@ -18,14 +18,89 @@ namespace Lite.Tests; /// SKUs' get_collection_health serve so half-rate collection stops being visible only as a service-log /// warning. This SAME table is pinned identically in Darling.Tests so the two SKUs cannot drift. /// -/// The load-bearing case is the motivating measurement: prod-pos-use2-multi-01's four heavy +/// The load-bearing case is the motivating measurement: prod-sql-use2-multi-01's four heavy /// collectors averaged 22,141 + 16,590 + 13,544 + 8,437 ms against a 60s cadence — the body could not /// fit, every relaunch was skipped (~50 warnings/hour), the server collected at half rate, and all 40 /// collectors read HEALTHY, because from each one's own seat nothing was wrong. +/// +/// #2446 added the second dimension and the second load-bearing case, which is the OPPOSITE shape: +/// prod-sql-use2-multi-49 logged six skipped relaunches in three hours while reading OK at 20.4%, because +/// its 37-second collector runs once a day and amortizes to 26 ms/min. The pins below assert both — that +/// the new dimension catches it, and that the VERDICT is unmoved by it, which is the whole reason the two +/// are separate fields. +/// +/// #2460 fixed the INPUT to that second dimension. #2446 built the aligned cycle out of each +/// collector's mean, which is a contradiction on a collector whose runs come in two sizes — and the same +/// server has one. query_store averaged 13,834 ms over 1,155 runs of which 958 carried the +/// empty-enumeration note and cost about 36 ms each, so its 197 PRODUCTIVE runs cost ~80,933 ms apiece: +/// each one, on its own, larger than the whole 60,000 ms budget. The collectors now carry a p95 beside +/// the mean and the cycle is charged . The pins below add +/// the three properties that has to have: it finds the ~67,000 ms #2446 could not see, it moves the +/// VERDICT not at all, and it can never compute a SMALLER cycle than #2446 did. /// public sealed class SweepPressureClassifierTests { - private static (string, double, int) C(string name, double avgMs, int freqMin) => (name, avgMs, freqMin); + /// + /// A UNIMODAL collector: every run costs about the mean, so its p95 IS its mean. Every #2446 fixture + /// below is built from this overload deliberately — with p95 == avg the classifier must reproduce + /// #2446's published numbers to the digit, which makes those pins a regression baseline for #2460 + /// rather than just history. + /// + private static (string, double, double, int) C(string name, double avgMs, int freqMin) => (name, avgMs, avgMs, freqMin); + + /// A collector whose runs come in two sizes: the mean it reports, and the p95 a heavy run costs. + private static (string, double, double, int) C(string name, double avgMs, double p95Ms, int freqMin) => (name, avgMs, p95Ms, freqMin); + + /* --- The two measured servers, as fixtures. Both come from get_collection_health on the dogfood + fleet; the tier roll-ups stand in for the ~35 cheap collectors whose individual names are not the + point, and are the real per-tier sums, so each fixture reconciles to the busy_ms_per_minute the + tool actually reported for that server. --- */ + + /// + /// prod-sql-use2-multi-49 (#2446): 12,248 ms/min sustained — comfortably OK — and a 73,408 ms body on + /// the cycle where every cadence coincides. index_object_stats is 37,207 ms of that in one run. + /// + /// Every collector here is UNIMODAL by default — p95 == mean — which is the conservative + /// assumption and the one #2446 made implicitly for all of them. + /// is the one that is known NOT to hold: pass 80,933 for the measured bimodal profile. That figure is + /// not invented. It is what the store's own numbers force — 958 runs carrying the empty-enumeration + /// note at the 36 ms prod-sql-use2-alpha-01 pays for the identical note, 197 productive runs, and a + /// reported mean of 13,834 ms over all 1,155 — and it reconciles: 958 x 36 + 197 x 80,933 over 1,155 + /// runs gives 13,834.02 ms, the mean the tool reported. + /// + private static IReadOnlyList<(string, double, double, int)> Multi49( + double indexObjectStatsMs = 37_207, + double queryStoreP95Ms = 13_834) => new[] + { + C("procedure_stats", 3_205, 1), + C("query_stats", 2_674, 1), + C("other_one_minute_collectors", 948, 1), + C("query_store", 13_834, queryStoreP95Ms, 5), + C("plan_correction", 11_246, 5), + C("other_five_minute_collectors", 1_680, 5), + C("hourly_collectors", 2_614, 60), + C("index_object_stats", indexObjectStatsMs, 1440), + }; + + /// + /// prod-sql-use2-alpha-01: the negative control, and a real one — same fleet, same collectors, same + /// box, no skipped relaunches in the log. Its index_object_stats averages 4,097 ms, not 37,207. + /// + /// Unimodal throughout, and its query_store measurably so: that collector yields nothing on ALL + /// 1,551 of its runs here and pays 36 ms for each, so mean, p95 and max are the same 36 ms. This is + /// the fixture that has to stay quiet after #2460 — a change that makes the healthy control fire is + /// not a sharper signal, it is a broken one. + /// + private static IReadOnlyList<(string, double, double, int)> Apex() => new[] + { + C("query_stats", 2_562, 1), + C("procedure_stats", 461, 1), + C("other_one_minute_collectors", 635, 1), + C("plan_correction", 1_501, 5), + C("other_five_minute_collectors", 1_174, 5), + C("hourly_collectors", 777, 60), + C("index_object_stats", 4_097, 1440), + }; /// The #2296 measurement verbatim: ~101% of the minute — SATURATED, not a warning-log easter egg. [Fact] @@ -42,9 +117,16 @@ public void TheMotivatingServerReadsSaturated() Assert.Equal(SweepPressureClassifier.Saturated, pressure.Verdict); Assert.Equal(60_712, pressure.BusyMsPerMinute, 3); Assert.True(pressure.BusyPercent > 100.0); + + /* #2446: on a server saturated by per-minute collectors the two dimensions AGREE — every cycle is + the aligned cycle when everything runs every minute. The dimensions being orthogonal does not + mean they must disagree; it means neither can be derived from the other. */ + Assert.Equal(60_712, pressure.PeakCycleMs, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, pressure.PeakCycleRisk); + Assert.Equal("procedure_stats", pressure.PeakCollectorName); } - /// An ordinary in-region profile sits far below every threshold. + /// An ordinary in-region profile sits far below every threshold, on both dimensions. [Fact] public void AHealthyProfileReadsOk() { @@ -58,6 +140,8 @@ public void AHealthyProfileReadsOk() Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); Assert.True(pressure.BusyPercent < 5.0); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); + Assert.True(pressure.PeakCyclePercent < 10.0); } /// @@ -94,6 +178,14 @@ public void OnLoadAndZeroDurationCollectorsAreExcluded() Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); Assert.Equal(300, pressure.BusyMsPerMinute, 3); + + /* #2446: the SAME exclusion on the peak cycle, and it is load-bearing there too. An on-load + collector runs on connect, not in any scheduled cycle, so a 500-second one must not manufacture + a BODY_OVERRUN on a server whose recurring body is 300 ms. It must also not be able to become + the peak collector, which would name the wrong thing at the top of the block. */ + Assert.Equal(300, pressure.PeakCycleMs, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); + Assert.Equal("wait_stats", pressure.PeakCollectorName); } /// @@ -108,16 +200,344 @@ public void SlowCollectorsAreAmortizedByTheirOwnCadence() Assert.Equal(500, pressure.BusyMsPerMinute, 3); Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); + + /* #2446 does NOT undo that: the peak cycle takes the same collector UNdivided, and 30s of a 60s + budget on a server with nothing else scheduled still fits. The new dimension is not "any slow + collector is bad". */ + Assert.Equal(30_000, pressure.PeakCycleMs, 3); + Assert.Equal(50.0, pressure.PeakCyclePercent, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); } /// No collectors — a server before first collection — is OK with zero demand, never a verdict from nothing. [Fact] public void AnEmptyWindowReadsOkWithZeroDemand() { - var pressure = SweepPressureClassifier.Compute(Array.Empty<(string, double, int)>()); + var pressure = SweepPressureClassifier.Compute(Array.Empty<(string, double, double, int)>()); Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); Assert.Equal(0, pressure.BusyMsPerMinute); Assert.Equal(0, pressure.BusyPercent); + + /* #2446: and no peak collector invented out of an empty set. Null, not "" and not a zero-cost + name, so a caller rendering the block has something to branch on. */ + Assert.Equal(0, pressure.PeakCycleMs); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); + Assert.Null(pressure.PeakCollectorName); + } + + /* --------------------------------------------------------------------------------------------- + #2446. The case the amortized model answers correctly and an operator still reads as wrong. + --------------------------------------------------------------------------------------------- */ + + /// + /// prod-sql-use2-multi-49 verbatim: six skipped relaunches in three hours while sweep_pressure read + /// busy_percent 20.4, verdict OK, every collector HEALTHY. The verdict is RIGHT — sustained demand + /// genuinely fits — and the server genuinely overruns, because index_object_stats takes 37,207 ms of + /// a 60,000 ms body and its 1440-minute cadence amortizes that to 26 ms/min. This is the pin that + /// fails on dev. + /// + [Fact] + public void AnInfrequentHeavyCollectorReadsOkAndBodyOverrun() + { + var pressure = SweepPressureClassifier.Compute(Multi49()); + + /* The verdict is unmoved, deliberately. An operator told SATURATED because of a once-daily + collector learns to ignore the next SATURATED, and the capacity lever that verdict recommends + is the wrong lever for a schedule-shape problem. */ + Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); + Assert.Equal(12_248.4, pressure.BusyMsPerMinute, 1); + Assert.Equal(20.4, pressure.BusyPercent, 1); + + /* The second dimension sees what the first cannot. */ + Assert.Equal(73_408, pressure.PeakCycleMs, 3); + Assert.Equal(122.3, pressure.PeakCyclePercent, 1); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, pressure.PeakCycleRisk); + + /* And it names the collector heaviest_collectors structurally cannot: that list ranks by + amortized contribution, and on this server index_object_stats ranks LAST of the eight on that + key while owning 62% of a single body. */ + Assert.Equal("index_object_stats", pressure.PeakCollectorName); + Assert.Equal(37_207, pressure.PeakCollectorAvgDurationMs, 3); + Assert.Equal(1440, pressure.PeakCollectorFrequencyMinutes); + + /* #2460: with p95 == mean for every collector this profile is #2446 verbatim, which is the point + of pinning it that way — the tail statistic must change nothing where there is no tail. */ + Assert.Equal(37_207, pressure.PeakCollectorPeakRunMs, 3); + + var amortized = pressure.PeakCollectorAvgDurationMs / pressure.PeakCollectorFrequencyMinutes; + Assert.True(amortized < 30, $"amortized share was {amortized} ms/min"); + Assert.True(pressure.PeakCollectorAvgDurationMs / SweepPressureClassifier.SweepBudgetMs > 0.6); + } + + /// + /// prod-sql-use2-alpha-01: the cry-wolf control, and the reason the threshold is where it is. Same + /// fleet, same collector set, same 60s budget, no skipped relaunches — and it must stay quiet on BOTH + /// dimensions. A second signal that fires on a healthy server is worth less than no second signal. + /// + [Fact] + public void AGenuinelyHealthyServerTripsNeitherDimension() + { + var pressure = SweepPressureClassifier.Compute(Apex()); + + Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); + Assert.Equal(4_208.8, pressure.BusyMsPerMinute, 1); + Assert.Equal(7.0, pressure.BusyPercent, 1); + + Assert.Equal(11_207, pressure.PeakCycleMs, 3); + Assert.Equal(18.7, pressure.PeakCyclePercent, 1); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); + Assert.Equal(string.Empty, SweepPressureClassifier.FormatPeakCycleNote(pressure)); + } + + /// + /// The variable isolated. One collector's single-run average is the ONLY difference between the two + /// fixtures — multi-49's index_object_stats at 37,207 ms against apex's 4,097 — and it moves the peak + /// cycle across the budget while moving the verdict not at all (20.4% to 20.4%). That is the property + /// the whole change exists for: the two dimensions are independent, and neither is derivable from the + /// other. + /// + [Fact] + public void OnlyTheSingleRunCostSeparatesThemAndTheVerdictDoesNotMove() + { + var heavy = SweepPressureClassifier.Compute(Multi49()); + var light = SweepPressureClassifier.Compute(Multi49(indexObjectStatsMs: 4_097)); + + Assert.Equal(SweepPressureClassifier.Ok, heavy.Verdict); + Assert.Equal(SweepPressureClassifier.Ok, light.Verdict); + Assert.Equal(heavy.BusyPercent, light.BusyPercent, 1); + + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, heavy.PeakCycleRisk); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, light.PeakCycleRisk); + + /* With the daily collector down to 4s the heaviest single run on the server is query_store, and + 13,834 ms of a 60,000 ms budget is genuinely fine — the block still names a peak collector, it + just no longer names a problem. */ + Assert.Equal("query_store", light.PeakCollectorName); + } + + /// + /// The peak-cycle edge, inclusive at exactly the budget for the same reason the amortized edges are — + /// and pinned on a 1440-minute collector so the case cannot be confused with saturation: 60,000 ms + /// once a day is 42 ms/min, which is 0.07% of the budget. The verdict reads OK on both sides of the + /// edge while the risk flips, which is the orthogonality stated as an assertion. + /// + [Fact] + public void ThePeakCycleEdgeIsInclusiveAndIndependentOfTheVerdict() + { + var under = SweepPressureClassifier.Compute(new[] { C("index_object_stats", 59_999, 1440) }); + var at = SweepPressureClassifier.Compute(new[] { C("index_object_stats", 60_000, 1440) }); + + Assert.Equal(SweepPressureClassifier.PeakCycleFits, under.PeakCycleRisk); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, at.PeakCycleRisk); + + Assert.Equal(SweepPressureClassifier.Ok, under.Verdict); + Assert.Equal(SweepPressureClassifier.Ok, at.Verdict); + Assert.True(at.BusyPercent < 1.0, $"busy_percent was {at.BusyPercent}"); + } + + /// + /// The note is composed in the classifier, not at either SKU's tool, so the two cannot render the same + /// finding differently — and it carries the numbers that make it actionable rather than restating the + /// risk string. Empty on the healthy side, because a note that fires on FITS is how a signal teaches + /// people to skip it. + /// + [Fact] + public void ThePeakCycleNoteNamesTheCollectorAndItsShareOfOneBody() + { + var note = SweepPressureClassifier.FormatPeakCycleNote(SweepPressureClassifier.Compute(Multi49())); + + Assert.Contains("index_object_stats", note, StringComparison.Ordinal); + Assert.Contains("73,408 ms", note, StringComparison.Ordinal); + Assert.Contains("122.3%", note, StringComparison.Ordinal); + Assert.Contains("37,207 ms on a heavy run", note, StringComparison.Ordinal); + Assert.Contains("62.0% of the budget", note, StringComparison.Ordinal); + Assert.Contains("every 1440 minutes", note, StringComparison.Ordinal); + Assert.Contains("26 ms per minute", note, StringComparison.Ordinal); + + /* #2460's bimodal clause must NOT fire here: index_object_stats' runs all cost about the same, + and "its MEAN run is only 37,207 ms ... understates one body by 0 ms" is how a note teaches + people to stop reading notes. */ + Assert.DoesNotContain("bimodal", note, StringComparison.Ordinal); + + /* The sustained figure is quoted too: the note has to explain why the verdict beside it disagrees, + or it reads as the two contradicting each other. */ + Assert.Contains("20.4%", note, StringComparison.Ordinal); + + Assert.Equal(string.Empty, + SweepPressureClassifier.FormatPeakCycleNote(SweepPressureClassifier.Compute(Apex()))); + } + + /* --------------------------------------------------------------------------------------------- + #2460. One mean over two populations, and what the cycle built from it was missing. + --------------------------------------------------------------------------------------------- */ + + /// + /// The floor rule on its own, because it is the one piece of #2460 that is not obvious and the one + /// that decides whether this change can retract an answer #2446 already gave. + /// + /// p95 is NOT guaranteed to sit above the mean. A collector with 19 runs at 100 ms and one + /// pathological 1,220,000 ms run has a mean of 61,095 ms and a p95 of 100 ms — the percentile is doing + /// exactly its job, discarding the outlier, and the mean is the number that happens to remember it. + /// Taking the p95 unconditionally would compute a 100 ms aligned cycle for that collector where #2446 + /// computed 61,095 ms, turning a BODY_OVERRUN it correctly caught into a FITS. Flooring at the mean + /// makes #2460 monotonic: the aligned cycle can only ever go up. + /// + [Fact] + public void ThePeakRunIsFlooredAtTheMeanSoTheCycleCanOnlyEverRise() + { + Assert.Equal(80_933, SweepPressureClassifier.PeakRunMs(13_834, 80_933), 3); + Assert.Equal(61_095, SweepPressureClassifier.PeakRunMs(61_095, 100), 3); + Assert.Equal(37_207, SweepPressureClassifier.PeakRunMs(37_207, 37_207), 3); + + var outlier = SweepPressureClassifier.Compute(new[] { C("index_object_stats", 61_095, 100, 1440) }); + + Assert.Equal(61_095, outlier.PeakCycleMs, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, outlier.PeakCycleRisk); + Assert.Equal(61_095, outlier.PeakCollectorPeakRunMs, 3); + } + + /// + /// prod-sql-use2-multi-49's query_store, as the store's own numbers force it. #2446 charged the aligned + /// cycle 13,834 ms for this collector — a mean blended 958/197 out of a 36 ms empty run and an ~80,933 ms + /// productive one, describing neither. Charged the p95 instead, the same server's aligned body goes from + /// 73,408 ms to 140,507 ms: the ~67,000 ms #2459 said out loud it was missing, found. + /// + /// And the collector NAMED changes with it, which is the half that would have closed #2446 on its + /// own. index_object_stats has the larger mean (37,207 against 13,834) and was therefore the peak + /// collector; query_store has by far the larger heavy run, runs every 5 minutes rather than once a day, + /// and is what actually puts that body over the budget. + /// + [Fact] + public void TheBimodalCollectorIsChargedItsHeavyRunAndBecomesThePeakCollector() + { + var pressure = SweepPressureClassifier.Compute(Multi49(queryStoreP95Ms: 80_933)); + + Assert.Equal(140_507, pressure.PeakCycleMs, 3); + Assert.Equal(234.2, pressure.PeakCyclePercent, 1); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, pressure.PeakCycleRisk); + + Assert.Equal("query_store", pressure.PeakCollectorName); + Assert.Equal(80_933, pressure.PeakCollectorPeakRunMs, 3); + Assert.Equal(13_834, pressure.PeakCollectorAvgDurationMs, 3); + Assert.Equal(5, pressure.PeakCollectorFrequencyMinutes); + + /* One run of this collector costs more than the entire budget by itself — which is the sentence + #2446 needed and could not reach, because 13,834 ms of 60,000 reads like 23% of a body. */ + Assert.True(pressure.PeakCollectorPeakRunMs > SweepPressureClassifier.SweepBudgetMs, + $"one heavy run was {pressure.PeakCollectorPeakRunMs} ms against a {SweepPressureClassifier.SweepBudgetMs} ms budget"); + } + + /// + /// The variable isolated again, in the OTHER dimension. Only query_store's p95 differs between the two + /// fixtures — the mean is 13,834 ms on both sides — so the sustained answer is bit-for-bit identical + /// while the aligned cycle moves by 67,099 ms. That is the property #2460 exists for, and it is also + /// the guard on the thing most likely to be "simplified" later: the verdict must keep amortizing the + /// MEAN. Sustained demand over a window IS the mean; a rate built from a tail claims work the server + /// never sustains. + /// + [Fact] + public void TheTailMovesThePeakCycleAndLeavesTheVerdictExactlyWhereItWas() + { + var blended = SweepPressureClassifier.Compute(Multi49()); + var measured = SweepPressureClassifier.Compute(Multi49(queryStoreP95Ms: 80_933)); + + Assert.Equal(blended.BusyMsPerMinute, measured.BusyMsPerMinute, 6); + Assert.Equal(blended.BusyPercent, measured.BusyPercent, 6); + Assert.Equal(SweepPressureClassifier.Ok, blended.Verdict); + Assert.Equal(SweepPressureClassifier.Ok, measured.Verdict); + + Assert.Equal(67_099, measured.PeakCycleMs - blended.PeakCycleMs, 3); + Assert.True(measured.PeakCycleMs > blended.PeakCycleMs, + "#2460 must never compute a smaller aligned cycle than #2446 did"); + } + + /// + /// The healthy control after the change, which matters more than the positive case. alpha-01's + /// query_store is measurably unimodal — it yields nothing on all 1,551 runs and pays 36 ms every time — + /// so nothing about it moves, and the server that logs no skipped relaunches still trips neither + /// dimension and still gets no note. A second signal that starts firing on the quiet server is worth + /// less than no second signal. + /// + [Fact] + public void TheHealthyControlIsUnmovedByTheTailStatistic() + { + var pressure = SweepPressureClassifier.Compute(Apex()); + + Assert.Equal(4_208.8, pressure.BusyMsPerMinute, 1); + Assert.Equal(11_207, pressure.PeakCycleMs, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleFits, pressure.PeakCycleRisk); + Assert.Equal(string.Empty, SweepPressureClassifier.FormatPeakCycleNote(pressure)); + } + + /// + /// The note has to explain the disagreement it is sitting next to, and after #2460 there are two of + /// them: the aligned cycle against the verdict, and the collector's heavy run against its own mean. + /// Both numbers appear, the gap is stated as a number rather than left for the reader to subtract, and + /// the amortized figure is still computed from the MEAN — 13,834 / 5, not 80,933 / 5 — because that is + /// what the verdict beside it is made of. + /// + [Fact] + public void TheNoteSaysTheMeanDescribesNeitherPopulation() + { + var note = SweepPressureClassifier.FormatPeakCycleNote( + SweepPressureClassifier.Compute(Multi49(queryStoreP95Ms: 80_933))); + + Assert.Contains("query_store", note, StringComparison.Ordinal); + Assert.Contains("140,507 ms", note, StringComparison.Ordinal); + Assert.Contains("234.2%", note, StringComparison.Ordinal); + Assert.Contains("80,933 ms on a heavy run", note, StringComparison.Ordinal); + Assert.Contains("134.9% of the budget", note, StringComparison.Ordinal); + Assert.Contains("every 5 minutes", note, StringComparison.Ordinal); + Assert.Contains("2,767 ms per minute", note, StringComparison.Ordinal); + + Assert.Contains("MEAN run is only 13,834 ms", note, StringComparison.Ordinal); + Assert.Contains("bimodal", note, StringComparison.Ordinal); + Assert.Contains("understates one body by 67,099 ms", note, StringComparison.Ordinal); + } + + /// + /// The threshold on that clause, pinned so it cannot drift into either uselessness. At exactly twice + /// the mean it fires — inclusive, like every other edge in this table — and a hair under it does not, + /// because run-to-run variance is not bimodality and a note that fires on ordinary variance is a note + /// nobody finishes reading. + /// + [Fact] + public void TheBimodalClauseNeedsARealGapNotOrdinaryVariance() + { + var atTheEdge = SweepPressureClassifier.Compute(new[] { C("a", 45_000, 90_000, 1) }); + var justUnder = SweepPressureClassifier.Compute(new[] { C("a", 45_000, 89_999, 1) }); + + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, atTheEdge.PeakCycleRisk); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, justUnder.PeakCycleRisk); + + Assert.Contains("bimodal", SweepPressureClassifier.FormatPeakCycleNote(atTheEdge), StringComparison.Ordinal); + Assert.DoesNotContain("bimodal", SweepPressureClassifier.FormatPeakCycleNote(justUnder), StringComparison.Ordinal); + } + + /// + /// A collector whose mean rounds away to nothing but whose tail does not is still part of the body. + /// The old guard skipped any collector with a non-positive mean, which after #2460 would have dropped + /// exactly the shape this change is about — the rare-and-enormous run — from the cycle entirely. The + /// on-load and genuinely-idle exclusions are unchanged. + /// + [Fact] + public void ACollectorWithNoMeanButARealTailStillCounts() + { + var pressure = SweepPressureClassifier.Compute(new[] + { + C("rare_and_enormous", 0, 70_000, 5), + C("never_runs", 0, 0, 1), + C("on_load_collector", 0, 500_000, 0), + }); + + Assert.Equal(70_000, pressure.PeakCycleMs, 3); + Assert.Equal(SweepPressureClassifier.PeakCycleBodyOverrun, pressure.PeakCycleRisk); + Assert.Equal("rare_and_enormous", pressure.PeakCollectorName); + + /* Nothing was measured as sustained demand, so the verdict says nothing — the two dimensions stay + independent even at this edge, and an on-load collector still cannot reach either number. */ + Assert.Equal(0, pressure.BusyMsPerMinute, 6); + Assert.Equal(SweepPressureClassifier.Ok, pressure.Verdict); } } diff --git a/Lite.Tests/TempDbCeilingStoreTests.cs b/Lite.Tests/TempDbCeilingStoreTests.cs new file mode 100644 index 000000000..8dd15288c --- /dev/null +++ b/Lite.Tests/TempDbCeilingStoreTests.cs @@ -0,0 +1,140 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System.Linq; +using System.Threading.Tasks; +using PerformanceMonitorLite.Analysis; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// #2515, Lite's half: the tempdb growth ceiling through the REAL DuckDB store and the REAL reads. +/// +/// Darling's TempDbCeilingStoreTests pins the same behaviour against live Postgres. Both are +/// needed and neither substitutes for the other: the two SKUs write their own SQL against their own engine, +/// share the TempDbSpaceInfo that does the arithmetic, and a column selected in one adapter but not +/// the other puts exactly one product silently back on the old denominator. +/// +/// These go through a real database rather than a source pin because the defect this guards against — +/// an ordinal off by one, a NULL arriving as something other than zero, a column added to the schema but not +/// to the read — is invisible to any assertion over query TEXT. +/// +public sealed class TempDbCeilingStoreTests : IClassFixture +{ + private readonly DuckDbInitializer _duckDb; + + public TempDbCeilingStoreTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + } + + /// + /// The migration is the Lite twin of Darling's V81 rung, and the version has to move with it or an + /// existing Lite database never gets the column and every read of it fails. + /// + [Fact] + public void TheSchemaVersionMovedWithTheColumn() + { + Assert.Equal(56, DuckDbInitializer.CurrentSchemaVersion); + + var ddl = DuckDbSchemaGenerator.CreateTable(PerformanceMonitor.Collectors.TempDbStatsCollector.Instance); + Assert.Contains("max_size_mb DECIMAL(18,2)", ddl, System.StringComparison.Ordinal); + } + + /// + /// The Azure shape from the issue's measurement, round-tripped: 59.75 MB reserved and 2.69 MB unallocated + /// inside four 16 MB files whose max_size sums to 65,536 MB. It must read as 0.09% full rather than + /// the 95.7% the allocation produces, which is the whole difference between silence and a page. + /// + [Fact] + public async Task TheAzureShape_RoundTripsItsCeiling_AndReadsAsEmpty() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.ClearTestDataAsync(); + await seeder.SeedTestServerAsync(); + await seeder.SeedTempDbAsync( + reservedMb: 59.75, unallocatedMb: 2.69, + userObjectMb: 5.44, internalObjectMb: 1.81, versionStoreMb: 0.01, + maxSizeMb: 65_536); + + var info = await new LocalDataService(_duckDb).GetLatestTempDbSpaceAsync(TestDataSeeder.TestServerId); + + Assert.NotNull(info); + Assert.Equal(65_536d, info!.MaxSizeMb, precision: 2); + Assert.Equal(0.0912, info.UsedPercent, precision: 4); + Assert.True(info.UsedPercent < 80, "62 MB allocated against a 65,536 MB cap must not clear the 80% default."); + } + + /// + /// Unlimited survives as -1 rather than being flattened, and takes the allocation as its denominator — + /// so every unlimited-growth on-prem and RDS target reports exactly the number it reports today. + /// + [Fact] + public async Task AnUnlimitedCeiling_SurvivesAsMinusOne_AndKeepsTheAllocationDenominator() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.ClearTestDataAsync(); + await seeder.SeedTestServerAsync(); + await seeder.SeedTempDbAsync(reservedMb: 800, unallocatedMb: 200, maxSizeMb: -1); + + var info = await new LocalDataService(_duckDb).GetLatestTempDbSpaceAsync(TestDataSeeder.TestServerId); + + Assert.Equal(-1d, info!.MaxSizeMb, precision: 2); + Assert.Equal(80d, info.UsedPercent, precision: 3); + } + + /// + /// And a row from before the migration, whose ceiling is genuinely NULL. It has to arrive as 0 — the + /// "not measured" state — rather than as a zero-megabyte cap, which would divide by nothing. This is what + /// every historical row in a real Lite database looks like the moment the upgrade lands. + /// + [Fact] + public async Task ANullCeiling_ReadsAsNotMeasured_AndTheNumberDoesNotMove() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.ClearTestDataAsync(); + await seeder.SeedTestServerAsync(); + await seeder.SeedTempDbAsync(reservedMb: 800, unallocatedMb: 200); + + var info = await new LocalDataService(_duckDb).GetLatestTempDbSpaceAsync(TestDataSeeder.TestServerId); + + Assert.Equal(0d, info!.MaxSizeMb, precision: 2); + Assert.Equal(80d, info.UsedPercent, precision: 3); + } + + /// + /// The ANALYSIS surface has to agree with the alert, or analyze_server scores the same Azure target + /// at 96% full while the pager stays quiet — two answers about one server, which is the shape of defect + /// this whole change exists to remove. The fact carries its own denominator, so it needed the same fix. + /// + [Fact] + public async Task TheAnalysisFact_ScoresAgainstTheCeilingToo() + { + using var seeder = new TestDataSeeder(_duckDb); + await seeder.ClearTestDataAsync(); + await seeder.SeedTestServerAsync(); + await seeder.SeedTempDbAsync( + reservedMb: 59.75, unallocatedMb: 2.69, + userObjectMb: 5.44, internalObjectMb: 1.81, versionStoreMb: 0.01, + maxSizeMb: 65_536); + + var facts = await new DuckDbFactCollector(_duckDb).CollectFactsAsync(TestDataSeeder.CreateTestContext()); + var tempdb = facts.First(f => f.Key == "TEMPDB_USAGE"); + + Assert.Equal(0.000912, tempdb.Value, precision: 6); + Assert.Equal(65_536d, tempdb.Metadata["max_size_mb"], precision: 2); + + /* The severity arm concerns at 0.75 and criticals at 0.90, so the corrected fraction scores nothing — + where the allocation fraction (0.957) would have pinned it at the top of the scale. */ + Assert.True(tempdb.Value < 0.75); + } +} diff --git a/Lite.Tests/TempDbStatsCollectorDefinitionTests.cs b/Lite.Tests/TempDbStatsCollectorDefinitionTests.cs index 0afc93a4e..07ee4c1c4 100644 --- a/Lite.Tests/TempDbStatsCollectorDefinitionTests.cs +++ b/Lite.Tests/TempDbStatsCollectorDefinitionTests.cs @@ -38,6 +38,10 @@ public void PayloadColumns_MatchSchemaOrder() "total_sessions_using_tempdb", "top_session_id", "top_session_tempdb_mb", + /* #2515, APPENDED. Both stores generate their DDL from this list in order and both row + writers are positional, so the ceiling could only ever go last — inserting it beside + unallocated_mb, where it belongs semantically, would re-map every historical row. */ + "max_size_mb", }, names); } @@ -52,11 +56,37 @@ public void Query_TargetsBothTempDbDmvs() Assert.Equal("tempdb_stats", TempDbStatsCollector.Instance.TargetTable); } + /// + /// #2515: the ceiling comes from tempdb's own catalog, and the two things that make it the RIGHT + /// ceiling are both in the query rather than in the reader — so they can only be pinned here. + /// + /// LOG files are excluded because dm_db_file_space_usage, which supplies every other + /// column, reports DATA allocation: folding the log's cap into the same denominator would understate + /// usage on every server, not just Azure. And max_size is an int of 8 KB pages that tops + /// out at 16 TB per file, so a wide tempdb can overflow a plain SUM — the widen has to happen + /// before the sum, not after it. + /// + [Fact] + public void Query_ReadsTheCeilingFromTheRowsFilesOnly_AndSumsItWideEnough() + { + var queryText = TempDbStatsCollector.Instance.BuildQuery(CollectorTestContext.Make(new RecordingCollectorDeltaCalculator())).Text; + + Assert.Contains("tempdb.sys.database_files AS df", queryText, System.StringComparison.Ordinal); + Assert.Contains("WHERE df.type = 0 /*ROWS*/", queryText, System.StringComparison.Ordinal); + Assert.Contains("SUM(CONVERT(bigint, df.max_size))", queryText, System.StringComparison.Ordinal); + + /* -1 on any one data file means tempdb as a whole grows without limit, and MIN is what finds it. */ + Assert.Contains("WHEN MIN(df.max_size) = -1", queryText, System.StringComparison.Ordinal); + + /* House convention, and it is load-bearing on a query that now carries a second aggregate. */ + Assert.Contains("OPTION(RECOMPILE)", queryText, System.StringComparison.Ordinal); + } + [Fact] public async Task ReadAsync_CombinesTwoResultSets_IntoOneRow() { using var reader = FakeCollectorDataReader.WithResultSets( - new[] { new object[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m } }, + new[] { new object[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 65536.0m } }, new[] { new object[] { 55, 12.25m, 9L } }); var context = CollectorTestContext.Make(new RecordingCollectorDeltaCalculator()); @@ -64,7 +94,46 @@ public async Task ReadAsync_CombinesTwoResultSets_IntoOneRow() var rows = await TempDbStatsCollector.Instance.ReadAsync(reader, context, CancellationToken.None); var row = Assert.Single(rows); - Assert.Equal(new TempDbStatsCollector.Row(1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m), row); + Assert.Equal(new TempDbStatsCollector.Row(1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m, 65536.0m), row); + } + + /// + /// #2515: a tempdb with no ROWS files visible makes the ceiling subquery return NULL, and NULL is not a + /// ceiling of zero. It has to land on 0 — the "not measured" state every consumer answers by dividing by + /// the allocation, exactly as it did before this column existed. A zero cap would divide the alert's + /// percentage by nothing at all. + /// + [Fact] + public async Task ReadAsync_NullCeiling_ReadsAsNotMeasured_NotAsAZeroCap() + { + using var reader = FakeCollectorDataReader.WithResultSets( + new[] { new object[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m, System.DBNull.Value } }, + new[] { new object[] { 55, 12.25m, 9L } }); + + var context = CollectorTestContext.Make(new RecordingCollectorDeltaCalculator()); + + var rows = await TempDbStatsCollector.Instance.ReadAsync(reader, context, CancellationToken.None); + + Assert.Equal(0m, Assert.Single(rows).MaxSizeMb); + } + + /// + /// And the unlimited answer survives the read AS -1 rather than being flattened to 0. They take the same + /// denominator, but they are different facts — "this tempdb has no ceiling" versus "nobody looked" — and + /// the alert detail says which. + /// + [Fact] + public async Task ReadAsync_UnlimitedCeiling_StaysMinusOne() + { + using var reader = FakeCollectorDataReader.WithResultSets( + new[] { new object[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m, -1m } }, + new[] { new object[] { 55, 12.25m, 9L } }); + + var context = CollectorTestContext.Make(new RecordingCollectorDeltaCalculator()); + + var rows = await TempDbStatsCollector.Instance.ReadAsync(reader, context, CancellationToken.None); + + Assert.Equal(-1m, Assert.Single(rows).MaxSizeMb); } [Fact] @@ -87,11 +156,80 @@ public void WritePayload_EmitsSchemaOrder_NoDeltas() { var deltas = new RecordingCollectorDeltaCalculator(); var writer = new RecordingCollectorRowWriter(); - var row = new TempDbStatsCollector.Row(1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m); + var row = new TempDbStatsCollector.Row(1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m, 65536.0m); TempDbStatsCollector.Instance.WritePayload(row, writer, CollectorTestContext.Make(deltas)); - Assert.Equal(new object?[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m }, writer.Values); + Assert.Equal(new object?[] { 1.5m, 2.5m, 3.5m, 7.5m, 10.0m, 9L, 55, 12.25m, 65536.0m }, writer.Values); Assert.Empty(deltas.Calls); } + + /// + /// #2512: the whole pipeline on an AZURE SQL DATABASE target, which this collector was gated off + /// until the gate's stated reason was checked and found false. + /// + /// Three things the gate meant nobody had ever exercised, and each fails differently: + /// + /// The query is target-independent. If Azure needed a variant, BuildQuery would + /// have to branch on the target — it does not, and the measurement says it does not need to. + /// Column typing. sys.dm_db_session_space_usage.session_id is smallint, so + /// the driver hands back a short. The definition reads it through + /// Convert.ToInt32(GetValue(0)) rather than GetInt32 precisely for that, and the payload + /// column is top_session_id INTEGER — so the widening has to happen and has to be pinned. A + /// fixture that feeds an int (as the parity pin above does) can never see this. + /// The fan-out shape. RunsPerDatabase is false, so on Azure SQL DB this takes the + /// plain single-connection path rather than the per-database loop file_io_stats and + /// index_object_stats take. Per #2220 a registration that names a database is scoped to that + /// database, so this collects one tempdb snapshot per registration — not N of them. + /// + /// + /// Values are the ones actually measured on GP_S_Gen5_2 (EngineEdition 5) on 2026-08-22, + /// not invented ones, so the row this asserts is a row the platform really produced. + /// + [Fact] + public async Task AzureSqlDb_MeasuredValues_ComposeThroughToThePayload() + { + var azure = new CollectorTargetInfo { IsAzureSqlDb = true, SqlMajorVersion = 12 }; + var context = CollectorTestContext.Make(new RecordingCollectorDeltaCalculator(), isAzureSqlDb: true); + + /* One query for every target — no Azure variant, which is the claim the gate rested on. */ + Assert.Equal( + TempDbStatsCollector.Instance.BuildQuery(CollectorTestContext.Make(new RecordingCollectorDeltaCalculator())).Text, + TempDbStatsCollector.Instance.BuildQuery(context).Text, + System.StringComparer.Ordinal); + + /* Plain path, not the Azure per-database loop. */ + Assert.False(TempDbStatsCollector.Instance.RunsPerDatabase(azure)); + Assert.Null(TempDbStatsCollector.Instance.BuildEnumerationQuery(context)); + + using var reader = FakeCollectorDataReader.WithResultSets( + /* result set 1: user 5.44 / internal 1.81 / version 0.00 / total 7.25 / unallocated 54.19 MB */ + new[] { new object[] { 5.44m, 1.81m, 0.00m, 7.25m, 54.19m, 65536.00m } }, + /* result set 2: session 74 as SMALLINT, 0.13 MB, 1 session over threshold as COUNT_BIG */ + new[] { new object[] { (short)74, 0.13m, 1L } }); + + var rows = await TempDbStatsCollector.Instance.ReadAsync(reader, context, CancellationToken.None); + + var row = Assert.Single(rows); + /* + The ceiling is the ninth member and the reason this test exists on Azure at all: these five + allocation figures are the real GP_S_Gen5_2 measurement, where 7.25 MB reserved inside 62.44 MB + allocated reads 10% full -- and against the 65,536 MB ROWS ceiling the platform will actually + grow to, 0.01%. Asserting the ceiling composes through is what stops the payload carrying the + allocation alone and the alert dividing by the wrong number again (#2515). + */ + Assert.Equal( + new TempDbStatsCollector.Row(5.44m, 1.81m, 0.00m, 7.25m, 54.19m, 1L, 74, 0.13m, 65536.00m), + row); + + var writer = new RecordingCollectorRowWriter(); + TempDbStatsCollector.Instance.WritePayload(row, writer, context); + + /* Positional AND typed: 74 must arrive as int, not short, or the INTEGER column takes a + narrowed write on the Darling COPY path. */ + Assert.Equal(TempDbStatsCollector.Instance.PayloadColumns.Count, writer.Values.Count); + Assert.Equal(new object?[] { 5.44m, 1.81m, 0.00m, 7.25m, 54.19m, 1L, 74, 0.13m, 65536.00m }, writer.Values); + Assert.IsType(writer.Values[6]); + Assert.IsType(writer.Values[5]); + } } diff --git a/Lite.Tests/TestDataSeeder.cs b/Lite.Tests/TestDataSeeder.cs index 7d7729891..cca1af50a 100644 --- a/Lite.Tests/TestDataSeeder.cs +++ b/Lite.Tests/TestDataSeeder.cs @@ -1367,10 +1367,13 @@ INSERT INTO file_io_stats } /// - /// Seeds tempdb_stats across 16 collection points. + /// Seeds tempdb_stats across 16 collection points. is the #2515 growth + /// ceiling: null seeds NULL, which is what every row collected before the v56 migration looks like and + /// the state the callers here want by default. /// internal async Task SeedTempDbAsync(double reservedMb, double unallocatedMb, - double userObjectMb = 0, double internalObjectMb = 0, double versionStoreMb = 0) + double userObjectMb = 0, double internalObjectMb = 0, double versionStoreMb = 0, + double? maxSizeMb = null) { if (userObjectMb == 0) userObjectMb = reservedMb * 0.6; if (internalObjectMb == 0) internalObjectMb = reservedMb * 0.3; @@ -1387,8 +1390,8 @@ internal async Task SeedTempDbAsync(double reservedMb, double unallocatedMb, INSERT INTO tempdb_stats (collection_id, collection_time, server_id, server_name, user_object_reserved_mb, internal_object_reserved_mb, - version_store_reserved_mb, total_reserved_mb, unallocated_mb) -VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9)"; + version_store_reserved_mb, total_reserved_mb, unallocated_mb, max_size_mb) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10)"; var t = TestPeriodStart.AddMinutes(i * 15); cmd.Parameters.Add(new DuckDBParameter { Value = _nextId-- }); @@ -1400,6 +1403,7 @@ INSERT INTO tempdb_stats cmd.Parameters.Add(new DuckDBParameter { Value = versionStoreMb }); cmd.Parameters.Add(new DuckDBParameter { Value = reservedMb }); cmd.Parameters.Add(new DuckDBParameter { Value = unallocatedMb }); + cmd.Parameters.Add(new DuckDBParameter { Value = maxSizeMb.HasValue ? maxSizeMb.Value : (object)DBNull.Value }); await cmd.ExecuteNonQueryAsync(); } diff --git a/Lite.Tests/TrendEmptyParityToolTests.cs b/Lite.Tests/TrendEmptyParityToolTests.cs new file mode 100644 index 000000000..40814645b --- /dev/null +++ b/Lite.Tests/TrendEmptyParityToolTests.cs @@ -0,0 +1,238 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Text.Json; +using System.Threading.Tasks; +using DuckDB.NET.Data; +using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Mcp; +using PerformanceMonitorLite.Models; +using PerformanceMonitorLite.Services; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// The three tools where the two SKUs disagreed (#2485): get_memory_trend, get_file_io_trend +/// and get_query_duration_trend returned a BARE empty array on Lite while their Darling twins +/// returned a status envelope. Same tool name, same client, two different answers depending on which SKU it +/// was pointed at — a parity break independent of the issue's main point and arguably the more urgent half. +/// +/// Darling's envelope was not the target either. "No memory trend data available" is true both of a +/// window that was simply quiet and of a server the collector has never touched, and those want opposite +/// next moves — widen the window, versus go find out why collection is not running, which widening will +/// never fix. Both SKUs now make that distinction, in the same two sentences. +/// +public sealed class TrendEmptyParityToolTests : IClassFixture, IDisposable +{ + private const string ServerName = "TrendEmptySrv"; + + private readonly DuckDbInitializer _duckDb; + private readonly string _configDir; + private readonly ServerManager _serverManager; + private readonly int _serverId; + private DuckDBConnection? _seedConn; + private long _nextId = 830000; + + public TrendEmptyParityToolTests(SharedDuckDbFixture fixture) + { + fixture.ResetData(); + _duckDb = fixture.DuckDb; + + _configDir = Path.Combine(Path.GetTempPath(), "pmlite-trendempty-" + Guid.NewGuid().ToString("N")); + Directory.CreateDirectory(_configDir); + _serverManager = new ServerManager(_configDir); + + var server = new ServerConnection + { + Id = Guid.NewGuid().ToString(), + ServerName = ServerName, + IsEnabled = true, + }; + _serverManager.AddServer(server); + + /* Derived, not stored -- seeding under a hardcoded id would write rows the tool looks past. */ + _serverId = RemoteCollectorService.GetDeterministicHashCode( + RemoteCollectorService.GetServerNameForStorage(server)); + } + + public void Dispose() + { + _seedConn?.Dispose(); + try { Directory.Delete(_configDir, recursive: true); } catch (IOException) { /* temp dir */ } + } + + [Fact] + public async Task MemoryTrend_NeverCollected_AndAQuietWindow_AreDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + AssertNeverCollected(await McpMemoryTools.GetMemoryTrend(service, _serverManager, ServerName, 4)); + + await SeedMemoryAsync(DateTime.UtcNow.AddHours(-48)); + AssertQuietWindow(await McpMemoryTools.GetMemoryTrend(service, _serverManager, ServerName, 1)); + + await SeedMemoryAsync(DateTime.UtcNow.AddMinutes(-10)); + AssertPayload(await McpMemoryTools.GetMemoryTrend(service, _serverManager, ServerName, 4)); + } + + [Fact] + public async Task FileIoTrend_NeverCollected_AndAQuietWindow_AreDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + AssertNeverCollected(await McpIoTools.GetFileIoTrend(service, _serverManager, ServerName, 4)); + + await SeedFileIoAsync(DateTime.UtcNow.AddHours(-48)); + AssertQuietWindow(await McpIoTools.GetFileIoTrend(service, _serverManager, ServerName, 1)); + + await SeedFileIoAsync(DateTime.UtcNow.AddMinutes(-10)); + AssertPayload(await McpIoTools.GetFileIoTrend(service, _serverManager, ServerName, 4)); + } + + [Fact] + public async Task QueryDurationTrend_NeverCollected_AndAQuietWindow_AreDifferentAnswers() + { + var service = new LocalDataService(_duckDb); + + AssertNeverCollected(await McpQueryTools.GetQueryDurationTrend(service, _serverManager, ServerName, 4)); + + await SeedQueryAsync(DateTime.UtcNow.AddHours(-48)); + AssertQuietWindow(await McpQueryTools.GetQueryDurationTrend(service, _serverManager, ServerName, 1)); + + await SeedQueryAsync(DateTime.UtcNow.AddMinutes(-10)); + AssertPayload(await McpQueryTools.GetQueryDurationTrend(service, _serverManager, ServerName, 4)); + } + + /// Nothing has ever been stored for this server: NOT an empty window, and widening it would + /// never help. + private static void AssertNeverCollected(string payload) + { + var root = JsonDocument.Parse(payload).RootElement; + Assert.Equal("unavailable", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("EVER", text, StringComparison.Ordinal); + Assert.Contains("not an empty window", text, StringComparison.Ordinal); + Assert.DoesNotContain("widen hours_back", text, StringComparison.Ordinal); + } + + /// The server collected and this window is simply quiet — the opposite next move, and it must + /// not share wording with the case above. + private static void AssertQuietWindow(string payload) + { + var root = JsonDocument.Parse(payload).RootElement; + Assert.Equal("empty", root.GetProperty("status").GetString()); + var text = root.GetProperty("message").GetString()!; + Assert.Contains("widen hours_back", text, StringComparison.Ordinal); + Assert.DoesNotContain("EVER", text, StringComparison.Ordinal); + } + + private static void AssertPayload(string payload) + { + var root = JsonDocument.Parse(payload).RootElement; + Assert.False(root.TryGetProperty("status", out _)); + Assert.True(root.GetProperty("trend").GetArrayLength() > 0); + } + + private async Task SeedConnectionAsync() + { + if (_seedConn is null) + { + _seedConn = _duckDb.CreateConnection(); + await _seedConn.OpenAsync(); + } + return _seedConn; + } + + private async Task SeedMemoryAsync(DateTime collectionTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO memory_stats + (collection_id, collection_time, server_id, server_name, + total_physical_memory_mb, available_physical_memory_mb, + target_server_memory_mb, total_server_memory_mb, buffer_pool_mb, plan_cache_mb) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = 65536.0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 8192.0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 49152.0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 40000.0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 35000.0 }); + cmd.Parameters.Add(new DuckDBParameter { Value = 5000.0 }); + await cmd.ExecuteNonQueryAsync(); + } + + /* delta_reads above zero on purpose: the trend's top_files CTE requires read or write activity, so a + row with zero deltas would leave the window empty for a reason that has nothing to do with #2485. */ + private async Task SeedFileIoAsync(DateTime collectionTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO file_io_stats + (collection_id, collection_time, server_id, server_name, + database_name, file_name, file_type, size_mb, + num_of_reads, num_of_writes, read_bytes, write_bytes, + io_stall_read_ms, io_stall_write_ms, + delta_reads, delta_writes, delta_stall_read_ms, delta_stall_write_ms) +VALUES ($1, $2, $3, $4, $5, $6, $7, 0, + $8, $9, 0, 0, $10, $11, $12, $13, $14, $15)"; + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified) }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = "AppDb" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "app.mdf" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "ROWS" }); + cmd.Parameters.Add(new DuckDBParameter { Value = 500L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 200L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 2500L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 400L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 500L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 200L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 2500L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 400L }); + await cmd.ExecuteNonQueryAsync(); + } + + private async Task SeedQueryAsync(DateTime collectionTimeUtc) + { + using var readLock = _duckDb.AcquireReadLock(); + var connection = await SeedConnectionAsync(); + using var cmd = connection.CreateCommand(); + cmd.CommandText = @" +INSERT INTO query_stats + (collection_id, collection_time, server_id, server_name, database_name, + query_hash, sql_handle, last_execution_time, delta_execution_count, + delta_worker_time, delta_elapsed_time, query_text) +VALUES ($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11, $12)"; + var naive = DateTime.SpecifyKind(collectionTimeUtc, DateTimeKind.Unspecified); + cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = _serverId }); + cmd.Parameters.Add(new DuckDBParameter { Value = ServerName }); + cmd.Parameters.Add(new DuckDBParameter { Value = "AppDb" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "0xEMPTYTRENDHASH" }); + cmd.Parameters.Add(new DuckDBParameter { Value = "0xSQLH" }); + cmd.Parameters.Add(new DuckDBParameter { Value = naive }); + cmd.Parameters.Add(new DuckDBParameter { Value = 10L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 10000L }); + cmd.Parameters.Add(new DuckDBParameter { Value = 20000L }); + cmd.Parameters.Add(new DuckDBParameter { Value = "SELECT * FROM Orders" }); + await cmd.ExecuteNonQueryAsync(); + } +} diff --git a/Lite.Tests/WatermarkPolicyTests.cs b/Lite.Tests/WatermarkPolicyTests.cs index b610e29b4..d5699599e 100644 --- a/Lite.Tests/WatermarkPolicyTests.cs +++ b/Lite.Tests/WatermarkPolicyTests.cs @@ -7,6 +7,11 @@ */ using System; +using System.Collections.Generic; +using System.IO; +using System.Linq; +using System.Runtime.CompilerServices; +using System.Text.RegularExpressions; using PerformanceMonitor.Collectors; using Xunit; @@ -131,4 +136,164 @@ public void ReadFloor_OnDefault_IsNull() { Assert.Null(WatermarkPolicy.ReadFloor(default)); } + + /// + /// The horizon's NUMBER lives in and nowhere else. + /// + /// Why this needed asserting. #2102 moved the horizon from 24 hours to one. The constant + /// moved, the behaviour moved, and the operator-facing WARNINGs moved with it because they interpolate + /// MaxCatchup.TotalHours instead of restating it. What did not move were nine comments across five + /// files, each calling it "the 24h catch-up clamp". Nothing misbehaved, no test went red, and no log line + /// disagreed — so there was no instrument in this repo that could see it. It was found the way stale prose + /// always is: someone read the source, believed it, and reasoned from a 24-hour lever that had not existed + /// for months (#2468, filed on exactly that premise). + /// + /// So the assertion is not "the comments say 1h" — that is the same defect with a fresher number in + /// it, and it would go stale the next time the horizon moves. It is that no source discussing the clamp + /// states an hours figure at all. The number has one home, and prose has to point at it. + /// + /// If this fails on something that has nothing to do with the horizon, read this first. The + /// trigger list is broader than the concept, on purpose and at a known cost. "catch-up" is ordinary English + /// in this codebase — the compression backlog catches up, a cold-start sweep body catches up, WAL replay + /// catches up — and none of those is . A future comment that pairs + /// one of them with an unrelated hours figure inside the window will fail this test, and it will be a + /// SCOPE problem in the list below, not the horizon drifting. Fix it by narrowing the trigger, never by + /// widening the window's tolerance for figures. + /// + /// Narrowing to "catch-up clamp" is the obvious escape and it does not work: two of the twelve sites + /// this caught named the clamp without the word "clamp" adjacent ("roll past 24h on their own", two lines + /// under "legitimate catch-up"), and two more named it without the words "catch-up" at all + /// ("24h-clamped outage holes"). A precise trigger would have missed the ones that hide best. The false + /// positive is the price of that, and it is the cheaper failure — it is loud, and it is this paragraph. + /// + [Fact] + public void TheCatchUpHorizon_IsWrittenDownInExactlyOnePlace() + { + /* The concept, wherever it is named. "clamp fires" / "-clamped" / "clamp-bounded" are in the list + because the site in DarlingWorker that carried this defect named the clamp without ever using the + words "catch-up" — a trigger list built only from the obvious phrase would have missed it. */ + const string mentions = @"catch-up|ClampCatchup|MaxCatchup|clamp fires|clamp-bounded|-clamped"; + + /* An hours figure in any spelling. Minutes are legitimate and common nearby (the 60-minute + first-run window, the 15-minute adaptive floor), so the pattern is deliberately hours-only. */ + const string hours = @"\b\d+\s*-?\s*(h\b|hr\b|hrs\b|hours?\b|hour-)"; + + var offenders = new List(); + + foreach (var file in ClampSources()) + { + var text = File.ReadAllText(file).Replace("\r\n", "\n", StringComparison.Ordinal); + + foreach (Match mention in Regex.Matches(text, mentions, RegexOptions.IgnoreCase)) + { + /* A window rather than a line: these are wrapped block comments, and the figure and the + concept routinely land on different lines. */ + var from = Math.Max(0, mention.Index - 220); + var to = Math.Min(text.Length, mention.Index + 220); + var figure = Regex.Match(text[from..to], hours, RegexOptions.IgnoreCase); + + if (figure.Success) + { + offenders.Add($"{Path.GetFileName(file)}: '{figure.Value}' near '{mention.Value}'"); + } + } + } + + Assert.True( + offenders.Count == 0, + "the catch-up horizon's number belongs to WatermarkPolicy.MaxCatchup and nowhere else — " + + "name the concept and let the constant carry the figure: " + string.Join("; ", offenders.Distinct())); + } + + /// Every shipped source that could describe the clamp: the shared collectors, the Darling + /// service's runners, and Lite's twin of them. WatermarkPolicy.cs is excluded because it is the + /// one place allowed to write the number down — including its own record of what the horizon used to + /// be, which is history worth keeping rather than prose that has gone stale. + private static IEnumerable ClampSources() + { + var repo = RepoRoot(); + + var roots = new[] + { + Path.Combine(repo, "PerformanceMonitor.Collectors"), + Path.Combine(repo, "Darling", "PerformanceMonitor.Darling.Service"), + Path.Combine(repo, "Lite", "Services"), + }; + + foreach (var root in roots) + { + Assert.True(Directory.Exists(root), $"{root} is gone — find where it moved before editing this test"); + + foreach (var file in Directory.EnumerateFiles(root, "*.cs", SearchOption.AllDirectories)) + { + /* Skip build output. A local build drops generated .cs under obj/, and a scan that reads + them is asserting about artifacts rather than about the source anyone will edit. */ + if (file.Contains($"{Path.DirectorySeparatorChar}obj{Path.DirectorySeparatorChar}", StringComparison.Ordinal) + || file.Contains($"{Path.DirectorySeparatorChar}bin{Path.DirectorySeparatorChar}", StringComparison.Ordinal)) + { + continue; + } + + if (!string.Equals(Path.GetFileName(file), "WatermarkPolicy.cs", StringComparison.Ordinal)) + { + yield return file; + } + } + } + } + + /// + /// A guard that CI skips on the PRs it guards is a guard that has silently stopped guarding — the + /// same failure the horizon pin above exists to close, one layer out. + /// + /// Raised by review on #2471. That pin lives in Lite.Tests and walks the Darling service + /// tree, but build.yml's "Run Lite tests" step gates on lite / core / root. + /// core covers PerformanceMonitor.Collectors and lite covers Lite/Services, + /// so two of the three trees were fine — but Darling/PerformanceMonitor.Darling.Service belongs + /// only to darling, and three of the twelve sites the pin was written for live there. A + /// Darling-only PR reintroducing one would have fired darling, skipped this suite, and been + /// caught a day later by the nightly. + /// + /// So the filter entry is load-bearing, and a filter entry is exactly the kind of thing that gets + /// tidied away by someone trimming what looks like an over-broad path. It is asserted rather than + /// commented — Darling.Tests' CI-worker-sizing guards already set the precedent for a test + /// reading these workflows. + /// + [Fact] + public void TheLiteSuite_RunsOnEveryTreeTheHorizonPinScans() + { + var yaml = File.ReadAllText(Path.Combine(RepoRoot(), ".github", "workflows", "build.yml")) + .Replace("\r\n", "\n", StringComparison.Ordinal); + + var at = yaml.IndexOf("\n lite:\n", StringComparison.Ordinal); + Assert.True(at > 0, "build.yml's 'lite' path filter is gone — find where it moved before editing this test"); + + /* The block runs to the next area key at the same indent; its own entries are indented deeper. */ + var rest = yaml[(at + 1)..]; + var next = Regex.Match(rest, "\n [a-z_]+:\n"); + var block = next.Success ? rest[..next.Index] : rest; + + Assert.Contains("Darling/PerformanceMonitor.Darling.Service/**/!(*.md)", block, StringComparison.Ordinal); + + /* The step that consumes it. If "Run Lite tests" ever stops reading `lite`, the entry above is + decoration and this test is the only thing that would notice. */ + var step = yaml.IndexOf("name: Run Lite tests", StringComparison.Ordinal); + Assert.True(step > 0, "the 'Run Lite tests' step is gone — find where it moved before editing this test"); + Assert.Contains("steps.filter.outputs.lite == 'true'", yaml[step..(step + 400)], StringComparison.Ordinal); + } + + /// The repo root, located by walking up from this file's compile-time path — the same idiom + /// the other source-scanning suites use. + private static string RepoRoot([CallerFilePath] string thisFile = "") + { + for (var dir = new DirectoryInfo(Path.GetDirectoryName(thisFile)!); dir is not null; dir = dir.Parent) + { + if (Directory.Exists(Path.Combine(dir.FullName, "PerformanceMonitor.Collectors"))) + { + return dir.FullName; + } + } + + throw new DirectoryNotFoundException($"could not locate the repo root walking up from {thisFile}"); + } } diff --git a/Lite.Tests/XeSessionPermissionLogLevelTests.cs b/Lite.Tests/XeSessionPermissionLogLevelTests.cs new file mode 100644 index 000000000..9925ea847 --- /dev/null +++ b/Lite.Tests/XeSessionPermissionLogLevelTests.cs @@ -0,0 +1,94 @@ +/* + * Copyright (c) 2026 Erik Darling, Darling Data LLC + * + * This file is part of the SQL Server Performance Monitor Lite. + * + * Licensed under the MIT License. See LICENSE file in the project root for full license information. + */ + +using System; +using System.IO; +using System.Linq; +using System.Text.RegularExpressions; +using Xunit; + +namespace PerformanceMonitorLite.Tests; + +/// +/// A denied XE session logs at Warn, not Error. +/// +/// Why this is worth pinning. Declining to grant ALTER ANY EVENT SESSION is a +/// least-privilege choice a customer is entitled to make (#1823), and the collector already treats it as +/// one: it classifies the outcome as PERMISSIONS and flags the collector so the scheduler stops +/// retrying it for the session. The logging did not agree. A field log from #2594 carried three consecutive +/// [ERROR] lines — two from the XE layer, one from the collector — for a login that was simply not +/// granted the permission, while every other permission denial in the same method logs at Warn. Someone +/// reading that log reasonably concludes something is broken. +/// +/// Anchored on the logger call adjacent to the permission test rather than on the message text, so a +/// reworded message does not fail this and a level change does. +/// +public class XeSessionPermissionLogLevelTests +{ + private static readonly string[] XeSources = + { + "RemoteCollectorService.Deadlocks.cs", + "RemoteCollectorService.BlockedProcessReport.cs", + }; + + [Fact] + public void ADeniedXeSession_LogsAtWarn_NotError() + { + foreach (var file in XeSources) + { + var source = ReadServiceSource(file); + + var guards = Regex.Matches( + source, + @"if \(SqlServerPermissionErrors\.IsPermissionDenied\(ex\.Number\)\)\s*\{\s*AppLogger\.(\w+)\(", + RegexOptions.Singleline); + + Assert.True( + guards.Count > 0, + $"{file} no longer routes a denied XE session through IsPermissionDenied, so a least-privilege " + + "login is back to logging as a fault."); + + foreach (Match guard in guards) + { + Assert.Equal("Warn", guard.Groups[1].Value); + } + } + } + + /// + /// The Error arm must survive. Downgrading everything would hide a genuine XE failure — a session that + /// cannot start for a reason that is not permissions is exactly what #1086 made loud. + /// + [Fact] + public void AnXeFailureThatIsNotPermissions_StillLogsAtError() + { + foreach (var file in XeSources) + { + var source = ReadServiceSource(file); + + Assert.Contains("AppLogger.Error(\"XeSession\"", source, StringComparison.Ordinal); + } + } + + private static string ReadServiceSource(string fileName) + { + var dir = new DirectoryInfo(AppContext.BaseDirectory); + + while (dir is not null && !Directory.Exists(Path.Combine(dir.FullName, "Lite", "Services"))) + { + dir = dir.Parent; + } + + Assert.NotNull(dir); + + var path = Path.Combine(dir!.FullName, "Lite", "Services", fileName); + Assert.True(File.Exists(path), $"could not locate {fileName}"); + + return File.ReadAllText(path); + } +} diff --git a/Lite/Analysis/AnalysisAbandon.cs b/Lite/Analysis/AnalysisAbandon.cs new file mode 100644 index 000000000..51b6c3af7 --- /dev/null +++ b/Lite/Analysis/AnalysisAbandon.cs @@ -0,0 +1,41 @@ +using System; +using System.Threading; + +namespace PerformanceMonitorLite.Analysis; + +/// +/// Classifies whether an analysis-pass failure is the residue of an abandonment we asked for, so the +/// pass's many catch sites can tell unfinished because we called it off from unfinished +/// because something broke (#2443). Lite's counterpart to Darling's +/// AnalysisShutdown.IsExpectedAbandon, and deliberately narrower than it. +/// +/// Both halves of the predicate are load-bearing, and the TYPE half especially so. Since #2419 +/// the pass token fires on an ordinary TIMEOUT as well as at shutdown, so it is signalled during +/// perfectly normal running — a filter that asked only "has the token fired?" would relabel any +/// genuine fault landing after the budget elapsed as an abandonment and swallow the one line of +/// evidence it left. That is the same defect #2419's first review round caught, one layer further +/// in. +/// +/// The shape list is one entry long because that is what was MEASURED, not what was assumed. +/// Darling's classifier also has to name a disposed data source and the 57P0x trio, because its +/// store is a separate postmaster that can go away underneath an in-flight read; Lite's is an +/// embedded file in this process. Against DuckDB.NET 1.5.5 a pre-cancelled token throws +/// without touching the database, and a token fired +/// mid-query reaches duckdb_interrupt through DuckDBCommand.Cancel() and surfaces as +/// too — not as a DuckDBException anyone would have +/// to pattern-match. A 2,790 ms query cancelled at 2,000 ms aborted at 2,004 ms, and both the +/// interrupted connection and a fresh one were fully usable afterwards. So naming the one shape +/// loses nothing the cancellation actually produces, and widening it would cost the fault +/// evidence. +/// +public static class AnalysisAbandon +{ + /// + /// True when this failure should be ABANDONED quietly rather than treated as a fault: the pass's + /// token has fired AND the exception is the shape an abandonment produces. Catch sites on the + /// pass use this in a when filter so the residue PROPAGATES — unwinding to the one line + /// AnalysisService logs for it — instead of being swallowed per collector. + /// + public static bool IsExpected(Exception ex, CancellationToken passToken) => + passToken.IsCancellationRequested && ex is OperationCanceledException; +} diff --git a/Lite/Analysis/AnalysisService.cs b/Lite/Analysis/AnalysisService.cs index 4cf978b21..f5748d2c0 100644 --- a/Lite/Analysis/AnalysisService.cs +++ b/Lite/Analysis/AnalysisService.cs @@ -1,6 +1,7 @@ using System; using System.Collections.Generic; using System.Linq; +using System.Threading; using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Analysis; @@ -76,9 +77,26 @@ public AnalysisService( /// Runs the full analysis pipeline for a server. /// Default time range is the last 4 hours. /// - public async Task> AnalyzeAsync(int serverId, string serverName, int hoursBack = 4) + /// #2412: abandons the pass at the scheduler's per-server + /// budget (and at app shutdown). Optional so the on-demand callers — the Recommendations tab + /// and the MCP tool, neither of which has a budget to enforce — keep the prior behavior. The + /// argument order matches the Darling twin's AnalyzeAsync so the two stay transplantable. + /// #2506: moves the END of the window off "now" while + /// stays its LENGTH, so an incident can be analyzed where it happened. + /// Null — every caller but the anchored MCP tool — is the pre-#2506 behaviour exactly. Anchoring + /// reaches the whole pipeline through the context, including the anomaly detector's hour-of-day × + /// day-of-week baseline, which is keyed off the window rather than off the clock. An anchored pass + /// does not persist — see . + [System.Diagnostics.CodeAnalysis.SuppressMessage("Design", "CA1068:CancellationToken parameters must come last", + Justification = "The token is at position 4 and the scheduler passes it POSITIONALLY. Moving asOfUtc ahead " + + "of it to satisfy the rule would silently rebind that call site's arguments — a compiling " + + "change of meaning on the one caller that matters. Appending is the only edit that cannot " + + "do that, and the Darling twin keeps the same order so the two stay transplantable.")] + public async Task> AnalyzeAsync( + int serverId, string serverName, int hoursBack = 4, CancellationToken cancellationToken = default, + DateTime? asOfUtc = null) { - var timeRangeEnd = DateTime.UtcNow; + var timeRangeEnd = asOfUtc ?? DateTime.UtcNow; var timeRangeStart = timeRangeEnd.AddHours(-hoursBack); var context = new AnalysisContext @@ -86,7 +104,9 @@ public async Task> AnalyzeAsync(int serverId, string serve ServerId = serverId, ServerName = serverName, TimeRangeStart = timeRangeStart, - TimeRangeEnd = timeRangeEnd + TimeRangeEnd = timeRangeEnd, + AsOfUtc = asOfUtc, + CancellationToken = cancellationToken }; return await AnalyzeAsync(context); @@ -105,9 +125,20 @@ public async Task> AnalyzeAsync(AnalysisContext context) try { + /* #2412: a checkpoint ahead of every store-touching stage, so the rule is simply that + no store read STARTS after the budget has gone. One check at the top of the method + would not deliver that — each stage below is a many-query phase, and a cancelled + pass would run out whichever one it was already inside. This first checkpoint earns + its place even though the read below is a single scalar: an already-cancelled + context arrives here whenever the budget is very short or the pass queued behind a + wedged server, and it should not buy a round-trip. The post-enrichment tail (action + build + insert) carries no check on purpose — by then the expensive work is paid + for and finishing is what preserves it. */ + context.CancellationToken.ThrowIfCancellationRequested(); + // 0. Check minimum data span — total history, not the analysis window. // A server with 100h of total history can be analyzed over a 4h window. - var dataSpanHours = await GetTotalDataSpanHoursAsync(context.ServerId); + var dataSpanHours = await GetTotalDataSpanHoursAsync(context.ServerId, context.CancellationToken); if (dataSpanHours < MinimumDataHours) { var needed = MinimumDataHours >= 24 @@ -128,6 +159,8 @@ public async Task> AnalyzeAsync(AnalysisContext context) return []; } + context.CancellationToken.ThrowIfCancellationRequested(); + // 1. Collect facts from DuckDB var facts = await _collector.CollectFactsAsync(context); @@ -137,6 +170,8 @@ public async Task> AnalyzeAsync(AnalysisContext context) return []; } + context.CancellationToken.ThrowIfCancellationRequested(); + // 1.5. Detect anomalies (compare analysis window against baseline) var anomalies = await _anomalyDetector.DetectAnomaliesAsync(context); facts.AddRange(anomalies); @@ -167,11 +202,15 @@ public async Task> AnalyzeAsync(AnalysisContext context) // dropped, only the incident tag is reconciled. AnomalyIncidentReconciler.Reconcile(stories); + context.CancellationToken.ThrowIfCancellationRequested(); + // 4. Mute-filter the stories into the surviving findings WITHOUT inserting yet (the // Darling twin's D2/P2 reorder) — enrichment + action-build happen on the survivors // first so the BUILT RemediationAction is persisted on each row. var findings = await _findingStore.FilterMutedFindingsAsync(stories, context); + context.CancellationToken.ThrowIfCancellationRequested(); + // 5. Enrich the survivors with drill-down data (ephemeral except through the built action). await _drillDown.EnrichFindingsAsync(findings, context); @@ -193,26 +232,61 @@ public async Task> AnalyzeAsync(AnalysisContext context) ?? FactRemediation.BuildMissingIndexAction(finding); // missing-index CREATE — copy-paste only } - // 7. Insert the survivors in one batched pass, persisting remediation_action_json. - await _findingStore.InsertFindingsAsync(findings, context); + // 7. Insert the survivors in one batched pass, persisting remediation_action_json — + // UNLESS the window was anchored at a past instant (#2506), in which case the pass is + // exploratory and writes nothing. The findings are still built, enriched and returned in + // full; only the row is withheld, because the row would claim to be a current + // observation. AnalysisContext.PersistFindings carries the whole argument. + if (context.PersistFindings) + { + await _findingStore.InsertFindingsAsync(findings, context); + } LastAnalysisTime = DateTime.UtcNow; - // 8. Notify listeners - AnalysisCompleted?.Invoke(this, new AnalysisCompletedEventArgs + // 8. Notify listeners. Gated with the insert for the same reason and not a weaker one: + // this event is how findings reach notification, and an alert about last Tuesday + // delivered today is the persistence problem with a shorter fuse. + if (context.PersistFindings) { - ServerId = context.ServerId, - ServerName = context.ServerName, - Findings = findings, - AnalysisTime = LastAnalysisTime.Value - }); + AnalysisCompleted?.Invoke(this, new AnalysisCompletedEventArgs + { + ServerId = context.ServerId, + ServerName = context.ServerName, + Findings = findings, + AnalysisTime = LastAnalysisTime.Value + }); + } AppLogger.Info("AnalysisService", $"Analysis complete for {context.ServerName}: {findings.Count} finding(s), " + - $"highest severity {(findings.Count > 0 ? findings.Max(f => f.Severity) : 0):F2}"); + $"highest severity {(findings.Count > 0 ? findings.Max(f => f.Severity) : 0):F2}" + + (context.PersistFindings ? string.Empty : " (anchored window — exploratory, not persisted)")); return findings; } + catch (Exception ex) when (AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* #2443: the predicate moved to AnalysisAbandon.IsExpected, unchanged — signalled token + AND OperationCanceledException — because the store layer now needs the same judgement + at forty-odd catch sites, and two copies of it would drift. */ + /* #2412: the pass was abandoned because it outlived its budget (or the app is + stopping), which is not a fault and must not read as one. Whatever this pass would + have written is gone; the next scheduled pass recomputes it from the store. + + Both halves of the filter are load-bearing, and the TYPE half especially so here. + This token fires on TIMEOUT as well as at shutdown, so it is signalled during + ordinary running — a blanket `Exception` filter would relabel any genuine fault + that happened to land after the budget elapsed as abandonment and drop it to Info, + burying the one line of evidence it left. OperationCanceledException is what the + checkpoints throw, and what DuckDB's own duckdb_interrupt surfaces, so nothing the + cancellation actually produces is lost by naming it. This is the Darling twin's + AnalysisShutdown.IsExpectedAbandon discipline — signalled token AND a shape the + cancellation really produces — narrowed to the one shape that arises here. */ + AppLogger.Info("AnalysisService", + $"Analysis abandoned for {context.ServerName} — this pass's findings are lost by design; the next pass recomputes them ({ex.Message})"); + return []; + } catch (Exception ex) { AppLogger.Error("AnalysisService", $"Analysis failed for {context.ServerName}: {ex.Message}"); @@ -227,10 +301,15 @@ public async Task> AnalyzeAsync(AnalysisContext context) /// /// Runs the collect + score pipeline without graph traversal. /// Returns raw scored facts with amplifier details for direct inspection. + /// + /// #2506: anchors the END of the window; null is "now", which is + /// every caller but the anchored MCP tool. Nothing here persists, so the anchor carries no + /// write-side question — this is a read that happens to score what it read. /// - public async Task> CollectAndScoreFactsAsync(int serverId, string serverName, int hoursBack = 4) + public async Task> CollectAndScoreFactsAsync( + int serverId, string serverName, int hoursBack = 4, DateTime? asOfUtc = null) { - var timeRangeEnd = DateTime.UtcNow; + var timeRangeEnd = asOfUtc ?? DateTime.UtcNow; var timeRangeStart = timeRangeEnd.AddHours(-hoursBack); var context = new AnalysisContext @@ -238,7 +317,8 @@ public async Task> CollectAndScoreFactsAsync(int serverId, string ser ServerId = serverId, ServerName = serverName, TimeRangeStart = timeRangeStart, - TimeRangeEnd = timeRangeEnd + TimeRangeEnd = timeRangeEnd, + AsOfUtc = asOfUtc }; try @@ -308,10 +388,16 @@ public async Task> GetLatestFindingsAsync(int serverId) /// Gets recent findings for a server within the given time range. The MCP findings read /// passes so its occurrence stats cover /// the whole window; the store's default 100 stays for everyone else. + /// + /// #2506: anchors the window's END. This one is a pure read of + /// rows the SCHEDULED passes already wrote, so anchoring it asks "what did analysis say about this + /// server at the time" — the only way to see findings the retention sweep has not yet reached but + /// the default 24-hour window has scrolled past. /// - public async Task> GetRecentFindingsAsync(int serverId, int hoursBack = 24, int limit = 100) + public async Task> GetRecentFindingsAsync( + int serverId, int hoursBack = 24, int limit = 100, DateTime? asOfUtc = null) { - return await _findingStore.GetRecentFindingsAsync(serverId, hoursBack, limit); + return await _findingStore.GetRecentFindingsAsync(serverId, hoursBack, limit, asOfUtc); } /// @@ -339,13 +425,13 @@ public async Task CleanupAsync(int retentionDays = 30) /// /* Internal for AnalysisDataSpanTests (#1809): the span must survive an archive/reset, which is only observable with a real DuckDB + parquet fixture. */ - internal async Task GetTotalDataSpanHoursAsync(int serverId) + internal async Task GetTotalDataSpanHoursAsync(int serverId, CancellationToken cancellationToken = default) { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(cancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(cancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -355,14 +441,18 @@ FROM v_wait_stats cmd.Parameters.Add(new DuckDBParameter { Value = serverId }); - var result = await cmd.ExecuteScalarAsync(); + var result = await cmd.ExecuteScalarAsync(cancellationToken); if (result == null || result is DBNull) return 0; return Convert.ToDouble(result); } - catch + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, cancellationToken)) { + /* A probe failure reads as "no data yet" — EXCEPT an abandonment, which must not be + allowed to masquerade as a 0-hour history (#2443). That would turn a cancelled pass + into an insufficient-data SKIP, which is a different and far calmer-looking answer + than the one the caller is about to log. */ return 0; } } diff --git a/Lite/Analysis/AnomalyDetector.cs b/Lite/Analysis/AnomalyDetector.cs index 2fbf797af..30e0ce0a1 100644 --- a/Lite/Analysis/AnomalyDetector.cs +++ b/Lite/Analysis/AnomalyDetector.cs @@ -1,6 +1,7 @@ using System; using System.Collections.Generic; using System.Linq; +using System.Threading; using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Analysis; @@ -78,8 +79,15 @@ public async Task> DetectAnomaliesAsync(AnalysisContext context) { var anomalies = new List(); - // Check if baseline period has any data at all — if not, skip all anomaly detection. - if (!await HasBaselineDataAsync(context.ServerId)) + /* Check if baseline period has any data at all — if not, skip all anomaly detection. + + #2506: the gate's 30 days are measured back from the WINDOW's end, not from the clock. Every + other bound in this class already comes off context.TimeRangeStart/End, and the baseline this + gate is guarding is computed at context.TimeRangeStart too — so asking "was anything collected + in the 30 days before now" while the baseline reads the 30 days before an anchored window was + the one place the two could disagree. Identical for an unanchored pass, whose TimeRangeEnd IS + now. */ + if (!await HasBaselineDataAsync(context.ServerId, context.TimeRangeEnd, context.CancellationToken)) return anomalies; // Existing detection methods (upgraded to time-bucketed baselines) @@ -108,9 +116,9 @@ private async Task DetectObjectStatsAnomalies(AnalysisContext context, List - private async Task HasBaselineDataAsync(int serverId) + private async Task HasBaselineDataAsync(int serverId, DateTime windowEnd, CancellationToken cancellationToken) { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(cancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(cancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -243,12 +251,18 @@ private async Task HasBaselineDataAsync(int serverId) + (SELECT COUNT(*) FROM v_cpu_utilization_stats WHERE server_id = $1 AND collection_time >= $2)"; cmd.Parameters.Add(new DuckDBParameter { Value = serverId }); - cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.UtcNow.AddDays(-30) }); + cmd.Parameters.Add(new DuckDBParameter { Value = windowEnd.AddDays(-30) }); - var count = Convert.ToInt64(await cmd.ExecuteScalarAsync() ?? 0); + var count = Convert.ToInt64(await cmd.ExecuteScalarAsync(cancellationToken) ?? 0); return count > 0; } - catch { return false; } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, cancellationToken)) + { + /* No baseline data reads as "skip anomaly detection", which is right — but an + abandonment is NOT that answer, and swallowing it here would let the pass go on + through nine detectors under a token that had already fired (#2443). */ + return false; + } } /// @@ -259,16 +273,16 @@ private async Task DetectCpuAnomalies(AnalysisContext context, List anomal try { var baseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.Cpu, context.TimeRangeStart); + context.ServerId, MetricNames.Cpu, context.TimeRangeStart, context.CancellationToken); if (baseline.SampleCount == 0) return; // No effectiveStdDev<=0 early return — an untrustworthy/zero-dispersion baseline falls // back to the absolute bar (below) rather than going silent. var effectiveStdDev = baseline.EffectiveStdDev; - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -286,8 +300,8 @@ FROM v_cpu_utilization_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var peakCpu = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var avgCpu = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -326,7 +340,7 @@ FROM v_cpu_utilization_stats Metadata = metadata }); } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"CPU anomaly detection failed: {ex.Message}"); } @@ -346,11 +360,11 @@ private async Task DetectWaitAnomalies(AnalysisContext context, List anoma try { var baseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.WaitMsPerSec, context.TimeRangeStart); + context.ServerId, MetricNames.WaitMsPerSec, context.TimeRangeStart, context.CancellationToken); - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); // Current window: all-types wait ms/sec per collection (interval via LAG, never an assumed // cadence — mirrors the WaitMsPerSec baseline), then PEAK across collections (matching the @@ -378,8 +392,8 @@ SELECT MAX(CASE WHEN interval_sec > 0 THEN total_wait_ms / interval_sec ELSE 0 E rateCmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); rateCmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var rateReader = await rateCmd.ExecuteReaderAsync(); - if (!await rateReader.ReadAsync()) return; + using var rateReader = await rateCmd.ExecuteReaderAsync(context.CancellationToken); + if (!await rateReader.ReadAsync(context.CancellationToken)) return; peakRate = rateReader.IsDBNull(0) ? 0.0 : Convert.ToDouble(rateReader.GetValue(0)); totalWaitMs = rateReader.IsDBNull(1) ? 0.0 : Convert.ToDouble(rateReader.GetValue(1)); collectionCount = rateReader.IsDBNull(2) ? 0L : Convert.ToInt64(rateReader.GetValue(2)); @@ -447,8 +461,8 @@ ORDER BY total_ms DESC contribCmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); contribCmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var contribReader = await contribCmd.ExecuteReaderAsync(); - while (await contribReader.ReadAsync()) + using var contribReader = await contribCmd.ExecuteReaderAsync(context.CancellationToken); + while (await contribReader.ReadAsync(context.CancellationToken)) { var waitType = contribReader.GetString(0); metadata[$"contrib_{waitType}"] = Convert.ToDouble(contribReader.GetValue(1)); @@ -464,7 +478,7 @@ ORDER BY total_ms DESC Metadata = metadata }); } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"Wait anomaly detection failed: {ex.Message}"); } @@ -479,13 +493,13 @@ private async Task DetectBlockingAnomalies(AnalysisContext context, List a try { var blockingBaseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.Blocking, context.TimeRangeStart); + context.ServerId, MetricNames.Blocking, context.TimeRangeStart, context.CancellationToken); var deadlockBaseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.Deadlock, context.TimeRangeStart); + context.ServerId, MetricNames.Deadlock, context.TimeRangeStart, context.CancellationToken); - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); /* current_blocking: prefer the blocked-process-report; fall back to the always-on DMV @@ -505,8 +519,8 @@ snapshot so RDS (where the BPR session is empty) still counts blocking. Mirrors cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var currentBlocking = Convert.ToInt64(reader.GetValue(0)); var currentDeadlocks = Convert.ToInt64(reader.GetValue(1)); @@ -577,7 +591,7 @@ so normalize them to per-hour before the ratio — otherwise the ratio scales wi }); } } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"Blocking anomaly detection failed: {ex.Message}"); } @@ -591,14 +605,14 @@ private async Task DetectIoAnomalies(AnalysisContext context, List anomali try { var baseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.IoLatency, context.TimeRangeStart); + context.ServerId, MetricNames.IoLatency, context.TimeRangeStart, context.CancellationToken); if (baseline.SampleCount == 0) return; var effectiveStdDev = baseline.EffectiveStdDev; - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -612,8 +626,8 @@ FROM v_file_io_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var currentReadLat = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var currentWriteLat = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -678,7 +692,7 @@ FROM v_file_io_stats }); } } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"I/O anomaly detection failed: {ex.Message}"); } @@ -692,14 +706,14 @@ private async Task DetectBatchRequestAnomalies(AnalysisContext context, List an try { var baseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.SessionCount, context.TimeRangeStart); + context.ServerId, MetricNames.SessionCount, context.TimeRangeStart, context.CancellationToken); if (baseline.SampleCount == 0) return; var effectiveStdDev = baseline.EffectiveStdDev; - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -794,8 +808,8 @@ SELECT AVG(total_connections) AS avg_connections, cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var avgConnections = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var peakConnections = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -832,7 +846,7 @@ SELECT AVG(total_connections) AS avg_connections, Metadata = metadata }); } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"Session anomaly detection failed: {ex.Message}"); } @@ -847,14 +861,14 @@ private async Task DetectQueryDurationAnomalies(AnalysisContext context, List ano try { var baseline = await _baselineProvider.GetBaselineAsync( - context.ServerId, MetricNames.Memory, context.TimeRangeStart); + context.ServerId, MetricNames.Memory, context.TimeRangeStart, context.CancellationToken); if (baseline.SampleCount == 0) return; var effectiveStdDev = baseline.EffectiveStdDev; - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -952,8 +966,8 @@ FROM v_memory_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var avgPressure = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var peakPressure = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -990,7 +1004,7 @@ FROM v_memory_stats Metadata = metadata }); } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("AnomalyDetector", $"Memory anomaly detection failed: {ex.Message}"); } diff --git a/Lite/Analysis/BaselineProvider.cs b/Lite/Analysis/BaselineProvider.cs index fe41f385d..a517765c8 100644 --- a/Lite/Analysis/BaselineProvider.cs +++ b/Lite/Analysis/BaselineProvider.cs @@ -2,6 +2,7 @@ using System.Collections.Concurrent; using System.Collections.Generic; using System.Linq; +using System.Threading; using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Analysis.Baselines; @@ -116,14 +117,14 @@ private void WarnIfRetentionUndercutsBaselineWindow(string metricName) /// Returns the most specific bucket available, collapsing as needed. /// public async Task GetBaselineAsync( - int serverId, string metricName, DateTime analysisTime) + int serverId, string metricName, DateTime analysisTime, CancellationToken cancellationToken = default) { WarnIfRetentionUndercutsBaselineWindow(metricName); var hourOfDay = analysisTime.Hour; var dayOfWeek = (int)analysisTime.DayOfWeek; // Sunday=0 - var baselines = await GetOrComputeBaselinesAsync(serverId, metricName, analysisTime); + var baselines = await GetOrComputeBaselinesAsync(serverId, metricName, analysisTime, cancellationToken); if (baselines == null || baselines.Count == 0) return BaselineBucket.Empty; @@ -142,7 +143,7 @@ public void InvalidateCache(int serverId) public void ClearCache() => _cache.Clear(); private async Task?> GetOrComputeBaselinesAsync( - int serverId, string metricName, DateTime analysisTime) + int serverId, string metricName, DateTime analysisTime, CancellationToken cancellationToken) { var cacheKey = $"{serverId}:{metricName}"; var roundedHour = new DateTime(analysisTime.Year, analysisTime.Month, analysisTime.Day, analysisTime.Hour, 0, 0); @@ -154,7 +155,7 @@ public void InvalidateCache(int serverId) return cached.Buckets; } - var buckets = await ComputeBaselinesAsync(serverId, metricName, analysisTime); + var buckets = await ComputeBaselinesAsync(serverId, metricName, analysisTime, cancellationToken); _cache[cacheKey] = new CachedBaseline { @@ -167,7 +168,7 @@ public void InvalidateCache(int serverId) } private async Task?> ComputeBaselinesAsync( - int serverId, string metricName, DateTime analysisTime) + int serverId, string metricName, DateTime analysisTime, CancellationToken cancellationToken) { var query = GetBaselineQuery(metricName); if (query == null) return null; @@ -177,9 +178,9 @@ public void InvalidateCache(int serverId) try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(cancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(cancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = query; @@ -189,13 +190,13 @@ public void InvalidateCache(int serverId) var buckets = new Dictionary<(int, int), BaselineBucket>(); - using var reader = await cmd.ExecuteReaderAsync(); + using var reader = await cmd.ExecuteReaderAsync(cancellationToken); /* #1743: the robust-scaffold metrics return eight columns (…, median_val, mad_val) and carry sentinel tier rows — every Lite arm except the two event-family ones, which keep the six-column classical shape (detected by column count) and whose buckets read Median=0/Mad=0, degrading the robust path to the classical one. */ var hasRobustColumns = reader.FieldCount >= 8; - while (await reader.ReadAsync()) + while (await reader.ReadAsync(cancellationToken)) { var hour = Convert.ToInt32(reader.GetValue(0)); var dow = Convert.ToInt32(reader.GetValue(1)); @@ -230,8 +231,12 @@ public void InvalidateCache(int serverId) return buckets; } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, cancellationToken)) { + /* #2443: a baseline that could not be computed leaves its metric with no anomaly + detection, which is worth an ERROR. An abandoned one is not that — it is the pass we + called off, and five of the seven lines #2299 was filed about came from exactly this + catch on the Darling twin. */ AppLogger.Error("BaselineProvider", $"Failed to compute baselines for {metricName}: {ex.Message}"); return null; } diff --git a/Lite/Analysis/BlockingPairRowQuery.cs b/Lite/Analysis/BlockingPairRowQuery.cs index fd13f1bb6..f5d692ce1 100644 --- a/Lite/Analysis/BlockingPairRowQuery.cs +++ b/Lite/Analysis/BlockingPairRowQuery.cs @@ -2,6 +2,7 @@ using System.Collections.Generic; using System.Data.Common; using System.Numerics; +using System.Threading; using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Analysis; @@ -101,8 +102,15 @@ public static long ToInt64(object value) /// fragments as the blocked-process-report queries, against v_dmv_blocking_snapshots, so /// maps it unchanged. Takes a command factory so the caller's connection (a LockedConnection in the viewer, /// a raw DuckDBConnection in the collectors) runs it on the read lock it already holds. + /// + /// #2443: the token is required, not defaulted. Three callers share this fetch and two of + /// them are on the analysis pass; a default would have let either keep passing nothing while the + /// signature claimed the read was abandonable. The viewer's call is the one that legitimately has + /// no pass to abandon, and it says so at its own call site rather than here. /// - internal static async Task AppendDmvSnapshotRowsAsync(Func createCommand, List rows, int serverId, DateTime start, DateTime end) + internal static async Task AppendDmvSnapshotRowsAsync( + Func createCommand, List rows, int serverId, DateTime start, DateTime end, + CancellationToken cancellationToken) { var dmv = new List(); using (var cmd = createCommand()) @@ -123,8 +131,8 @@ ORDER BY event_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = start }); cmd.Parameters.Add(new DuckDBParameter { Value = end }); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) dmv.Add(Read(reader)); } diff --git a/Lite/Analysis/DrillDownCollector.Blocking.cs b/Lite/Analysis/DrillDownCollector.Blocking.cs index b34898924..730a1fdd0 100644 --- a/Lite/Analysis/DrillDownCollector.Blocking.cs +++ b/Lite/Analysis/DrillDownCollector.Blocking.cs @@ -18,9 +18,9 @@ public partial class DrillDownCollector { private async Task CollectTopDeadlocks(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -37,8 +37,8 @@ ORDER BY collection_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { /* #1140: parse the involved objects from the graph for the dedup fingerprint + a readable Objects field. The raw graph XML is NOT surfaced (it would bloat the alert detail). */ @@ -59,9 +59,9 @@ Objects field. The raw graph XML is NOT surfaced (it would bloat the alert detai private async Task CollectTopBlockingChains(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); /* BPR + always-on DMV blocking snapshot, so the flat top-blocking list isn't empty when the @@ -98,8 +98,8 @@ ORDER BY wait_time_ms DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -126,9 +126,9 @@ ORDER BY wait_time_ms DESC /// private async Task CollectReconstructedBlockingChains(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); // SpidFilter keeps Lite's drill-down, fact collector, and viewer fetch in lockstep on the apex @@ -153,16 +153,17 @@ ORDER BY event_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var rows = new List(); - using (var reader = await cmd.ExecuteReaderAsync()) + using (var reader = await cmd.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) rows.Add(BlockingPairRowQuery.Read(reader)); } // Always-on DMV blocking snapshot fallback. Merge BEFORE the empty check so DMV-only blocking // (blocked-process-report unavailable, e.g. AWS RDS) still reconstructs. await BlockingPairRowQuery.AppendDmvSnapshotRowsAsync( - connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd); + connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd, + context.CancellationToken); if (rows.Count == 0) return; @@ -199,9 +200,9 @@ await BlockingPairRowQuery.AppendDmvSnapshotRowsAsync( private async Task CollectLockModeBreakdown(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -221,8 +222,8 @@ ORDER BY total_wait_ms DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Lite/Analysis/DrillDownCollector.Config.cs b/Lite/Analysis/DrillDownCollector.Config.cs index 20abac7fa..6fe49f3d9 100644 --- a/Lite/Analysis/DrillDownCollector.Config.cs +++ b/Lite/Analysis/DrillDownCollector.Config.cs @@ -18,9 +18,9 @@ public partial class DrillDownCollector { private async Task CollectConfigIssues(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -36,8 +36,8 @@ FROM v_database_config cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var issues = new List(); if (!reader.IsDBNull(2) && reader.GetBoolean(2)) issues.Add("auto_shrink ON"); diff --git a/Lite/Analysis/DrillDownCollector.Plans.cs b/Lite/Analysis/DrillDownCollector.Plans.cs index 23bcc5cd1..67c03d77b 100644 --- a/Lite/Analysis/DrillDownCollector.Plans.cs +++ b/Lite/Analysis/DrillDownCollector.Plans.cs @@ -37,9 +37,9 @@ private async Task CollectPlanAnalysis(AnalysisFinding finding, AnalysisContext string? planHandle = null; try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -54,16 +54,21 @@ ORDER BY collection_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); cmd.Parameters.Add(new DuckDBParameter { Value = queryHash }); - using var reader = await cmd.ExecuteReaderAsync(); - if (await reader.ReadAsync() && !reader.IsDBNull(0)) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (await reader.ReadAsync(context.CancellationToken) && !reader.IsDBNull(0)) planHandle = reader.GetString(0); } - catch { return; } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* No plan_handle for this hash — the fetch below has nothing to ask for. An abandonment + is NOT swallowed here (#2443). */ + return; + } if (string.IsNullOrEmpty(planHandle)) return; // Fetch plan XML live from SQL Server - var planXml = await _planFetcher.FetchPlanXmlAsync(context.ServerId, planHandle); + var planXml = await _planFetcher.FetchPlanXmlAsync(context.ServerId, planHandle, context.CancellationToken); if (string.IsNullOrEmpty(planXml)) return; try @@ -128,10 +133,10 @@ private async Task CollectPlanAdvisoryDetail(AnalysisFinding finding, AnalysisCo { var planXmls = new List(); - using (var readLock = _duckDb.AcquireReadLock()) + using (var readLock = _duckDb.AcquireReadLock(context.CancellationToken)) using (var connection = _duckDb.CreateConnection()) { - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -147,8 +152,8 @@ ORDER BY delta_worker_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { if (!reader.IsDBNull(0)) planXmls.Add(reader.GetString(0)); @@ -188,9 +193,10 @@ ORDER BY delta_worker_time DESC .ToList(); } } - catch + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { // Plan read/parse can fail on malformed XML — skip, the detail is best-effort. + // An abandonment is NOT swallowed here (#2443). } } } diff --git a/Lite/Analysis/DrillDownCollector.Queries.cs b/Lite/Analysis/DrillDownCollector.Queries.cs index 566655dcc..be5224a84 100644 --- a/Lite/Analysis/DrillDownCollector.Queries.cs +++ b/Lite/Analysis/DrillDownCollector.Queries.cs @@ -19,9 +19,9 @@ public partial class DrillDownCollector private async Task CollectQueriesAtSpike(AnalysisFinding finding, AnalysisContext context) { // Find the peak CPU time, then get queries active within 2 minutes of it - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); // Step 1: Find when the spike occurred using var peakCmd = connection.CreateCommand(); @@ -38,9 +38,9 @@ ORDER BY sqlserver_cpu_utilization DESC DateTime? peakTime = null; int peakCpu = 0; - using (var peakReader = await peakCmd.ExecuteReaderAsync()) + using (var peakReader = await peakCmd.ExecuteReaderAsync(context.CancellationToken)) { - if (await peakReader.ReadAsync()) + if (await peakReader.ReadAsync(context.CancellationToken)) { peakTime = peakReader.GetDateTime(0); peakCpu = peakReader.GetInt32(1); @@ -69,9 +69,9 @@ ORDER BY cpu_time_ms DESC queryCmd.Parameters.Add(new DuckDBParameter { Value = peakTime.Value.AddMinutes(2) }); var items = new List(); - using (var reader = await queryCmd.ExecuteReaderAsync()) + using (var reader = await queryCmd.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -103,9 +103,9 @@ ORDER BY cpu_time_ms DESC private async Task CollectTopCpuQueries(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -127,8 +127,8 @@ ORDER BY total_cpu_us DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -148,9 +148,9 @@ ORDER BY total_cpu_us DESC private async Task CollectTopSpillingQueries(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -170,8 +170,8 @@ ORDER BY total_spills DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -193,9 +193,9 @@ ORDER BY total_spills DESC /// private async Task CollectParameterSensitiveQueries(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -251,8 +251,8 @@ ORDER BY worker_ratio DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -281,9 +281,9 @@ ORDER BY worker_ratio DESC /// private async Task CollectRegressedQueries(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -489,8 +489,8 @@ ORDER BY regression_factor DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -527,9 +527,9 @@ private async Task CollectBadActorDetail(AnalysisFinding finding, AnalysisContex var queryHash = finding.RootFactKey.Replace("BAD_ACTOR_", ""); if (string.IsNullOrEmpty(queryHash)) return; - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -561,8 +561,8 @@ FROM v_query_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); cmd.Parameters.Add(new DuckDBParameter { Value = queryHash }); - using var reader = await cmd.ExecuteReaderAsync(); - if (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (await reader.ReadAsync(context.CancellationToken)) { finding.DrillDown!["bad_actor_query"] = new { @@ -583,9 +583,9 @@ FROM v_query_stats private async Task CollectPendingGrants(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -605,8 +605,8 @@ ORDER BY waiter_count DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Lite/Analysis/DrillDownCollector.Storage.cs b/Lite/Analysis/DrillDownCollector.Storage.cs index e76122e49..8d3eb0629 100644 --- a/Lite/Analysis/DrillDownCollector.Storage.cs +++ b/Lite/Analysis/DrillDownCollector.Storage.cs @@ -18,9 +18,9 @@ public partial class DrillDownCollector { private async Task CollectFileLatencyBreakdown(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -41,8 +41,8 @@ ORDER BY avg_read_ms DESC NULLS LAST cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { @@ -69,9 +69,9 @@ ORDER BY avg_read_ms DESC NULLS LAST /// private async Task CollectAutogrowthPercentFiles(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -93,8 +93,8 @@ ORDER BY total_size_mb DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var database = reader.IsDBNull(0) ? "" : reader.GetString(0); var fileType = reader.IsDBNull(1) ? "" : reader.GetString(1); @@ -122,9 +122,9 @@ ORDER BY total_size_mb DESC private async Task CollectTempDbBreakdown(AnalysisFinding finding, AnalysisContext context) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -140,8 +140,8 @@ ORDER BY (user_object_reserved_mb + internal_object_reserved_mb + version_store_ cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var items = new List(); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { items.Add(new { diff --git a/Lite/Analysis/DrillDownCollector.cs b/Lite/Analysis/DrillDownCollector.cs index 4c6ce318c..a59532b0d 100644 --- a/Lite/Analysis/DrillDownCollector.cs +++ b/Lite/Analysis/DrillDownCollector.cs @@ -41,6 +41,11 @@ public async Task EnrichFindingsAsync(List findings, AnalysisCo { foreach (var finding in findings) { + /* #2443: between findings is the natural abandon point — the per-finding catch below + deliberately does NOT swallow an abandonment, so this throw (and any residue from a + drill-down mid-read) unwinds the pass to the service's single Information line. */ + context.CancellationToken.ThrowIfCancellationRequested(); + try { finding.DrillDown = new Dictionary(); @@ -124,7 +129,7 @@ 0.5 display gate. (Lite is advise/copy-paste only — no Apply — but still if (finding.DrillDown.Count == 0) finding.DrillDown = null; } - catch (Exception ex) + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { AppLogger.Error("DrillDownCollector", $"Drill-down failed for {finding.StoryPath}: {ex.GetType().Name}: {ex.Message}"); diff --git a/Lite/Analysis/DuckDbFactCollector.Activity.cs b/Lite/Analysis/DuckDbFactCollector.Activity.cs index b2fd66fc6..48c0e7f20 100644 --- a/Lite/Analysis/DuckDbFactCollector.Activity.cs +++ b/Lite/Analysis/DuckDbFactCollector.Activity.cs @@ -22,9 +22,9 @@ private async Task CollectBadActorFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -61,8 +61,8 @@ ORDER BY SUM(delta_worker_time)::DOUBLE PRECISION / GREATEST(SUM(delta_execution cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var dbName = reader.IsDBNull(0) ? "" : reader.GetString(0); var queryHash = reader.IsDBNull(1) ? "" : reader.GetString(1); @@ -100,7 +100,10 @@ ORDER BY SUM(delta_worker_time)::DOUBLE PRECISION / GREATEST(SUM(delta_execution }); } } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -110,9 +113,9 @@ private async Task CollectActiveQueryFactsAsync(AnalysisContext context, List @@ -174,9 +180,9 @@ private async Task CollectRunningJobFactsAsync(AnalysisContext context, List @@ -229,9 +238,9 @@ private async Task CollectSessionFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -256,8 +265,8 @@ FROM v_session_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var totalConns = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (totalConns == 0) return; @@ -285,7 +294,10 @@ FROM v_session_stats } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Lite/Analysis/DuckDbFactCollector.Config.cs b/Lite/Analysis/DuckDbFactCollector.Config.cs index aafe3908b..ee2ce4edb 100644 --- a/Lite/Analysis/DuckDbFactCollector.Config.cs +++ b/Lite/Analysis/DuckDbFactCollector.Config.cs @@ -19,9 +19,9 @@ public partial class DuckDbFactCollector /// private async Task CollectServerConfigFactsAsync(AnalysisContext context, List facts) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); // Latest value PER configuration_name (ROW_NUMBER, not LIMIT N): server_config accumulates @@ -57,9 +57,9 @@ FROM latest double? maxMemoryMb = null; double? minMemoryMb = null; - using (var reader = await cmd.ExecuteReaderAsync()) + using (var reader = await cmd.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { var configName = reader.GetString(0); var value = Convert.ToDouble(reader.GetValue(1)); @@ -110,9 +110,9 @@ private async Task CollectServerMetadataFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -125,8 +125,8 @@ ORDER BY collection_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var edition = reader.IsDBNull(0) ? 0 : Convert.ToInt32(reader.GetValue(0)); var majorVersion = reader.IsDBNull(1) ? 0 : Convert.ToInt32(reader.GetValue(1)); @@ -136,7 +136,10 @@ ORDER BY collection_time DESC if (majorVersion > 0) facts.Add(new Fact { Source = "config", Key = "SERVER_MAJOR_VERSION", Value = majorVersion, ServerId = context.ServerId }); } - catch { /* Columns may not exist yet (pre-migration) */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Columns may not exist yet (pre-migration). An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -147,9 +150,9 @@ private async Task CollectDatabaseConfigFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -178,8 +181,8 @@ FROM v_database_config cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var dbCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (dbCount == 0) return; @@ -215,7 +218,10 @@ FROM v_database_config } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -225,9 +231,9 @@ private async Task CollectTraceFlagFactsAsync(AnalysisContext context, List(); var flagCount = 0; - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { var flag = Convert.ToInt32(reader.GetValue(0)); metadata[$"TF_{flag}"] = 1; @@ -268,7 +274,10 @@ SELECT trace_flag Metadata = metadata }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -279,9 +288,9 @@ private async Task CollectServerPropertiesFactsAsync(AnalysisContext context, Li { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -295,8 +304,8 @@ ORDER BY collection_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var cpuCount = reader.IsDBNull(0) ? 0 : Convert.ToInt32(reader.GetValue(0)); var htRatio = reader.IsDBNull(1) ? 0 : Convert.ToInt32(reader.GetValue(1)); @@ -333,7 +342,10 @@ ORDER BY collection_time DESC // emitted (noise control). FactCollectorHelpers.EmitServerHealthFacts(context, facts, edition, physicalMemMb, lpim, ifi, dumpCount); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Lite/Analysis/DuckDbFactCollector.QueryPerf.cs b/Lite/Analysis/DuckDbFactCollector.QueryPerf.cs index 2ac62f305..e616f76c3 100644 --- a/Lite/Analysis/DuckDbFactCollector.QueryPerf.cs +++ b/Lite/Analysis/DuckDbFactCollector.QueryPerf.cs @@ -20,9 +20,9 @@ private async Task CollectQueryStatsFactsAsync(AnalysisContext context, List @@ -99,9 +102,9 @@ private async Task CollectParameterSensitivityFactsAsync(AnalysisContext context { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -157,8 +160,8 @@ ORDER BY worker_ratio DESC var worstGrantRatio = 0.0; var worstSpillDivergence = 0; - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { // Rows arrive ordered by worker_ratio DESC — the first row is the worst offender. if (offenderCount == 0) @@ -192,7 +195,10 @@ ORDER BY worker_ratio DESC } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -207,9 +213,9 @@ private async Task CollectPlanRegressionFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -370,8 +376,8 @@ ORDER BY regression_factor DESC var worstLatestForced = 0; var worstForceFailures = 0L; - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { // Rows arrive ordered by regression_factor DESC — the first row is the worst offender. if (offenderCount == 0) @@ -415,7 +421,10 @@ ORDER BY regression_factor DESC } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -425,9 +434,9 @@ private async Task CollectProcedureStatsFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -447,8 +456,8 @@ FROM v_procedure_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var distinctProcs = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); var totalExecs = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -475,7 +484,10 @@ FROM v_procedure_stats } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -491,10 +503,10 @@ private async Task CollectPlanAdvisoryFactsAsync(AnalysisContext context, List(); - using (var readLock = _duckDb.AcquireReadLock()) + using (var readLock = _duckDb.AcquireReadLock(context.CancellationToken)) using (var connection = _duckDb.CreateConnection()) { - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var command = connection.CreateCommand(); command.CommandText = @" @@ -510,8 +522,8 @@ ORDER BY delta_worker_time DESC command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await command.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { if (!reader.IsDBNull(0)) planXmls.Add(reader.GetString(0)); @@ -556,9 +568,10 @@ ORDER BY delta_worker_time DESC }); } } - catch + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) { // query_stats / plan parse may be unavailable — skip, the advisory is best-effort. + // An abandonment is NOT swallowed here (#2443). } } diff --git a/Lite/Analysis/DuckDbFactCollector.Resources.cs b/Lite/Analysis/DuckDbFactCollector.Resources.cs index 7bd98fada..7060f81bd 100644 --- a/Lite/Analysis/DuckDbFactCollector.Resources.cs +++ b/Lite/Analysis/DuckDbFactCollector.Resources.cs @@ -20,9 +20,9 @@ private async Task CollectMemoryFactsAsync(AnalysisContext context, List f { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -36,8 +36,8 @@ ORDER BY collection_time DESC cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var totalPhysical = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var bufferPool = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -50,7 +50,10 @@ ORDER BY collection_time DESC if (targetMemory > 0) facts.Add(new Fact { Source = "memory", Key = "MEMORY_TARGET_MB", Value = targetMemory, ServerId = context.ServerId }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -70,9 +73,9 @@ private async Task CollectRunnableTaskFactsAsync(AnalysisContext context, List @@ -118,9 +124,9 @@ private async Task CollectMemoryGrantFactsAsync(AnalysisContext context, List @@ -177,9 +186,9 @@ private async Task CollectMemoryClerkFactsAsync(AnalysisContext context, List(); var totalMb = 0.0; var clerkCount = 0; - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) { var clerkType = reader.GetString(0); var memoryMb = Convert.ToDouble(reader.GetValue(1)); @@ -226,7 +235,10 @@ ORDER BY memory_mb DESC Metadata = metadata }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -237,9 +249,9 @@ private async Task CollectCpuUtilizationFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -258,8 +270,8 @@ FROM v_cpu_utilization_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var avgSqlCpu = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var maxSqlCpu = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -303,7 +315,10 @@ FROM v_cpu_utilization_stats }); } } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -314,9 +329,9 @@ private async Task CollectPerfmonFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -336,8 +351,8 @@ AND counter_name IN ('Batch Requests/sec', 'SQL Compilations/sec', 'SQL Re-Com cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var counterName = reader.GetString(0); var cntrValue = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -370,7 +385,10 @@ AND counter_name IN ('Batch Requests/sec', 'SQL Compilations/sec', 'SQL Re-Com }); } } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -388,9 +406,9 @@ private async Task CollectPlanCacheFactsAsync(AnalysisContext context, List @@ -463,9 +484,9 @@ private async Task CollectMemoryPressureEventFactsAsync(AnalysisContext context, { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -482,8 +503,8 @@ FROM v_memory_pressure_events cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var pressureEventCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); var maxProcess = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -508,7 +529,10 @@ FROM v_memory_pressure_events } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Lite/Analysis/DuckDbFactCollector.Storage.cs b/Lite/Analysis/DuckDbFactCollector.Storage.cs index f8411b136..7b01d09e3 100644 --- a/Lite/Analysis/DuckDbFactCollector.Storage.cs +++ b/Lite/Analysis/DuckDbFactCollector.Storage.cs @@ -20,9 +20,9 @@ private async Task CollectDatabaseSizeFactAsync(AnalysisContext context, List 0) facts.Add(new Fact { Source = "config", Key = "DATABASE_TOTAL_SIZE_MB", Value = totalSize, ServerId = context.ServerId }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -59,9 +62,9 @@ private async Task CollectIoLatencyFactsAsync(AnalysisContext context, List @@ -135,9 +141,9 @@ private async Task CollectTempDbFactsAsync(AnalysisContext context, List f { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -147,7 +153,12 @@ private async Task CollectTempDbFactsAsync(AnalysisContext context, List f MAX(internal_object_reserved_mb) AS max_internal_object_mb, MAX(version_store_reserved_mb) AS max_version_store_mb, MIN(unallocated_mb) AS min_unallocated_mb, - AVG(total_reserved_mb) AS avg_total_reserved_mb + AVG(total_reserved_mb) AS avg_total_reserved_mb, + /* #2515: the growth ceiling, and MIN is the right aggregate for it twice over. -1 (at least one + unlimited data file) sorts below every real cap, so MIN finds the unlimited case anywhere in the + window; and where the cap was raised mid-window MIN keeps the more conservative of the two. NULL + rows — collected before the v56 migration — are skipped rather than dragging the answer down. */ + MIN(max_size_mb) AS min_max_size_mb FROM v_tempdb_stats WHERE server_id = $1 AND collection_time >= $2 @@ -157,8 +168,8 @@ FROM v_tempdb_stats cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var maxReserved = reader.IsDBNull(0) ? 0.0 : Convert.ToDouble(reader.GetValue(0)); var maxUserObj = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -166,11 +177,18 @@ FROM v_tempdb_stats var maxVersionStore = reader.IsDBNull(3) ? 0.0 : Convert.ToDouble(reader.GetValue(3)); var minUnallocated = reader.IsDBNull(4) ? 0.0 : Convert.ToDouble(reader.GetValue(4)); var avgReserved = reader.IsDBNull(5) ? 0.0 : Convert.ToDouble(reader.GetValue(5)); + var maxSizeMb = reader.IsDBNull(6) ? 0.0 : Convert.ToDouble(reader.GetValue(6)); if (maxReserved <= 0) return; - // TempDB usage as fraction of total space (reserved + unallocated) - var totalSpace = maxReserved + minUnallocated; + /* #2515: against the CEILING where tempdb has one, against the current allocation where it does + not (-1 unlimited, or 0 for a window collected before the ceiling was captured). reserved + + unallocated is the files AS ALLOCATED, so on its own this fraction measures distance to the + next autogrow — which reads as 96% full on an Azure SQL Database target holding one temp + table. TempDbSpaceInfo.CapacityMb is the alert's twin of this, and the two must agree or + analyze_server and the pager describe the same server differently. */ + var allocatedMb = maxReserved + minUnallocated; + var totalSpace = maxSizeMb > 0 ? Math.Max(maxSizeMb, allocatedMb) : allocatedMb; var usageFraction = totalSpace > 0 ? maxReserved / totalSpace : 0; facts.Add(new Fact @@ -187,11 +205,17 @@ FROM v_tempdb_stats ["max_internal_object_mb"] = maxInternalObj, ["max_version_store_mb"] = maxVersionStore, ["min_unallocated_mb"] = minUnallocated, + /* Added, never redefined — the existing keys are a consumer surface (FactAdvice reads + them by name). -1 unlimited, 0 not measured, positive = the ROWS ceiling in MB. */ + ["max_size_mb"] = maxSizeMb, ["usage_fraction"] = usageFraction } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -206,9 +230,9 @@ private async Task CollectFileAutogrowthFactsAsync(AnalysisContext context, List { try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var cmd = connection.CreateCommand(); cmd.CommandText = @" @@ -229,8 +253,8 @@ FROM latest cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var fileCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (fileCount == 0) return; @@ -250,7 +274,10 @@ FROM latest } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } /// @@ -260,9 +287,9 @@ private async Task CollectDiskSpaceFactsAsync(AnalysisContext context, List 0 cmd.Parameters.Add(new DuckDBParameter { Value = context.ServerId }); cmd.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await cmd.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await cmd.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var minFreePct = reader.IsDBNull(0) ? 1.0 : Convert.ToDouble(reader.GetValue(0)); var minFreeMb = reader.IsDBNull(1) ? 0.0 : Convert.ToDouble(reader.GetValue(1)); @@ -312,6 +339,9 @@ AND volume_total_mb > 0 } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Lite/Analysis/DuckDbFactCollector.Waits.cs b/Lite/Analysis/DuckDbFactCollector.Waits.cs index caa2f8c0a..16b577da0 100644 --- a/Lite/Analysis/DuckDbFactCollector.Waits.cs +++ b/Lite/Analysis/DuckDbFactCollector.Waits.cs @@ -18,9 +18,9 @@ public partial class DuckDbFactCollector /// private async Task CollectWaitStatsFactsAsync(AnalysisContext context, List facts) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var command = connection.CreateCommand(); command.CommandText = @" @@ -41,8 +41,8 @@ GROUP BY wait_type command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await command.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + while (await reader.ReadAsync(context.CancellationToken)) { var waitType = reader.GetString(0); var waitingTasks = reader.IsDBNull(1) ? 0L : ToInt64(reader.GetValue(1)); @@ -80,9 +80,9 @@ GROUP BY wait_type /// private async Task CollectBlockingFactsAsync(AnalysisContext context, List facts) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var command = connection.CreateCommand(); command.CommandText = @" @@ -101,8 +101,8 @@ FROM v_blocked_process_reports command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await command.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var eventCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (eventCount <= 0) return; @@ -141,9 +141,9 @@ FROM v_blocked_process_reports /// private async Task CollectDeadlockFactsAsync(AnalysisContext context, List facts) { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var command = connection.CreateCommand(); command.CommandText = @" @@ -157,8 +157,8 @@ FROM v_deadlocks command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeStart }); command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); - using var reader = await command.ExecuteReaderAsync(); - if (!await reader.ReadAsync()) return; + using var reader = await command.ExecuteReaderAsync(context.CancellationToken); + if (!await reader.ReadAsync(context.CancellationToken)) return; var deadlockCount = reader.IsDBNull(0) ? 0L : ToInt64(reader.GetValue(0)); if (deadlockCount <= 0) return; @@ -194,9 +194,9 @@ private async Task CollectBlockingChainFactsAsync(AnalysisContext context, List< try { - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); using var command = connection.CreateCommand(); // SpidFilter keeps this in lockstep with the drill-down + viewer fetch on the apex @@ -222,16 +222,17 @@ ORDER BY event_time DESC command.Parameters.Add(new DuckDBParameter { Value = context.TimeRangeEnd }); var rows = new List(); - using (var reader = await command.ExecuteReaderAsync()) + using (var reader = await command.ExecuteReaderAsync(context.CancellationToken)) { - while (await reader.ReadAsync()) + while (await reader.ReadAsync(context.CancellationToken)) rows.Add(BlockingPairRowQuery.Read(reader)); } // Always-on DMV blocking snapshot fallback (works when the blocked-process-report XE is empty, // e.g. AWS RDS). Merge BEFORE the empty check so DMV-only blocking still produces facts. await BlockingPairRowQuery.AppendDmvSnapshotRowsAsync( - connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd); + connection.CreateCommand, rows, context.ServerId, context.TimeRangeStart, context.TimeRangeEnd, + context.CancellationToken); if (rows.Count == 0) return; @@ -262,7 +263,10 @@ await BlockingPairRowQuery.AppendDmvSnapshotRowsAsync( } }); } - catch { /* Table may not exist or have no data */ } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* Table may not exist or have no data. An abandonment is NOT swallowed here (#2443). */ + } } } diff --git a/Lite/Analysis/DuckDbFactCollector.cs b/Lite/Analysis/DuckDbFactCollector.cs index dad6bad12..6103c9970 100644 --- a/Lite/Analysis/DuckDbFactCollector.cs +++ b/Lite/Analysis/DuckDbFactCollector.cs @@ -26,39 +26,52 @@ public async Task> CollectFactsAsync(AnalysisContext context) { var facts = new List(); - await CollectWaitStatsFactsAsync(context, facts); + /* #2412: one cancellation checkpoint per collector, not one for the phase. This is the + longest stage of an analysis pass — thirty-one collectors, each opening the store and + running its own reads — so a check only at the phase boundary would let a pass that has + already blown its budget run every remaining collector before noticing. Routing the + calls through a local step is what buys the per-collector check without writing it out + thirty-one times. The two grouping helpers below rewrite facts already in hand rather + than reading the store, so they stay inline where their ordering matters. */ + async Task RunCollectorAsync(Func, Task> collect) + { + context.CancellationToken.ThrowIfCancellationRequested(); + await collect(context, facts); + } + + await RunCollectorAsync(CollectWaitStatsFactsAsync); FactCollectorHelpers.GroupGeneralLockWaits(facts, context); FactCollectorHelpers.GroupParallelismWaits(facts, context); - await CollectBlockingFactsAsync(context, facts); - await CollectBlockingChainFactsAsync(context, facts); - await CollectDeadlockFactsAsync(context, facts); - await CollectServerConfigFactsAsync(context, facts); - await CollectMemoryFactsAsync(context, facts); - await CollectDatabaseSizeFactAsync(context, facts); - await CollectServerMetadataFactsAsync(context, facts); - await CollectCpuUtilizationFactsAsync(context, facts); - await CollectRunnableTaskFactsAsync(context, facts); - await CollectIoLatencyFactsAsync(context, facts); - await CollectTempDbFactsAsync(context, facts); - await CollectMemoryGrantFactsAsync(context, facts); - await CollectQueryStatsFactsAsync(context, facts); - await CollectParameterSensitivityFactsAsync(context, facts); - await CollectPlanRegressionFactsAsync(context, facts); - await CollectBadActorFactsAsync(context, facts); - await CollectPerfmonFactsAsync(context, facts); - await CollectMemoryClerkFactsAsync(context, facts); - await CollectPlanCacheFactsAsync(context, facts); - await CollectMemoryPressureEventFactsAsync(context, facts); - await CollectDatabaseConfigFactsAsync(context, facts); - await CollectFileAutogrowthFactsAsync(context, facts); - await CollectProcedureStatsFactsAsync(context, facts); - await CollectActiveQueryFactsAsync(context, facts); - await CollectRunningJobFactsAsync(context, facts); - await CollectSessionFactsAsync(context, facts); - await CollectTraceFlagFactsAsync(context, facts); - await CollectServerPropertiesFactsAsync(context, facts); - await CollectDiskSpaceFactsAsync(context, facts); - await CollectPlanAdvisoryFactsAsync(context, facts); + await RunCollectorAsync(CollectBlockingFactsAsync); + await RunCollectorAsync(CollectBlockingChainFactsAsync); + await RunCollectorAsync(CollectDeadlockFactsAsync); + await RunCollectorAsync(CollectServerConfigFactsAsync); + await RunCollectorAsync(CollectMemoryFactsAsync); + await RunCollectorAsync(CollectDatabaseSizeFactAsync); + await RunCollectorAsync(CollectServerMetadataFactsAsync); + await RunCollectorAsync(CollectCpuUtilizationFactsAsync); + await RunCollectorAsync(CollectRunnableTaskFactsAsync); + await RunCollectorAsync(CollectIoLatencyFactsAsync); + await RunCollectorAsync(CollectTempDbFactsAsync); + await RunCollectorAsync(CollectMemoryGrantFactsAsync); + await RunCollectorAsync(CollectQueryStatsFactsAsync); + await RunCollectorAsync(CollectParameterSensitivityFactsAsync); + await RunCollectorAsync(CollectPlanRegressionFactsAsync); + await RunCollectorAsync(CollectBadActorFactsAsync); + await RunCollectorAsync(CollectPerfmonFactsAsync); + await RunCollectorAsync(CollectMemoryClerkFactsAsync); + await RunCollectorAsync(CollectPlanCacheFactsAsync); + await RunCollectorAsync(CollectMemoryPressureEventFactsAsync); + await RunCollectorAsync(CollectDatabaseConfigFactsAsync); + await RunCollectorAsync(CollectFileAutogrowthFactsAsync); + await RunCollectorAsync(CollectProcedureStatsFactsAsync); + await RunCollectorAsync(CollectActiveQueryFactsAsync); + await RunCollectorAsync(CollectRunningJobFactsAsync); + await RunCollectorAsync(CollectSessionFactsAsync); + await RunCollectorAsync(CollectTraceFlagFactsAsync); + await RunCollectorAsync(CollectServerPropertiesFactsAsync); + await RunCollectorAsync(CollectDiskSpaceFactsAsync); + await RunCollectorAsync(CollectPlanAdvisoryFactsAsync); return facts; } diff --git a/Lite/Analysis/FindingStore.cs b/Lite/Analysis/FindingStore.cs index e8e089aeb..e06e86d27 100644 --- a/Lite/Analysis/FindingStore.cs +++ b/Lite/Analysis/FindingStore.cs @@ -1,10 +1,13 @@ using System; using System.Collections.Generic; +using System.Threading; using System.Threading.Tasks; using DuckDB.NET.Data; using PerformanceMonitor.Analysis; +using PerformanceMonitor.Collectors; using PerformanceMonitor.Notifications; using PerformanceMonitorLite.Database; +using PerformanceMonitorLite.Services; namespace PerformanceMonitorLite.Analysis; @@ -22,18 +25,88 @@ namespace PerformanceMonitorLite.Analysis; /// Recommendations reader can render the copy-paste command byte-identically to the Darling viewer. /// remains as a single-pass convenience wrapper (no action attached). /// +/// +/// +/// The read lock around the WRITES is deliberate (#2455). It is the first thing in this file +/// that looks like a bug, so it is answered here rather than left to cost every reader the same five +/// minutes. INSERTs, INSERTs and +/// DELETEs, and all three hold +/// DuckDbInitializer.AcquireReadLock. That lock is not a table lock and does not claim "I am +/// reading": it coordinates everyone against MAINTENANCE — CHECKPOINT, archive DELETEs, compaction — +/// which take the exclusive write lock and reorganize the file underneath whatever is running. A held +/// read lock blocks EnterWriteLock, so holding one IS how a write says "not while I am in +/// flight", and that is the only exclusion these three need. Concurrency between writers is DuckDB's +/// own job and it does it: two connections writing to analysis_findings at the same time both +/// succeed, and the retention DELETE overlaps an insert batch cleanly on disjoint rows — both +/// measured on DuckDB.NET 1.5.5 rather than assumed, and measured with each batch wrapped in a +/// transaction, which is the longer-held and therefore stronger case (and the shape #2448 gives +/// them). +/// +/// +/// +/// Taking the write lock instead would be a real cost for a benefit nobody has shown a need for: it +/// would serialize every finding batch against every UI read for the length of the batch, and #2443 +/// had just made the read-lock WAIT abandonable so the analysis pass could yield to a long archival. +/// Turning the pass into the thing archival and the UI queue behind reverses that. What a SHARED lock +/// genuinely cannot protect is a read-modify-write, because it admits concurrent holders — which is +/// why the ids below do not use one. +/// +/// +/// +/// A reader who checks the neighbours will find DuckDbAlertHistoryStore and +/// DuckDbMuteRuleStore taking the WRITE lock, and should not conclude that one of the two is +/// wrong. #2463 settled it: the lock is chosen by what a statement must EXCLUDE, not by whether it +/// reads or writes, and the rule with its measurements is on DuckDbInitializer.s_dbLock. The +/// short of it is that these three are on the correct side of it — +/// and APPEND new rows, which DuckDB lets run concurrently, and +/// 's retention DELETE is disjoint from them. What the write lock +/// buys is exclusion of another writer of the SAME rows, and nothing else writes these tables. +/// /// public class FindingStore { private readonly DuckDbInitializer _duckDb; - private long _nextId; + + /// + /// What 's upper bound is when the caller did NOT anchor: an + /// instant no row can reach, i.e. no bound at all. Matches the Darling twin's PgFindingStore. + /// + /// Why not "now". Two reasons, and the second is the one that would have hurt. First, + /// #2495's promise is that a caller sending only hours_back gets byte-for-byte the window it + /// always got, and this read has been half-open for its whole life. Second, analysis_time is + /// stamped by the WRITER's clock and would be filtered by the READER's; those are the same process + /// today, and the day they are not, a bounded default read would intermittently drop the newest run + /// — a findings read that "sometimes misses the analysis that just finished", with nothing in it to + /// point at a clock. An anchored read has a caller-supplied end and neither problem. + /// + private static readonly DateTime NoUpperBound = new(9999, 12, 31, 23, 59, 59); public FindingStore(DuckDbInitializer duckDb) { _duckDb = duckDb; - _nextId = DateTime.UtcNow.Ticks; } + /// + /// #2455: finding and mute ids come from the shared process-wide + /// rather than a per-instance counter. + /// + /// The _nextId++ this replaces was seeded from DateTime.UtcNow.Ticks per + /// INSTANCE and unprotected twice over. The increment is a read-modify-write under a lock that + /// admits concurrent holders, which is the half that was filed — but Interlocked.Increment + /// on the field would not have fixed the worse half: Lite constructs TWO of these, one in + /// AnalysisService and one in RecommendationsTab, and two instances built inside the + /// same timer tick start from the SAME value and then hand out the same ids independently, with no + /// shared state to make atomic. finding_id and mute_id are both PRIMARY KEY in + /// DuckDB, so a collision is a hard INSERT failure — and since #2448 made the batch atomic, one + /// such row costs the whole analysis rather than itself. + /// + /// The generator is process-wide and Interlocked-incremented, and it is what the + /// Darling twin's PgFindingStore has always used — whose own class doc names "the twins' + /// per-instance tick counters" as the thing it deliberately moved away from. So this closes a + /// stated parity gap rather than inventing a mechanism. + /// + private static long NextId() => CollectionIdGenerator.Next(); + /// /// Mute-filters the stories and materializes the SURVIVING findings WITHOUT inserting them /// (mirrors Darling's PgFindingStore.FilterMutedFindingsAsync — the recommendations @@ -52,11 +125,11 @@ public async Task> FilterMutedFindingsAsync( /* One read lock + one connection for the mute-filter read only. Released before the caller enriches + builds actions; InsertFindingsAsync then re-acquires for the batched insert. The lock is NoRecursion, so the helper below operates on the passed connection. */ - using var readLock = _duckDb.AcquireReadLock(); + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); - var mutedHashes = await GetMutedHashesAsync(connection, context.ServerId); + var mutedHashes = await GetMutedHashesAsync(connection, context.ServerId, context.CancellationToken); foreach (var story in stories) { @@ -69,7 +142,7 @@ public async Task> FilterMutedFindingsAsync( survivors.Add(new AnalysisFinding { - FindingId = _nextId++, + FindingId = NextId(), AnalysisTime = analysisTime, ServerId = context.ServerId, ServerName = context.ServerName, @@ -101,6 +174,40 @@ public async Task> FilterMutedFindingsAsync( /// as remediation_action_json via the shared , so a Lite /// finding's persisted action round-trips byte-identically to a Darling / Dashboard one. Returns the /// same list for caller convenience; the in-memory findings are unchanged. + /// + /// #2448: the transaction is what makes this method's promise true, and it is the whole + /// reason for the shape. A finding set is one indivisible statement about a server: every row + /// shares one analysis_time, and reads the newest + /// analysis_time. So a batch that lands four of forty rows before the store faults does + /// NOT read as truncated — it reads as a complete analysis that found four problems, and the + /// server looks HEALTHIER for the store having failed. Rolling the batch back instead leaves the + /// PREVIOUS pass as the newest complete set: stale, stamped with its own older + /// analysis_time, and incapable of misleading anyone. The rollback needs no explicit + /// call — a row that throws skips the commit and DuckDBTransaction.Dispose discards the + /// batch, measured on DuckDB.NET 1.5.5. + /// + /// Deliberately identical to the Darling twin, which reached it the same way; a divergence + /// here would be a parity bug rather than a local choice. Worth knowing when reading the loop: + /// once one statement in a DuckDB transaction fails, every later one is refused + /// ("TransactionContext Error: Current transaction is aborted") and there is no SAVEPOINT + /// to escape it — 1.5.5 does not parse the keyword — so per-row failure isolation is not + /// available inside a transaction on this engine even if it were wanted. It is not: a batch that + /// silently drops row 5 and commits the other 39 is this same defect at row granularity. + /// Commit on an already-aborted transaction also RETURNS NORMALLY while committing + /// nothing, on both this driver and Npgsql, so reaching the commit is never evidence of a + /// write. + /// + /// The batch is small enough for this to be free: a busy production server persists ~10 + /// rows per pass, as small INSERTs into an embedded file in this process. Two passes on + /// different servers can hold open append transactions on this table at once — the read lock + /// admits concurrent holders — and both commit; measured, because widening the write window + /// under a shared lock is the one way this change could have cost something. + /// + /// #2443: the read lock and the connection open are still the LAST cancellation points on + /// this pass, unchanged by the above. Past them the batch runs to completion, and cancelling + /// before the first row costs this cycle's findings and says so. Same call the Darling twin + /// makes, and the same call AnalysisService made a layer up in #2419 — "the + /// post-enrichment tail carries no check on purpose" — restated at the write it protects. /// public async Task> InsertFindingsAsync( List findings, AnalysisContext context) @@ -108,13 +215,51 @@ public async Task> InsertFindingsAsync( if (findings.Count == 0) return findings; - /* One read lock + one connection for the whole batch, reused for every insert. */ - using var readLock = _duckDb.AcquireReadLock(); + /* One read lock + one connection for the whole batch, reused for every insert. A READ lock + around a WRITE is deliberate — it excludes compaction/archival, which is the only exclusion + this needs; see the class note (#2455). */ + using var readLock = _duckDb.AcquireReadLock(context.CancellationToken); using var connection = _duckDb.CreateConnection(); - await connection.OpenAsync(); + await connection.OpenAsync(context.CancellationToken); + + /* #2448: one transaction for the whole set — the batch commits complete or not at all. + No token check between rows: see the note above. */ + using var transaction = connection.BeginTransaction(); + + var row = 0; + var everyRowAccepted = false; + + try + { + foreach (var finding in findings) + { + row++; + await InsertFindingAsync(connection, transaction, finding); + } - foreach (var finding in findings) - await InsertFindingAsync(connection, finding); + everyRowAccepted = true; + transaction.Commit(); + } + catch (Exception ex) when (!AnalysisAbandon.IsExpected(ex, context.CancellationToken)) + { + /* #2448: the same line the Darling twin logs, for the same reason and in the same two + shapes. Without it a Lite operator gets only AnalysisService's generic "Analysis failed + for {server}", which cannot answer the question the issue said should decide this — + how often a batch actually fails partway — because it does not say a batch was even + involved, let alone which row. + + The commit gets its own branch because `row` sits at findings.Count once the loop ends, + so sharing one line would report "failed at row N of N" for a commit fault and name the + last finding as the bad one when every row had in fact been accepted. + + Logged and RETHROWN, not swallowed: letting it out is what stops AnalysisService + announcing a completed analysis for a set the store does not have, and it is the + behaviour this store has always had. Only the diagnostic is new. */ + AppLogger.Error("AnalysisService", everyRowAccepted + ? $"Finding batch for {context.ServerName} had all {findings.Count} row(s) accepted and then failed to COMMIT them, so the batch was rolled back — this analysis persisted NO findings, deliberately: a partial set would have read as a complete analysis that found fewer problems. {ex.Message}" + : $"Finding batch for {context.ServerName} failed at row {row} of {findings.Count} and was rolled back — this analysis persisted NO findings, deliberately: a partial set would have read as a complete analysis that found fewer problems. {ex.Message}"); + throw; + } return findings; } @@ -136,17 +281,37 @@ public async Task> SaveFindingsAsync( /// /// Returns the most recent findings for a server within the given time range. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. Matches the Darling twin's PgFindingStore exactly. /// public async Task> GetRecentFindingsAsync( - int serverId, int hoursBack = 24, int limit = 100) + int serverId, int hoursBack = 24, int limit = 100, DateTime? asOfUtc = null) { var findings = new List(); + /* #2506: the window's END, from which the START is measured. Null is "now" — the pre-#2506 + read exactly. */ + var windowEnd = asOfUtc ?? DateTime.UtcNow; + using var readLock = _duckDb.AcquireReadLock(); using var connection = _duckDb.CreateConnection(); await connection.OpenAsync(); using var cmd = connection.CreateCommand(); + + /* + #2506 added the UPPER bound ($3). Without it the read had a start and no end, so an as_of + anchor could only ever move the window's start EARLIER and every anchored read would still + return everything up to now — the anchor validated, the caller told the window had moved, + and the answer unchanged. That is the exact defect this convention exists to prevent, so + the bound is part of the read rather than something the caller filters afterwards. + + It binds ONLY when the caller anchored; unanchored, $3 is NoUpperBound and the read is the + half-open window it has always been. See that field for why "now" is the wrong default. + */ cmd.CommandText = @" SELECT finding_id, analysis_time, server_id, server_name, database_name, time_range_start, time_range_end, severity, confidence, category, @@ -156,11 +321,13 @@ public async Task> GetRecentFindingsAsync( FROM analysis_findings WHERE server_id = $1 AND analysis_time >= $2 +AND analysis_time <= $3 ORDER BY analysis_time DESC, severity DESC -LIMIT $3"; +LIMIT $4"; cmd.Parameters.Add(new DuckDBParameter { Value = serverId }); - cmd.Parameters.Add(new DuckDBParameter { Value = DateTime.UtcNow.AddHours(-hoursBack) }); + cmd.Parameters.Add(new DuckDBParameter { Value = windowEnd.AddHours(-hoursBack) }); + cmd.Parameters.Add(new DuckDBParameter { Value = asOfUtc is null ? NoUpperBound : windowEnd }); cmd.Parameters.Add(new DuckDBParameter { Value = limit }); using var reader = await cmd.ExecuteReaderAsync(); @@ -199,6 +366,11 @@ FROM analysis_findings /// /// Returns the latest analysis run's findings for a server (most recent analysis_time). + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. Matches the Darling twin's PgFindingStore exactly. /// public async Task> GetLatestFindingsAsync(int serverId) { @@ -260,9 +432,17 @@ SELECT MAX(analysis_time) FROM analysis_findings WHERE server_id = $1 /// /// Mutes a story pattern so it won't appear in future analysis runs. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. Matches the Darling twin's PgFindingStore exactly. /// public async Task MuteStoryAsync(int serverId, string storyPathHash, string storyPath, string? reason = null) { + /* Read lock around an INSERT: deliberate, see the class note (#2455). It admits concurrent + holders, which is exactly why the mute_id below comes from the shared generator and not + from per-instance state this lock could not have protected. */ using var readLock = _duckDb.AcquireReadLock(); using var connection = _duckDb.CreateConnection(); await connection.OpenAsync(); @@ -272,7 +452,7 @@ public async Task MuteStoryAsync(int serverId, string storyPathHash, string stor INSERT INTO analysis_muted (mute_id, server_id, story_path_hash, story_path, muted_date, reason) VALUES ($1, $2, $3, $4, $5, $6)"; - cmd.Parameters.Add(new DuckDBParameter { Value = _nextId++ }); + cmd.Parameters.Add(new DuckDBParameter { Value = NextId() }); // serverId 0 is the MCP "mute across all servers" sentinel; persist it as NULL, the // canonical global marker every reader filters on (legacy 0 rows are still honored). cmd.Parameters.Add(new DuckDBParameter { Value = serverId == 0 ? (object)DBNull.Value : serverId }); @@ -286,9 +466,17 @@ INSERT INTO analysis_muted (mute_id, server_id, story_path_hash, story_path, mut /// /// Cleans up old findings beyond the retention period. + /// + /// #2443 exempt: off the analysis pass. This surface serves the viewer, the MCP and the + /// retention sweep — lifetimes with no per-pass budget and no wedged analysis to abandon — so + /// its store calls take no pass token. Threading one here would mean inventing a caller that + /// does not exist. Matches the Darling twin's PgFindingStore exactly. /// public async Task CleanupOldFindingsAsync(int retentionDays = 30) { + /* Read lock around a DELETE: deliberate, see the class note (#2455). The rows it removes are + older than the retention cutoff and an insert batch only ever writes rows stamped now, so + the sweep and a pass overlap on disjoint rows — measured, not assumed. */ using var readLock = _duckDb.AcquireReadLock(); using var connection = _duckDb.CreateConnection(); await connection.OpenAsync(); @@ -304,7 +492,8 @@ public async Task CleanupOldFindingsAsync(int retentionDays = 30) /// lock and connection (NoRecursion lock — do not re-acquire here). Used by /// FilterMutedFindingsAsync so the mute-filter read reuses that phase's connection. /// - private static async Task> GetMutedHashesAsync(DuckDBConnection connection, int serverId) + private static async Task> GetMutedHashesAsync( + DuckDBConnection connection, int serverId, CancellationToken cancellationToken) { var hashes = new HashSet(); @@ -317,21 +506,34 @@ SELECT story_path_hash FROM analysis_muted cmd.Parameters.Add(new DuckDBParameter { Value = serverId }); - using var reader = await cmd.ExecuteReaderAsync(); - while (await reader.ReadAsync()) + using var reader = await cmd.ExecuteReaderAsync(cancellationToken); + while (await reader.ReadAsync(cancellationToken)) hashes.Add(reader.GetString(0)); return hashes; } /// - /// Inserts one finding on an already-open connection. The caller owns the read lock - /// and connection, so a batch of inserts in one InsertFindingsAsync call shares a - /// single lock acquisition and connection. + /// Inserts one finding on an already-open connection, enlisted in the batch's transaction. The + /// caller owns the read lock, the connection and the transaction, so a batch of inserts in one + /// InsertFindingsAsync call shares a single lock acquisition, connection and transaction. + /// + /// #2448: this throws rather than logging and continuing, and that is the point. Letting + /// it out skips the commit, so the batch is discarded and the pass logs its one line instead of + /// a line per remaining row for a single event. Continuing would not work here in any case — + /// DuckDB refuses every later statement once one has failed inside the transaction. + /// + /// #2443 exempt: this write deliberately takes no token. Cancelling inside a single-row + /// INSERT buys nothing — the row is microseconds of work in-process — and costs a definite + /// answer about whether it landed: DuckDB's cancel is duckdb_interrupt, so an interrupted + /// INSERT can leave a row that did commit. carries the full + /// reasoning and the abandonment point that replaces this one. /// - private static async Task InsertFindingAsync(DuckDBConnection connection, AnalysisFinding finding) + private static async Task InsertFindingAsync( + DuckDBConnection connection, DuckDBTransaction transaction, AnalysisFinding finding) { using var cmd = connection.CreateCommand(); + cmd.Transaction = transaction; cmd.CommandText = @" INSERT INTO analysis_findings (finding_id, analysis_time, server_id, server_name, database_name, diff --git a/Lite/Analysis/SqlPlanFetcher.cs b/Lite/Analysis/SqlPlanFetcher.cs index ec5b94e6e..f572ebc23 100644 --- a/Lite/Analysis/SqlPlanFetcher.cs +++ b/Lite/Analysis/SqlPlanFetcher.cs @@ -1,5 +1,6 @@ using System; using System.Linq; +using System.Threading; using System.Threading.Tasks; using Microsoft.Data.SqlClient; using PerformanceMonitor.Analysis; @@ -20,7 +21,7 @@ public SqlPlanFetcher(ServerManager serverManager) _serverManager = serverManager; } - public async Task FetchPlanXmlAsync(int serverId, string planHandle) + public async Task FetchPlanXmlAsync(int serverId, string planHandle, CancellationToken cancellationToken) { if (string.IsNullOrEmpty(planHandle)) return null; @@ -43,7 +44,7 @@ public SqlPlanFetcher(ServerManager serverManager) }; await using var connection = new SqlConnection(builder.ConnectionString); - await connection.OpenAsync(); + await connection.OpenAsync(cancellationToken); await using var cmd = new SqlCommand(@" SET NOCOUNT ON; @@ -53,13 +54,20 @@ SELECT query_plan cmd.CommandTimeout = 15; cmd.Parameters.AddWithValue("@plan_handle", planHandle); - var result = await cmd.ExecuteScalarAsync(); + var result = await cmd.ExecuteScalarAsync(cancellationToken); if (result == null || result is DBNull) return null; return result.ToString(); } - catch (Exception ex) + catch (Exception ex) when (ex is not OperationCanceledException) { + /* #2443: review caught this, and it is a real behaviour gap rather than a tidiness one. + Lite arms a genuine per-pass budget (#2412), so the token this method now takes can and + does fire mid-fetch on a healthy running app. An unconditional catch would log + "Failed to fetch plan for handle …: The operation was canceled." as an ERROR for work + we ourselves called off, AND swallow it, so the pass would carry on enriching under a + fired token instead of unwinding to AnalysisService's one quiet line. Same arm the + Darling twin's PgPlanFetcher carries for the identical call shape. */ AppLogger.Error("SqlPlanFetcher", $"Failed to fetch plan for handle {planHandle}: {ex.Message}"); return null; diff --git a/Lite/App.xaml.cs b/Lite/App.xaml.cs index 3941539ae..5a30566cb 100644 --- a/Lite/App.xaml.cs +++ b/Lite/App.xaml.cs @@ -20,6 +20,7 @@ using PerformanceMonitor.Notifications; using System.Windows.Threading; using PerformanceMonitorLite.Services; +using PerformanceMonitor.Common; using PerformanceMonitor.Ui; namespace PerformanceMonitorLite; @@ -120,6 +121,14 @@ being handed back the stale in-memory version. The coordinator owns the mutex + /// workload-dependent, and because a legitimate post-resume catch-up spike would otherwise page. public static long AgRedoQueueAlertKb { get; set; } + /// #2426: re-announce a replica that is STILL disconnected every N minutes (0 = off, the + /// shipped default, so nothing starts re-alerting on upgrade). "AG Replica Disconnected" is otherwise a + /// pure edge, which means a replica down for a week announces itself exactly once and a week-long outage + /// is indistinguishable from a momentary blip in the alert history. Re-fires deliver under the SAME + /// metric name so webhook automation keyed on it re-triggers. Darling's ag_disconnect_refire_minutes, + /// same 0-1440 clamp. + public static int AgDisconnectRefireMinutes { get; set; } + public static bool AlertCpuEnabled { get; set; } = true; public static int AlertCpuThreshold { get; set; } = 80; /// Which CPU metric the alert evaluates against. Total = sql_server_cpu + other_process_cpu (matches OS user+system). SqlOnly = SQL Server scheduler %. @@ -474,6 +483,11 @@ Connection and every collector connection are covered without per-site wiring. T // Create and show main window (StartupUri removed for Velopack custom Main) _mainWindow = new MainWindow(); _mainWindow.Show(); + + /* #2425. The log line for this was written back in LoadDefaultTimeRange/LoadAlertSettings and has + been sitting in AppLogger's buffer since; this is the visible half, and it is here rather than + beside the loaders so that it has a window behind it and so that startup order is untouched. */ + ReportUnreadableSettingsToUser(); } /// @@ -645,20 +659,205 @@ private static void TryGrantAuthenticatedUsersModify(string directoryPath) } } + /// + /// Why settings.json could not be parsed, set by whichever loader hit it first and consumed once by + /// (#2425). + /// + /// Both loaders run about thirteen statements BEFORE AppLogger.Initialize, which reads + /// like there is nowhere to report this to. There is: AppLogger.Log enqueues unconditionally and + /// only Flush is gated on initialization, so a line written from here lands in the log file a + /// few statements later — which is exactly how DataRootMigration's failure lines above already + /// reach disk. That is the deferred diagnostic, and it costs no reordering of startup. What cannot be + /// deferred to a buffer is the VISIBLE signal, which is what this field carries: someone looking at an + /// app that has forgotten its configuration should not have to find a log to learn why. + /// + private static string? s_unreadableSettingsProblem; + + /// + /// Records that settings.json is present but unparseable: to the log immediately (buffered until + /// AppLogger.Initialize) and to for the single dialog + /// shown once the main window is up. First caller wins, because both loaders read the same file and + /// would otherwise say the same thing twice. + /// + /// There is deliberately no counterpart for an ABSENT file. A first run has no settings.json, + /// defaults are the correct answer, and a warning there would be pure noise — which is precisely why + /// the old bare catch looked reasonable, since it could not tell the two apart. + /// + private static void ReportUnreadableSettings(string? problem) + { + if (s_unreadableSettingsProblem != null) + { + return; + } + + s_unreadableSettingsProblem = problem ?? "the reason could not be determined"; + + AppLogger.Error("Settings", + $"settings.json could not be parsed ({s_unreadableSettingsProblem}). Every setting it holds is " + + "at its default for this session -- including mcp_enabled, so the MCP server is OFF and nothing is " + + "listening on its port (#2431). The file has NOT been changed: fix it and restart to get the " + + "settings back, or save from the Settings window, which copies the unreadable file aside first."); + } + + /// + /// Every settings.json key whose VALUE was the wrong shape, accumulated across both loaders and consumed + /// once by (#2444). + /// + /// Separate from because they are different failures with + /// different costs, and the difference is the whole of #2444. An unparseable DOCUMENT costs every setting + /// and there is nothing to enumerate. A badly-shaped VALUE now costs exactly its own key, so the useful + /// thing to say is WHICH keys — the message that used to be shown sent someone to proofread an + /// eighty-seven-key file with no idea which line was wrong. + /// + /// They cannot both be set on one run: a document that will not parse never reaches a value read. + /// Kept as two fields anyway rather than one string, because collapsing them would make the loaders + /// responsible for deciding which message wins, and that decision belongs where the message is shown. + /// + private static readonly List s_badSettingValues = new(); + + /// + /// Records the keys a loader could not read (#2444): to the log immediately, and to + /// for the one startup dialog. + /// + /// Reports the whole SET rather than the first. A user who hand-edited one line probably has one + /// mistake, but a user who pasted a block out of settings.sample.json has several — and failing on the + /// first means they fix it, restart, and discover the next one, several times over. The old code could + /// not have done this even in principle: it threw on the first bad value and never reached the rest. + /// + private static void ReportBadSettingValues(IReadOnlyList problems) + { + if (problems.Count == 0) + { + return; + } + + foreach (var problem in problems) + { + if (!s_badSettingValues.Contains(problem.ToString(), StringComparer.Ordinal)) + { + s_badSettingValues.Add(problem.ToString()); + } + } + + /* Says "every other key this loader read", not "every other setting in the file": both loaders call + this and LoadDefaultTimeRange runs first, so a claim about the whole file would be written before + the alert settings had been read at all. The dialog CAN make the whole-file claim, because it is + shown once, after both loaders have run. */ + AppLogger.Error("Settings", + $"settings.json parsed, but {problems.Count} value(s) in it could not be read and are at their " + + "defaults for this session. Only the keys named here fell back -- every other key this loader " + + "read was applied.\n " + + string.Join("\n ", problems)); + } + + /// + /// Shows the one-time modal for an unreadable settings.json. Called after the main window exists so it + /// cannot be the app's entire first impression, and so a later startup failure still fails the way it + /// did before. + /// + /// Modal rather than a tray balloon on purpose. The app is running on settings the user did not + /// choose, and the likeliest next move is to reconfigure by hand over a file that still holds the real + /// answers — a balloon that has already faded does not stop that. + /// + private static void ReportUnreadableSettingsToUser() + { + var problem = s_unreadableSettingsProblem; + if (problem == null) + { + /* #2444: the document parsed, so if anything went wrong it was individual VALUES, and this is the + only place that says so where someone will see it. Deliberately a second message rather than a + second dialog — the two cannot both happen on one run, and an app that opens two modals before + its first window is used is an app people click through. */ + ReportBadSettingValuesToUser(); + return; + } + + MessageBox.Show( + "settings.json could not be read, so Performance Monitor Lite started with default settings.\n\n" + + $"{Path.Combine(ConfigDirectory, "settings.json")}\n{problem}\n\n" + + "The MCP server is one of those defaults, so it is OFF for this session. If you had it enabled, " + + "anything that connects to it -- an agent or a client, usually on another machine -- gets a refused " + + "connection until this is fixed, and nothing over there can tell you why.\n\n" + + "Your file has not been changed. Fix it and restart to get your settings back. If you save from " + + "the Settings window instead, the unreadable file is copied aside as " + + "settings.json.unreadable- first, so it is recoverable either way.", + "Settings", MessageBoxButton.OK, MessageBoxImage.Warning); + } + + /// + /// The visible half of #2444: the keys that fell back, by name, once the main window is up. + /// + /// Visible at all because a log line is not a startup message. The value half used to log and stop + /// there, so a user whose thresholds had silently reverted had no signal unless they went looking — and + /// #2425 had already decided, for the document half, that someone looking at an app which has forgotten + /// its configuration should not have to find a log to learn why. The same argument applies to a setting + /// that reverted; what is new is that the message can now be specific enough to act on, which is what + /// makes it worth a dialog rather than noise. + /// + /// It says what SURVIVED as well as what did not, because "a value could not be read" over an + /// eighty-seven-key file reads as "everything is gone" — which is what it used to mean. + /// + private static void ReportBadSettingValuesToUser() + { + if (s_badSettingValues.Count == 0) + { + return; + } + + MessageBox.Show( + $"{s_badSettingValues.Count} setting(s) in settings.json could not be read and are at their " + + "defaults for this session. Everything else in the file loaded normally.\n\n" + + $"{Path.Combine(ConfigDirectory, "settings.json")}\n\n" + + string.Join("\n", s_badSettingValues) + "\n\n" + + "Your file has not been changed. Fix these values and restart, or save from the Settings window " + + "to write the current values back over them.", + "Settings", MessageBoxButton.OK, MessageBoxImage.Warning); + } + private static void LoadDefaultTimeRange() { - try + var settings = SettingsFileGuard.Read(Path.Combine(ConfigDirectory, "settings.json")); + if (settings.State == SettingsFileState.Unreadable) + { + ReportUnreadableSettings(settings.Problem); + return; + } + + if (settings.Text == null) { - var path = Path.Combine(ConfigDirectory, "settings.json"); - if (!File.Exists(path)) return; + return; + } - using var doc = System.Text.Json.JsonDocument.Parse(File.ReadAllText(path)); - if (doc.RootElement.TryGetProperty("default_time_range_hours", out var val)) + try + { + /* Reads through a SettingsReader rather than JsonNode.TryGetPropertyValue, which would be the + more natural pairing with the guard's parsed Root. Two reasons: it keeps this loader the same + shape as LoadAlertSettings below, and #2418's SettingsSampleTests extracts the documented key + list by regexing the TryGetProperty key literals out of this very file. Spelled any other way + this key silently drops out of that guard, and settings.sample.json is then free to document a + setting nothing reads. (Which is also why the call is not written out in this comment: the + extractor cannot tell a comment from code, and would read the example as a real key.) */ + using var doc = System.Text.Json.JsonDocument.Parse(settings.Text); + var read = new SettingsReader(doc.RootElement); + + if (read.TryGetProperty("default_time_range_hours", out var val)) { - DefaultTimeRangeHours = val.GetInt32(); + DefaultTimeRangeHours = val.Int(DefaultTimeRangeHours); } + + /* #2444: this loader already named its one key, which is the behaviour LoadAlertSettings could + not manage across eighty-seven. It now says so through the shared reporter instead of its own + log line, so the startup dialog can name this key beside the others. */ + ReportBadSettingValues(read.Problems); + } + catch (Exception ex) + { + /* The document parsed and every value read is shape-checked rather than caught, so nothing + EXPECTED lands here any more. Kept because an unexpected throw must not take startup down. */ + AppLogger.Warn("Settings", + $"settings.json key 'default_time_range_hours' could not be read ({ex.Message}); the " + + $"default of {DefaultTimeRangeHours} hours is in use."); } - catch { /* Use default */ } } public static void LoadAlertSettings() => LoadAlertSettings(ConfigDirectory, GetWebhookUrl, SaveWebhookUrl); @@ -687,105 +886,144 @@ so opening Settings once after the data loss destroyed the surviving webhook URL GenericWebhookHeadersJson = readSecret(GenericWebhookHeadersCredentialKey); PagerDutyRoutingKey = readSecret(PagerDutyWebhookCredentialKey); - try + /* #2425: the PARSE is decided before the reads below, and separately from them. It used to share + their single try, so one trailing comma anywhere in the file threw before the first + TryGetProperty ran and reverted all eighty-eight settings at once, silently. An unreadable file + is now reported; an absent one still is not, because a first run legitimately has no file. */ + var settings = SettingsFileGuard.Read(Path.Combine(configDirectory, "settings.json")); + if (settings.State == SettingsFileState.Unreadable) { - var path = Path.Combine(configDirectory, "settings.json"); - if (!File.Exists(path)) return; + ReportUnreadableSettings(settings.Problem); + return; + } - using var doc = System.Text.Json.JsonDocument.Parse(File.ReadAllText(path)); - var root = doc.RootElement; + if (settings.Text == null) + { + return; + } - if (root.TryGetProperty("alerts_enabled", out var v)) AlertsEnabled = v.GetBoolean(); - if (root.TryGetProperty("notify_connection_changes", out v)) NotifyConnectionChanges = v.GetBoolean(); - if (root.TryGetProperty("notify_connection_down_at_startup", out v)) NotifyConnectionDownAtStartup = v.GetBoolean(); - if (root.TryGetProperty("connection_refire_minutes", out v)) ConnectionRefireMinutes = Math.Clamp(v.GetInt32(), 0, 1440); + try + { + using var doc = System.Text.Json.JsonDocument.Parse(settings.Text); + + /* #2444: every read below goes through SettingsReader, which checks ValueKind and RECORDS a key + it cannot read instead of throwing on it. Before this, all eighty-seven shared one try, so one + key of the wrong shape abandoned every read after it and which settings survived depended on + where the bad key sat in the file — an implementation detail of this method's line order. + + The reads keep their existing call shape on purpose and not by habit: #2418's + SettingsSampleTests regexes the key literals out of this file by the method NAME, and requires + every key it finds to be documented in settings.sample.json and vice versa. A helper spelled + any other way makes all eighty-seven vanish from that extraction and lets the sample drift + exactly the way #2418 was filed about -- which is what PR #2428 hit. So SettingsReader's method + carries the same name deliberately, and the extractor needed no change. (The call is not + written out here for the reason the sibling loader above gives: the extractor cannot tell a + comment from code, and would read the example as a real key.) + + The clamps moved into the reader with the reads. They were repeated inline at every call site, + and two of the forms — (int)Math.Max(0, v.GetInt64()) — narrowed AFTER the floor, so a + hand-typed value beyond int range wrapped negative. Int(fallback, min, max) clamps first. */ + var read = new SettingsReader(doc.RootElement); + + if (read.TryGetProperty("alerts_enabled", out var v)) AlertsEnabled = v.Bool(AlertsEnabled); + if (read.TryGetProperty("notify_connection_changes", out v)) NotifyConnectionChanges = v.Bool(NotifyConnectionChanges); + if (read.TryGetProperty("notify_connection_down_at_startup", out v)) NotifyConnectionDownAtStartup = v.Bool(NotifyConnectionDownAtStartup); + if (read.TryGetProperty("connection_refire_minutes", out v)) ConnectionRefireMinutes = v.Int(ConnectionRefireMinutes, 0, 1440); /* #1696 AG knobs, clamped on READ to the same ranges Darling clamps, so a hand-edited settings.json cannot drive a nonsense threshold in either app. */ - if (root.TryGetProperty("notify_ag_health", out v)) NotifyAgHealth = v.GetBoolean(); - if (root.TryGetProperty("ag_lag_alert_seconds", out v)) AgLagAlertSeconds = Math.Clamp(v.GetInt32(), 0, 86400); - if (root.TryGetProperty("ag_redo_queue_alert_kb", out v)) AgRedoQueueAlertKb = Math.Clamp(v.GetInt64(), 0L, 1073741824L); - if (root.TryGetProperty("alert_cpu_enabled", out v)) AlertCpuEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_cpu_threshold", out v)) AlertCpuThreshold = v.GetInt32(); - if (root.TryGetProperty("alert_cpu_mode", out v) && Enum.TryParse(v.GetString(), out var mode)) + if (read.TryGetProperty("notify_ag_health", out v)) NotifyAgHealth = v.Bool(NotifyAgHealth); + if (read.TryGetProperty("ag_lag_alert_seconds", out v)) AgLagAlertSeconds = v.Int(AgLagAlertSeconds, 0, 86400); + if (read.TryGetProperty("ag_redo_queue_alert_kb", out v)) AgRedoQueueAlertKb = v.Long(AgRedoQueueAlertKb, 0L, 1073741824L); + if (read.TryGetProperty("ag_disconnect_refire_minutes", out v)) AgDisconnectRefireMinutes = v.Int(AgDisconnectRefireMinutes, 0, 1440); + if (read.TryGetProperty("alert_cpu_enabled", out v)) AlertCpuEnabled = v.Bool(AlertCpuEnabled); + if (read.TryGetProperty("alert_cpu_threshold", out v)) AlertCpuThreshold = v.Int(AlertCpuThreshold); + if (read.TryGetProperty("alert_cpu_mode", out v) && Enum.TryParse(v.TextOrNull(), out var mode)) AlertCpuMode = mode; - if (root.TryGetProperty("alert_blocking_enabled", out v)) AlertBlockingEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_blocking_threshold", out v)) AlertBlockingThreshold = v.GetInt32(); - if (root.TryGetProperty("alert_blocking_wait_seconds_threshold", out v)) AlertBlockingWaitSecondsThreshold = v.GetInt32(); - if (root.TryGetProperty("alert_deadlock_enabled", out v)) AlertDeadlockEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_deadlock_threshold", out v)) AlertDeadlockThreshold = v.GetInt32(); - if (root.TryGetProperty("alert_poison_wait_enabled", out v)) AlertPoisonWaitEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_poison_wait_threshold_ms", out v)) AlertPoisonWaitThresholdMs = v.GetInt32(); - if (root.TryGetProperty("alert_long_running_query_enabled", out v)) AlertLongRunningQueryEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_query_threshold_minutes", out v)) AlertLongRunningQueryThresholdMinutes = v.GetInt32(); - if (root.TryGetProperty("alert_long_running_query_max_results", out v)) AlertLongRunningQueryMaxResults = (int)Math.Clamp(v.GetInt64(), 1, 1000); - if (root.TryGetProperty("alert_long_running_query_exclude_sp_server_diagnostics", out v)) AlertLongRunningQueryExcludeSpServerDiagnostics = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_query_exclude_waitfor", out v)) AlertLongRunningQueryExcludeWaitFor = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_query_exclude_backups", out v)) AlertLongRunningQueryExcludeBackups = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_query_exclude_misc_waits", out v)) AlertLongRunningQueryExcludeMiscWaits = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_query_exclude_cdc", out v)) AlertLongRunningQueryExcludeCdc = v.GetBoolean(); - if (root.TryGetProperty("alert_excluded_databases", out v) && v.ValueKind == System.Text.Json.JsonValueKind.Array) + if (read.TryGetProperty("alert_blocking_enabled", out v)) AlertBlockingEnabled = v.Bool(AlertBlockingEnabled); + if (read.TryGetProperty("alert_blocking_threshold", out v)) AlertBlockingThreshold = v.Int(AlertBlockingThreshold); + if (read.TryGetProperty("alert_blocking_wait_seconds_threshold", out v)) AlertBlockingWaitSecondsThreshold = v.Int(AlertBlockingWaitSecondsThreshold); + if (read.TryGetProperty("alert_deadlock_enabled", out v)) AlertDeadlockEnabled = v.Bool(AlertDeadlockEnabled); + if (read.TryGetProperty("alert_deadlock_threshold", out v)) AlertDeadlockThreshold = v.Int(AlertDeadlockThreshold); + if (read.TryGetProperty("alert_poison_wait_enabled", out v)) AlertPoisonWaitEnabled = v.Bool(AlertPoisonWaitEnabled); + if (read.TryGetProperty("alert_poison_wait_threshold_ms", out v)) AlertPoisonWaitThresholdMs = v.Int(AlertPoisonWaitThresholdMs); + if (read.TryGetProperty("alert_long_running_query_enabled", out v)) AlertLongRunningQueryEnabled = v.Bool(AlertLongRunningQueryEnabled); + if (read.TryGetProperty("alert_long_running_query_threshold_minutes", out v)) AlertLongRunningQueryThresholdMinutes = v.Int(AlertLongRunningQueryThresholdMinutes); + if (read.TryGetProperty("alert_long_running_query_max_results", out v)) AlertLongRunningQueryMaxResults = v.Int(AlertLongRunningQueryMaxResults, 1, 1000); + if (read.TryGetProperty("alert_long_running_query_exclude_sp_server_diagnostics", out v)) AlertLongRunningQueryExcludeSpServerDiagnostics = v.Bool(AlertLongRunningQueryExcludeSpServerDiagnostics); + if (read.TryGetProperty("alert_long_running_query_exclude_waitfor", out v)) AlertLongRunningQueryExcludeWaitFor = v.Bool(AlertLongRunningQueryExcludeWaitFor); + if (read.TryGetProperty("alert_long_running_query_exclude_backups", out v)) AlertLongRunningQueryExcludeBackups = v.Bool(AlertLongRunningQueryExcludeBackups); + if (read.TryGetProperty("alert_long_running_query_exclude_misc_waits", out v)) AlertLongRunningQueryExcludeMiscWaits = v.Bool(AlertLongRunningQueryExcludeMiscWaits); + if (read.TryGetProperty("alert_long_running_query_exclude_cdc", out v)) AlertLongRunningQueryExcludeCdc = v.Bool(AlertLongRunningQueryExcludeCdc); + if (read.TryGetProperty("alert_excluded_databases", out v) && v.IsArray()) { AlertExcludedDatabases = new List(); - foreach (var elem in v.EnumerateArray()) + + /* The ELEMENT kind is filtered rather than read, which the fleet-group list below already + did: elem.GetString() throws on a number in the array, and inside the old single try that + one element cost every setting after it. It is skipped rather than reported — an element + has no key to name, and the list itself is not the thing that failed. */ + foreach (var elem in v.Element.EnumerateArray()) { - var db = elem.GetString(); + var db = elem.ValueKind == JsonValueKind.String ? elem.GetString() : null; if (!string.IsNullOrWhiteSpace(db)) AlertExcludedDatabases.Add(db); } } - if (root.TryGetProperty("alert_tempdb_space_enabled", out v)) AlertTempDbSpaceEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_tempdb_space_threshold_percent", out v)) AlertTempDbSpaceThresholdPercent = v.GetInt32(); - if (root.TryGetProperty("alert_low_disk_enabled", out v)) AlertLowDiskEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_low_disk_threshold_percent", out v)) AlertLowDiskThresholdPercent = (int)Math.Clamp(v.GetInt64(), 0, 100); - if (root.TryGetProperty("alert_low_disk_threshold_gb", out v)) AlertLowDiskThresholdGb = (int)Math.Max(0, v.GetInt64()); + if (read.TryGetProperty("alert_tempdb_space_enabled", out v)) AlertTempDbSpaceEnabled = v.Bool(AlertTempDbSpaceEnabled); + if (read.TryGetProperty("alert_tempdb_space_threshold_percent", out v)) AlertTempDbSpaceThresholdPercent = v.Int(AlertTempDbSpaceThresholdPercent); + if (read.TryGetProperty("alert_low_disk_enabled", out v)) AlertLowDiskEnabled = v.Bool(AlertLowDiskEnabled); + if (read.TryGetProperty("alert_low_disk_threshold_percent", out v)) AlertLowDiskThresholdPercent = v.Int(AlertLowDiskThresholdPercent, 0, 100); + if (read.TryGetProperty("alert_low_disk_threshold_gb", out v)) AlertLowDiskThresholdGb = v.Int(AlertLowDiskThresholdGb, 0, int.MaxValue); /* #2107: the CRITICAL tier floors, clamped like the WARNING thresholds above. */ - if (root.TryGetProperty("alert_disk_critical_free_percent", out v)) AlertDiskCriticalFreePercent = Math.Clamp(v.GetInt32(), 0, 100); - if (root.TryGetProperty("alert_disk_critical_free_gb", out v)) AlertDiskCriticalFreeGb = (int)Math.Max(0, v.GetInt64()); - if (root.TryGetProperty("alert_pvs_enabled", out v)) AlertPvsEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_pvs_threshold_percent", out v)) AlertPvsThresholdPercent = (int)Math.Clamp(v.GetInt64(), 0, 100); - if (root.TryGetProperty("alert_file_growth_enabled", out v)) AlertFileGrowthEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_file_growth_rise_mb", out v)) AlertFileGrowthRiseMb = (int)Math.Max(0, v.GetInt64()); - if (root.TryGetProperty("alert_file_growth_volume_percent", out v)) AlertFileGrowthVolumePercent = (int)Math.Clamp(v.GetInt64(), 0, 100); - if (root.TryGetProperty("alert_file_growth_lookback_minutes", out v)) AlertFileGrowthLookbackMinutes = (int)Math.Clamp(v.GetInt64(), 5, 1440); - if (root.TryGetProperty("alert_pvs_floor_gb", out v)) AlertPvsFloorGb = (int)Math.Max(0, v.GetInt64()); - if (root.TryGetProperty("alert_long_running_job_enabled", out v)) AlertLongRunningJobEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_long_running_job_multiplier", out v)) AlertLongRunningJobMultiplier = v.GetInt32(); - if (root.TryGetProperty("alert_failed_job_enabled", out v)) AlertFailedJobEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_failed_job_lookback_minutes", out v)) AlertFailedJobLookbackMinutes = (int)Math.Clamp(v.GetInt64(), 1, 1440); - if (root.TryGetProperty("alert_database_state_enabled", out v)) AlertDatabaseStateEnabled = v.GetBoolean(); - if (root.TryGetProperty("alert_cooldown_minutes", out v)) AlertCooldownMinutes = (int)Math.Clamp(v.GetInt64(), 1, 120); - if (root.TryGetProperty("email_cooldown_minutes", out v)) EmailCooldownMinutes = (int)Math.Clamp(v.GetInt64(), 1, 120); - if (root.TryGetProperty("alert_delivery_mode", out v) && Enum.TryParse(v.GetString(), out var deliveryMode)) + if (read.TryGetProperty("alert_disk_critical_free_percent", out v)) AlertDiskCriticalFreePercent = v.Int(AlertDiskCriticalFreePercent, 0, 100); + if (read.TryGetProperty("alert_disk_critical_free_gb", out v)) AlertDiskCriticalFreeGb = v.Int(AlertDiskCriticalFreeGb, 0, int.MaxValue); + if (read.TryGetProperty("alert_pvs_enabled", out v)) AlertPvsEnabled = v.Bool(AlertPvsEnabled); + if (read.TryGetProperty("alert_pvs_threshold_percent", out v)) AlertPvsThresholdPercent = v.Int(AlertPvsThresholdPercent, 0, 100); + if (read.TryGetProperty("alert_file_growth_enabled", out v)) AlertFileGrowthEnabled = v.Bool(AlertFileGrowthEnabled); + if (read.TryGetProperty("alert_file_growth_rise_mb", out v)) AlertFileGrowthRiseMb = v.Int(AlertFileGrowthRiseMb, 0, int.MaxValue); + if (read.TryGetProperty("alert_file_growth_volume_percent", out v)) AlertFileGrowthVolumePercent = v.Int(AlertFileGrowthVolumePercent, 0, 100); + if (read.TryGetProperty("alert_file_growth_lookback_minutes", out v)) AlertFileGrowthLookbackMinutes = v.Int(AlertFileGrowthLookbackMinutes, 5, 1440); + if (read.TryGetProperty("alert_pvs_floor_gb", out v)) AlertPvsFloorGb = v.Int(AlertPvsFloorGb, 0, int.MaxValue); + if (read.TryGetProperty("alert_long_running_job_enabled", out v)) AlertLongRunningJobEnabled = v.Bool(AlertLongRunningJobEnabled); + if (read.TryGetProperty("alert_long_running_job_multiplier", out v)) AlertLongRunningJobMultiplier = v.Int(AlertLongRunningJobMultiplier); + if (read.TryGetProperty("alert_failed_job_enabled", out v)) AlertFailedJobEnabled = v.Bool(AlertFailedJobEnabled); + if (read.TryGetProperty("alert_failed_job_lookback_minutes", out v)) AlertFailedJobLookbackMinutes = v.Int(AlertFailedJobLookbackMinutes, 1, 1440); + if (read.TryGetProperty("alert_database_state_enabled", out v)) AlertDatabaseStateEnabled = v.Bool(AlertDatabaseStateEnabled); + if (read.TryGetProperty("alert_cooldown_minutes", out v)) AlertCooldownMinutes = v.Int(AlertCooldownMinutes, 1, 120); + if (read.TryGetProperty("email_cooldown_minutes", out v)) EmailCooldownMinutes = v.Int(EmailCooldownMinutes, 1, 120); + if (read.TryGetProperty("alert_delivery_mode", out v) && Enum.TryParse(v.TextOrNull(), out var deliveryMode)) AlertDeliveryMode = deliveryMode; - if (root.TryGetProperty("alert_per_event_max_per_cycle", out v)) AlertPerEventMaxPerCycle = (int)Math.Clamp(v.GetInt64(), 1, 100); - if (root.TryGetProperty("mute_rule_default_expiration", out v)) + if (read.TryGetProperty("alert_per_event_max_per_cycle", out v)) AlertPerEventMaxPerCycle = v.Int(AlertPerEventMaxPerCycle, 1, 100); + if (read.TryGetProperty("mute_rule_default_expiration", out v)) { - var exp = v.GetString(); + var exp = v.TextOrNull(); if (exp is "1 hour" or "24 hours" or "7 days" or "Never") MuteRuleDefaultExpiration = exp; } - if (root.TryGetProperty("log_alert_dismissals", out v)) LogAlertDismissals = v.GetBoolean(); + if (read.TryGetProperty("log_alert_dismissals", out v)) LogAlertDismissals = v.Bool(LogAlertDismissals); /* Connection settings */ - if (root.TryGetProperty("connection_timeout_seconds", out v)) + if (read.TryGetProperty("connection_timeout_seconds", out v)) { - var timeout = v.GetInt32(); + /* Rejected rather than clamped, which is why it does not use the reader's clamping overload: + an out-of-range timeout here has always been ignored in favour of the current value. */ + var timeout = v.Int(ConnectionTimeoutSeconds); if (timeout >= 5 && timeout <= 60) ConnectionTimeoutSeconds = timeout; } /* CSV export settings */ - if (root.TryGetProperty("csv_separator", out v)) + if (read.TryGetProperty("csv_separator", out v)) { - var sep = v.GetString(); + var sep = v.TextOrNull(); if (sep == "," || sep == ";" || sep == "\t") CsvSeparator = sep; } /* System tray settings */ - if (root.TryGetProperty("minimize_to_tray", out v)) MinimizeToTray = v.GetBoolean(); + if (read.TryGetProperty("minimize_to_tray", out v)) MinimizeToTray = v.Bool(MinimizeToTray); /* Time display mode */ - if (root.TryGetProperty("time_display_mode", out v)) + if (read.TryGetProperty("time_display_mode", out v)) { - var t = v.GetString(); + var t = v.TextOrNull(); if (t == "ServerTime" || t == "LocalTime" || t == "UTC") { TimeDisplayMode = t; @@ -795,62 +1033,62 @@ cannot drive a nonsense threshold in either app. */ } /* Color theme */ - if (root.TryGetProperty("color_theme", out v)) + if (read.TryGetProperty("color_theme", out v)) { - var t = v.GetString(); + var t = v.TextOrNull(); if (t == "Dark" || t == "Light" || t == "CoolBreeze") ColorTheme = t; } /* NOC Overview tile sort */ - if (root.TryGetProperty("overview_sort_mode", out v)) OverviewSortMode = ServerOverviewSort.ParseMode(v.GetString()); + if (read.TryGetProperty("overview_sort_mode", out v)) OverviewSortMode = ServerOverviewSort.ParseMode(v.TextOrNull()); /* Sidebar fleet-tree collapsed groups (#2020 2b-i-b) */ - if (root.TryGetProperty("collapsed_fleet_groups", out v) && v.ValueKind == JsonValueKind.Array) + if (read.TryGetProperty("collapsed_fleet_groups", out v) && v.IsArray()) { - CollapsedFleetGroups = v.EnumerateArray() + CollapsedFleetGroups = v.Element.EnumerateArray() .Where(e => e.ValueKind == JsonValueKind.String) .Select(e => e.GetString()!) .ToList(); } /* Update check settings */ - if (root.TryGetProperty("check_for_updates_on_startup", out v)) CheckForUpdatesOnStartup = v.GetBoolean(); + if (read.TryGetProperty("check_for_updates_on_startup", out v)) CheckForUpdatesOnStartup = v.Bool(CheckForUpdatesOnStartup); /* Teams webhook settings */ - if (root.TryGetProperty("teams_webhook_enabled", out v)) TeamsWebhookEnabled = v.GetBoolean(); - if (root.TryGetProperty("teams_proxy_address", out v)) TeamsProxyAddress = v.GetString() ?? ""; + if (read.TryGetProperty("teams_webhook_enabled", out v)) TeamsWebhookEnabled = v.Bool(TeamsWebhookEnabled); + if (read.TryGetProperty("teams_proxy_address", out v)) TeamsProxyAddress = v.Text(TeamsProxyAddress); /* Slack webhook settings */ - if (root.TryGetProperty("slack_webhook_enabled", out v)) SlackWebhookEnabled = v.GetBoolean(); - if (root.TryGetProperty("slack_proxy_address", out v)) SlackProxyAddress = v.GetString() ?? ""; + if (read.TryGetProperty("slack_webhook_enabled", out v)) SlackWebhookEnabled = v.Bool(SlackWebhookEnabled); + if (read.TryGetProperty("slack_proxy_address", out v)) SlackProxyAddress = v.Text(SlackProxyAddress); /* Generic webhook settings (#1506). The URL + headers JSON are secrets and load from Credential Manager below; only these three are plain prefs. */ - if (root.TryGetProperty("generic_webhook_enabled", out v)) GenericWebhookEnabled = v.GetBoolean(); - if (root.TryGetProperty("generic_proxy_address", out v)) GenericWebhookProxyAddress = v.GetString() ?? ""; - if (root.TryGetProperty("generic_body_template", out v)) GenericWebhookBodyTemplate = v.GetString() ?? ""; + if (read.TryGetProperty("generic_webhook_enabled", out v)) GenericWebhookEnabled = v.Bool(GenericWebhookEnabled); + if (read.TryGetProperty("generic_proxy_address", out v)) GenericWebhookProxyAddress = v.Text(GenericWebhookProxyAddress); + if (read.TryGetProperty("generic_body_template", out v)) GenericWebhookBodyTemplate = v.Text(GenericWebhookBodyTemplate); /* PagerDuty webhook settings. The routing key is a secret and loads from Credential Manager below; only the enable flag and EU-region toggle are plain prefs. */ - if (root.TryGetProperty("pagerduty_webhook_enabled", out v)) PagerDutyWebhookEnabled = v.GetBoolean(); - if (root.TryGetProperty("pagerduty_use_eu_region", out v)) PagerDutyUseEuRegion = v.GetBoolean(); - if (root.TryGetProperty("pagerduty_proxy_address", out v)) PagerDutyProxyAddress = v.GetString() ?? ""; + if (read.TryGetProperty("pagerduty_webhook_enabled", out v)) PagerDutyWebhookEnabled = v.Bool(PagerDutyWebhookEnabled); + if (read.TryGetProperty("pagerduty_use_eu_region", out v)) PagerDutyUseEuRegion = v.Bool(PagerDutyUseEuRegion); + if (read.TryGetProperty("pagerduty_proxy_address", out v)) PagerDutyProxyAddress = v.Text(PagerDutyProxyAddress); /* Migrate webhook URLs from plaintext settings.json to Credential Manager. A legacy plaintext URL still wins over whatever the store held, matching the old order (save, then read back); the live property is set here rather than re-reading, since we just wrote the value. */ - if (root.TryGetProperty("teams_webhook_url", out v)) + if (read.TryGetProperty("teams_webhook_url", out v)) { - var legacyUrl = v.GetString() ?? ""; + var legacyUrl = v.Text(""); if (!string.IsNullOrWhiteSpace(legacyUrl)) { writeSecret(TeamsWebhookCredentialKey, legacyUrl); TeamsWebhookUrl = legacyUrl; } } - if (root.TryGetProperty("slack_webhook_url", out v)) + if (read.TryGetProperty("slack_webhook_url", out v)) { - var legacyUrl = v.GetString() ?? ""; + var legacyUrl = v.Text(""); if (!string.IsNullOrWhiteSpace(legacyUrl)) { writeSecret(SlackWebhookCredentialKey, legacyUrl); @@ -859,41 +1097,124 @@ only the enable flag and EU-region toggle are plain prefs. */ } /* SMTP settings */ - if (root.TryGetProperty("smtp_enabled", out v)) SmtpEnabled = v.GetBoolean(); - if (root.TryGetProperty("smtp_server", out v)) SmtpServer = v.GetString() ?? ""; - if (root.TryGetProperty("smtp_port", out v)) SmtpPort = v.GetInt32(); - if (root.TryGetProperty("smtp_use_ssl", out v)) SmtpUseSsl = v.GetBoolean(); - if (root.TryGetProperty("smtp_username", out v)) SmtpUsername = v.GetString() ?? ""; - if (root.TryGetProperty("smtp_from_address", out v)) SmtpFromAddress = v.GetString() ?? ""; - if (root.TryGetProperty("smtp_recipients", out v)) SmtpRecipients = v.GetString() ?? ""; - - if (root.TryGetProperty("analysis_enabled", out v)) AnalysisEnabled = v.GetBoolean(); - if (root.TryGetProperty("query_store_backfill_enabled", out v)) QueryStoreBackfillEnabled = v.GetBoolean(); - if (root.TryGetProperty("analysis_notifications_enabled", out v)) AnalysisNotificationsEnabled = v.GetBoolean(); - if (root.TryGetProperty("analysis_interval_minutes", out v)) AnalysisIntervalMinutes = (int)Math.Clamp(v.GetInt64(), 5, 360); - if (root.TryGetProperty("analysis_notify_severity", out v)) AnalysisNotifySeverity = Math.Clamp(v.GetDouble(), 0.0, 2.0); - if (root.TryGetProperty("analysis_notify_cooldown_minutes", out v)) AnalysisNotifyCooldownMinutes = (int)Math.Clamp(v.GetInt64(), 30, 10080); - if (root.TryGetProperty("analysis_timeout_seconds", out v)) AnalysisTimeoutSeconds = (int)Math.Clamp(v.GetInt64(), 30, 600); - } - catch { /* Use defaults */ } + if (read.TryGetProperty("smtp_enabled", out v)) SmtpEnabled = v.Bool(SmtpEnabled); + if (read.TryGetProperty("smtp_server", out v)) SmtpServer = v.Text(SmtpServer); + if (read.TryGetProperty("smtp_port", out v)) SmtpPort = v.Int(SmtpPort); + if (read.TryGetProperty("smtp_use_ssl", out v)) SmtpUseSsl = v.Bool(SmtpUseSsl); + if (read.TryGetProperty("smtp_username", out v)) SmtpUsername = v.Text(SmtpUsername); + if (read.TryGetProperty("smtp_from_address", out v)) SmtpFromAddress = v.Text(SmtpFromAddress); + if (read.TryGetProperty("smtp_recipients", out v)) SmtpRecipients = v.Text(SmtpRecipients); + + if (read.TryGetProperty("analysis_enabled", out v)) AnalysisEnabled = v.Bool(AnalysisEnabled); + if (read.TryGetProperty("query_store_backfill_enabled", out v)) QueryStoreBackfillEnabled = v.Bool(QueryStoreBackfillEnabled); + if (read.TryGetProperty("analysis_notifications_enabled", out v)) AnalysisNotificationsEnabled = v.Bool(AnalysisNotificationsEnabled); + if (read.TryGetProperty("analysis_interval_minutes", out v)) AnalysisIntervalMinutes = v.Int(AnalysisIntervalMinutes, 5, 360); + if (read.TryGetProperty("analysis_notify_severity", out v)) AnalysisNotifySeverity = v.Double(AnalysisNotifySeverity, 0.0, 2.0); + if (read.TryGetProperty("analysis_notify_cooldown_minutes", out v)) AnalysisNotifyCooldownMinutes = v.Int(AnalysisNotifyCooldownMinutes, 30, 10080); + if (read.TryGetProperty("analysis_timeout_seconds", out v)) AnalysisTimeoutSeconds = v.Int(AnalysisTimeoutSeconds, 30, 600); + + /* #2444: reported AFTER every read, which is the point — the whole set, named, and every key that + was fine applied. Empty on a healthy file, so this costs nothing on the normal path. */ + ReportBadSettingValues(read.Problems); + } + catch (Exception ex) + { + /* SettingsFileGuard already parsed this exact text, so nothing reaching here is a document-level + fault — and since #2444 nothing reaching here is a badly-shaped VALUE either, because those are + checked by kind and recorded rather than thrown. So this catch no longer has a known cause and + is kept only so that an unforeseen one cannot take startup down with it. Note what it costs + when it does fire: the reads are ordered, so anything after the throw is still lost, which is + the residue this fix removes for every case it can name. */ + AppLogger.Error("Settings", + "settings.json parsed, but reading its values failed unexpectedly, so some alert settings " + + "are at their defaults for this session", ex); + } } /// - /// Reads settings.json (or starts fresh), applies , and writes it back - /// indented; logs and swallows any error under . Shared by the single-value - /// Save* methods (and MainWindow's Overview sort selector) so the read/merge/write/catch boilerplate - /// lives in one place. + /// The JSON object every settings.json writer in Lite merges into — this one, the four Save* methods in + /// SettingsWindow, and anything added later. + /// + /// Every Save here rewrites the WHOLE document, so the read in front of it decides whether a Save + /// merges or replaces (#2425). When the file is present but unparseable, merging is impossible and + /// replacing destroys the only copy of what the user actually configured — including through saves + /// nobody thinks of as saves, such as collapsing a sidebar group. So the unreadable file is copied + /// aside first and the caller merges into a fresh object. + /// + /// When even the copy cannot be made this throws rather than handing back an empty object. Every + /// call site already wraps its write in a catch that logs, so the file is left exactly as it was and the + /// failure is reported — which is the right way round when the alternative is permanent loss. /// - public static void WriteSetting(string what, Action mutate) + internal static JsonObject SettingsRootForWrite() + { + var settingsPath = Path.Combine(ConfigDirectory, "settings.json"); + var forWrite = SettingsFileGuard.RootForWrite(settingsPath, DateTime.Now); + + if (forWrite.Problem == null) + { + return forWrite.Root; + } + + if (forWrite.QuarantinedTo == null) + { + throw new IOException( + $"settings.json cannot be parsed ({forWrite.Problem}) and no copy of it could be made, so it " + + "has been left untouched rather than overwritten with defaults. Fix it, or move it aside by " + + "hand, and save again."); + } + + AppLogger.Warn("Settings", + $"settings.json could not be parsed ({forWrite.Problem}), so this save rewrites it from defaults. " + + $"The unreadable original was copied to '{Path.GetFileName(forWrite.QuarantinedTo)}' first — the " + + "settings it held are recoverable from there."); + + return forWrite.Root; + } + + /// + /// The one place settings.json is written, and the one that answers whether the write happened (#2433). + /// + /// Before this, five methods each ended in their own File.WriteAllText inside their own + /// catch, and not one of them could tell its caller anything. The Settings window said "Settings saved." + /// whether or not a single byte reached disk, because the only place the truth existed was the log, and + /// nobody reads a log after a dialog says it worked. A bool is the smallest thing that fixes that, and + /// having exactly one writer is what makes the bool mean something: there is no longer a save that + /// half happened. + /// + /// names the thing being saved so the log line is about the operator's + /// action rather than about a filename. + /// + internal static bool WriteSettingsDocument(JsonNode root, string what) { var settingsPath = Path.Combine(ConfigDirectory, "settings.json"); + try { - JsonNode root = File.Exists(settingsPath) - ? JsonNode.Parse(File.ReadAllText(settingsPath)) ?? new JsonObject() - : new JsonObject(); - mutate(root); File.WriteAllText(settingsPath, root.ToJsonString(new JsonSerializerOptions { WriteIndented = true })); + return true; + } + catch (Exception ex) + { + AppLogger.Error("Settings", $"Failed to save {what}: settings.json could not be written ({ex.Message})"); + return false; + } + } + + /// + /// Applies to settings.json and writes it back indented; logs and swallows any + /// error under . This is the SINGLE-VALUE path — MainWindow's Overview sort + /// selector and anything else that changes one knob outside the Settings window's Save button, which + /// opens the document once for all of its writers instead (#2433). The read is + /// , which is what keeps an unparseable file from being replaced + /// unrecorded. + /// + public static void WriteSetting(string what, Action mutate) + { + try + { + JsonNode root = SettingsRootForWrite(); + mutate(root); + WriteSettingsDocument(root, what); } catch (Exception ex) { diff --git a/Lite/Controls/CorrelatedTimelineLanesControl.xaml.cs b/Lite/Controls/CorrelatedTimelineLanesControl.xaml.cs index 5801a1168..ffe56849c 100644 --- a/Lite/Controls/CorrelatedTimelineLanesControl.xaml.cs +++ b/Lite/Controls/CorrelatedTimelineLanesControl.xaml.cs @@ -581,10 +581,16 @@ private void SyncXAxes(int hoursBack, DateTime? fromDate, DateTime? toDate, doub var charts = new[] { CpuChart, WaitStatsChart, BlockingChart, MemoryChart, FileIoChart }; foreach (var chart in charts) - { chart.Plot.Axes.SetLimitsX(xMin, xMax); + + /* #2535: same defect as #2533 on the Darling viewer, same fix. Every lane's X and Y limits are + final by now, so this is the one place that can see all five gutters at once and hand them the + widest. It has to run on EVERY refresh, not once at Initialize: ClearChart() calls + WpfPlot.Reset(), which swaps in a fresh Plot and takes any axis floor with it. */ + LaneAxisAligner.AlignLeftGutters(charts); + + foreach (var chart in charts) chart.Refresh(); - } } /// diff --git a/Lite/Controls/FinOpsTab.xaml b/Lite/Controls/FinOpsTab.xaml index 649285e17..b7a9cd64f 100644 --- a/Lite/Controls/FinOpsTab.xaml +++ b/Lite/Controls/FinOpsTab.xaml @@ -23,6 +23,22 @@ + + + + + + + + + + + + + + @@ -983,7 +1006,8 @@ + SelectionMode="Single" RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -1041,7 +1065,8 @@ - @@ -1159,7 +1184,8 @@ CanUserResizeColumns="True" HeadersVisibility="Column" SelectionMode="Extended" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -1362,7 +1388,8 @@ CanUserResizeColumns="True" HeadersVisibility="Column" SelectionMode="Extended" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -1589,7 +1616,8 @@ CanUserResizeColumns="True" HeadersVisibility="Column" SelectionMode="Extended" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -1861,7 +1889,8 @@ CanUserResizeColumns="True" HeadersVisibility="Column" SelectionMode="Extended" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -2049,7 +2078,8 @@ HeadersVisibility="Column" SelectionMode="Extended" MaxHeight="250" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -2111,7 +2141,8 @@ HeadersVisibility="Column" SelectionMode="Extended" MaxHeight="200" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> @@ -2187,7 +2218,8 @@ HeadersVisibility="Column" SelectionMode="Extended" MaxHeight="250" - RowStyle="{StaticResource DefaultRowStyle}"> + RowStyle="{StaticResource DefaultRowStyle}" + LoadingRow="MarkedGrid_LoadingRow"> - public IDisposable AcquireReadLock() + public IDisposable AcquireReadLock() => AcquireReadLock(CancellationToken.None); + + /// + /// The same read lock, abandonable (#2443). has taken a timeout + /// since it was written and this one took nothing, which left a real hole in the analysis budget: + /// every store read on the pass takes this lock BEFORE it opens its connection, so a pass queued + /// behind a long archival or a compaction sat in an uninterruptible EnterReadLock() however + /// carefully its reads were threaded. Cancellation reached the reads and stopped at the door. + /// + /// A token rather than a timeout, because a timeout would need a NUMBER and there is no + /// honest one to pick — the right budget for waiting on this lock is exactly the caller's remaining + /// budget, which the caller already holds. A poll is the only way to say that to + /// , which has no token-taking overload; the interval is only + /// reached under contention, since an uncontended TryEnterReadLock returns immediately. + /// takes the original uninterruptible path, so the fourteen + /// callers outside the analysis pass are unchanged rather than quietly re-timed. + /// + public IDisposable AcquireReadLock(CancellationToken cancellationToken) { try { - s_dbLock.EnterReadLock(); + if (!cancellationToken.CanBeCanceled) + { + s_dbLock.EnterReadLock(); + } + else + { + while (!s_dbLock.TryEnterReadLock(ReadLockPollInterval)) + { + cancellationToken.ThrowIfCancellationRequested(); + } + } } catch (LockRecursionException) { @@ -47,6 +139,56 @@ exception that prevented Dispose(). Since we're already protected by a read lock return new LockReleaser(s_dbLock, write: false); } + /// + /// How long a cancellable read-lock wait blocks before re-checking its token. Only reached under + /// contention, so it costs nothing on the normal path; 50 ms keeps the worst-case overshoot far + /// below the smallest budget anyone would set while leaving the wait almost entirely in the kernel + /// rather than spinning. + /// + private static readonly TimeSpan ReadLockPollInterval = TimeSpan.FromMilliseconds(50); + + /// + /// The read lock for a caller that would rather have NO answer than wait for one — returns null if + /// the lock is not free within , instead of blocking. + /// + /// This exists for the status bar (#2594). Every other read here takes the lock and waits, + /// which is right when the caller needs the data; the status-bar size figure is cosmetic, refreshed + /// on a 30-second UI timer, and runs on the dispatcher thread — so waiting on a lock held by a long + /// archival would freeze the window to render a number nobody is reading. The two honest options for + /// that caller are "skip the lock" and "skip the read", and skipping the lock is how a connection + /// ends up open against a file the reset path deletes. + /// + /// A timeout rather than a token here, unlike , + /// because this caller genuinely has a number to pick: it is bounded by what a UI thread may spend, + /// not by a caller's remaining budget. + /// + public IDisposable? TryAcquireReadLock(TimeSpan timeout) + { + try + { + if (!s_dbLock.TryEnterReadLock(timeout)) + { + return null; + } + } + catch (LockRecursionException) + { + /* Already held by this thread — same reasoning as AcquireReadLock: we are protected, so + hand back a no-op rather than reporting a failure the caller cannot act on. */ + return NoOpDisposable.Instance; + } + + return new LockReleaser(s_dbLock, write: false); + } + + /// + /// How long the status bar will wait for the read lock before giving up and reporting no used-size + /// figure. Short on purpose: this runs on the dispatcher thread, and the caller already renders the + /// file size alone when this returns null, so the degraded answer is a smaller status bar rather + /// than a stalled window. + /// + private static readonly TimeSpan StatusBarReadLockTimeout = TimeSpan.FromMilliseconds(100); + /// /// Acquires an exclusive write lock on the database. Blocks until all readers finish. /// Dispose the returned object to release the lock. @@ -74,6 +216,35 @@ private sealed class NoOpDisposable : IDisposable public void Dispose() { } } + /// + /// Releases the lock entry its constructor was handed. + /// + /// Thread affinity is a LATENT hazard here, and the obvious mitigation is worse than the + /// hazard (#2463). is thread-affine: ExitReadLock throws + /// when called from a thread that did not enter. Every one + /// of these locks is held across await — AcquireReadLock, then OpenAsync, then the + /// reads — and the analysis pass runs under Task.Run with no SynchronizationContext, so a + /// continuation is free to resume on a different pool thread. It does not, and the reason is measured + /// rather than lucky: DuckDB.NET 1.5.5's async methods complete synchronously, so no await ever + /// yields and the entering thread is always the exiting thread. Thread id across OpenAsync, a + /// DDL statement, 200 ExecuteNonQueryAsync calls and a reader drain: 4, 4, 4, 4, 4. + /// + /// Do not "harden" this with an IsReadLockHeld guard. It looks like the cheap fix + /// and it is a bug amplifier, measured: on a thread that did not enter, IsReadLockHeld is + /// false, so the guard SKIPS the exit — and the entry the original thread took is then held + /// forever. With it still held, TryEnterWriteLock(400 ms) fails. So the guard would convert a + /// loud, attributable into a silently leaked reader that + /// permanently wedges every maintenance operation in the process: no CHECKPOINT, no archival, no + /// compaction, for the life of the app. Throwing is the better failure. + /// + /// A real fix would have to make the releaser not thread-affine at all — replacing + /// with something like a SemaphoreSlim, which would also + /// retire 's poll — and that is a change to the lock + /// primitive, not to this class. It is not worth doing while the premise holds. What would break the + /// premise is a DuckDB.NET release whose async genuinely yields, so + /// DuckDbLockModelTests.DuckDbAsyncStillCompletesOnTheCallingThread is a tripwire on exactly + /// that: if it goes red after a driver bump, this paragraph is where to start. + /// private sealed class LockReleaser : IDisposable { private readonly ReaderWriterLockSlim _lock; @@ -98,7 +269,7 @@ public void Dispose() /// /// Current schema version. Increment this when schema changes require table rebuilds. /// - internal const int CurrentSchemaVersion = 54; + internal const int CurrentSchemaVersion = 56; private readonly string _archivePath; @@ -1244,6 +1415,65 @@ information rather than a broken alert path. */ _logger?.LogWarning("Migration to v54 encountered an error (non-fatal): {Error}", ex.Message); } } + + if (fromVersion < 55) + { + /* v55 (#2472): the per-database fan-out rollup on collection_log, twinning Darling's V80. One + collector run that fans out over N databases writes ONE row whose duration_ms is the sum, so + "eight databases at 10.1s" and "one at 62s beside seven at 2.7s" are the same number and want + opposite fixes. These three carry the ratio that separates them. + + Fresh installs get the columns from GetAllTableStatements(); this is for an existing database + and is idempotent. Nothing to backfill and nothing that COULD be: a row written before the + upgrade genuinely does not know its fan-out, so NULL is the honest value rather than a zero + that would read as "fanned out over nothing". + + Non-fatal, matching the blocks above: without the columns the writer's INSERT would fail and + take collection logging with it, so a failed ADD COLUMN must not also break the run — the + write path treats a log failure as non-fatal already. */ + _logger?.LogInformation("Running migration to v55: adding collection_log fan-out rollup columns"); + try + { + await ExecuteNonQueryAsync(connection, + "ALTER TABLE collection_log ADD COLUMN IF NOT EXISTS fanout_item_count INTEGER"); + await ExecuteNonQueryAsync(connection, + "ALTER TABLE collection_log ADD COLUMN IF NOT EXISTS slowest_item VARCHAR"); + await ExecuteNonQueryAsync(connection, + "ALTER TABLE collection_log ADD COLUMN IF NOT EXISTS slowest_item_ms INTEGER"); + } + catch (Exception ex) + { + _logger?.LogWarning("Migration to v55 encountered an error (non-fatal): {Error}", ex.Message); + } + } + + if (fromVersion < 56) + { + /* v56 (#2515): tempdb's growth CEILING on tempdb_stats, twinning Darling's V81. Every other + column here comes from dm_db_file_space_usage, which reports the data files AS CURRENTLY + ALLOCATED — so the tempdb Space alert's percentage measured distance to the next AUTOGROW + rather than distance to the point where tempdb cannot grow at all. SUM(max_size) over the + ROWS files is the denominator that means the same thing on every engine. + + Fresh installs get the column from GetAllTableStatements(); this is for an existing database + and is idempotent. Nothing to backfill: a row collected before the upgrade genuinely does not + know the ceiling, and NULL says so — the read maps it to 0, which is the "no ceiling measured" + state where the denominator stays the allocation and the reported percentage does not move. + + The v_tempdb_stats view needs no work here: Lite rebuilds every v_ passthrough on start + (CreateArchiveViewsAsync, called after this), which is the difference from Darling, where the + view's SELECT * column list is frozen at CREATE and the rung has to refresh it. */ + _logger?.LogInformation("Running migration to v56: adding max_size_mb to tempdb_stats"); + try + { + await ExecuteNonQueryAsync(connection, + "ALTER TABLE tempdb_stats ADD COLUMN IF NOT EXISTS max_size_mb DECIMAL(18,2)"); + } + catch (Exception ex) + { + _logger?.LogWarning("Migration to v56 encountered an error (non-fatal): {Error}", ex.Message); + } + } } /// @@ -1508,6 +1738,17 @@ public double GetDatabaseSizeMb() /// public double? GetUsedDataSizeMb() { + /* Under the read lock (#2594). This was the one periodic connection site that opened without it, + on the UI's 30-second timer, which put a live handle on the database file at moments the + archival path may be deleting and recreating it. Bounded rather than blocking - see + StatusBarReadLockTimeout for why this caller may give up where others may not. */ + using var readLock = TryAcquireReadLock(StatusBarReadLockTimeout); + + if (readLock is null) + { + return null; + } + try { using var connection = CreateConnection(); diff --git a/Lite/Database/Schema.cs b/Lite/Database/Schema.cs index 92e7db1c6..3a53cfbbf 100644 --- a/Lite/Database/Schema.cs +++ b/Lite/Database/Schema.cs @@ -56,7 +56,10 @@ CREATE TABLE IF NOT EXISTS collection_log ( error_message VARCHAR, rows_collected INTEGER, sql_duration_ms INTEGER, - duckdb_duration_ms INTEGER + duckdb_duration_ms INTEGER, + fanout_item_count INTEGER, + slowest_item VARCHAR, + slowest_item_ms INTEGER )"; public const string CreateCollectionLogIndex = @" diff --git a/Lite/MainWindow.AlertEngine.cs b/Lite/MainWindow.AlertEngine.cs index a6dadf881..f6d3fd7f5 100644 --- a/Lite/MainWindow.AlertEngine.cs +++ b/Lite/MainWindow.AlertEngine.cs @@ -302,7 +302,12 @@ evaluating on some launches with nothing in the log. Capturing also removes the var (replicaTimeUtc, replicas) = await data.GetLatestAgReplicaStatesAsync(serverId); if (IsFresh(replicaTimeUtc)) { - alerts.AddRange(_agAlertEvaluator.EvaluateReplicas(serverId, replicas)); + alerts.AddRange(_agAlertEvaluator.EvaluateReplicas( + serverId, + replicas, + App.AgDisconnectRefireMinutes > 0 + ? TimeSpan.FromMinutes(App.AgDisconnectRefireMinutes) + : null)); } var (databaseTimeUtc, databases) = await data.GetLatestAgDatabaseReplicaStatesAsync(serverId); @@ -324,6 +329,12 @@ evaluating on some launches with nothing in the log. Capturing also removes the foreach (var alert in alerts) { SendAgAlert(serverId, serverName, alert); + + /* #2426: the re-fire window opens on DELIVERY, which is what the suppressed return above + makes necessary — an acknowledged or silenced server still evaluates every sweep and + sends nothing, and windows consumed by alerts nobody received would turn the re-fire back + into the silence it exists to end. */ + _agAlertEvaluator.NoteDelivered(alert); } bool IsFresh(DateTime? snapshotUtc) => diff --git a/Lite/MainWindow.xaml b/Lite/MainWindow.xaml index 03c4dd7c3..cddf8fd02 100644 --- a/Lite/MainWindow.xaml +++ b/Lite/MainWindow.xaml @@ -65,12 +65,23 @@ - + +