Skip to content

Hyperscale log file reads n/a (log service) and stays out of allocated database sizes - #4880

Merged
erikdarlingdata merged 5 commits into
devfrom
fix/hyperscale-log-size-not-allocated
Oct 1, 2026
Merged

erikdarlingdata merged 5 commits into
devfrom
fix/hyperscale-log-size-not-allocated

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

What does this PR do?

On Azure SQL Database Hyperscale, Database Sizes now shows the log file's size as "n/a (log service)" and leaves the log out of every allocated total. The data file keeps its real size. Storage Growth leaves the log file out of its current, 7-day, and 30-day sizes, so its growth is the data file's own. SQL Server, Managed Instance, and other Azure SQL Database tiers do not change. The fix covers Lite and Darling (the service, the viewer, and the web page).

Which component(s) does this affect?

  • Lite
  • Darling
  • Lite Tests
  • Darling Tests
  • SQL collection scripts
  • Documentation
  • Full Dashboard (deprecated)
  • CLI Installer (deprecated)

Why

A Hyperscale database keeps its transaction log in a separate log service. sys.database_files still lists a LOG row for it, sized at about 1 TB (1,046,528 MB). The Hyperscale FAQ gives that 1 TB as the cap on the active log. It is not storage that the database holds. The Hyperscale service tier page says you pay for storage based on actual allocation. It also says storage is allocated automatically between 10 GB and 128 TB: https://learn.microsoft.com/en-us/azure/azure-sql/database/service-tier-hyperscale

Before this change, one Hyperscale database showed about 1,056,768 MB allocated against 315 MB used. That total is the 10,240 MB data file plus the 1,046,528 MB log row. The log row made the allocation look about 100 times too large.

What changes

The collector (shared by Lite and Darling):

  • The Azure SQL Database query now detects Hyperscale with DATABASEPROPERTYEX(DB_NAME(), N'Edition') = N'Hyperscale'.
  • On Hyperscale, the LOG row (df.type = 1) returns NULL for total_size_mb, max_size_mb, auto_growth_mb, growth_pct, and is_percent_growth. The max size and growth settings are log-service values like the 1 TB size, and no reader needs them.
  • used_size_mb and vlf_count stay as collected. The data file row is unchanged.
  • The SQL Server query is unchanged.

The stores:

  • Lite (DuckDB): total_size_mb was NOT NULL, so DuckDB rejected the NULL row. Schema version 67 drops that constraint. DuckDB refuses the change while the table has an index, so the migration drops idx_database_size_stats_time first, as version 57 does. The startup index pass recreates it.
  • Darling (PostgreSQL): the generated column is already nullable, so there is no migration. A test pins that.

Every reader, in both apps:

  • The shared words live in PerformanceMonitor.Common (HyperscaleLogSize), so both apps use the same text: n/a (log service). FileIoStatsCollector.NoSizeLabel, the File I/O text for the same file, now points at HyperscaleLogSize.Display, so one constant owns the words.
  • Lite and the Darling viewer: the Database Sizes grid shows that text for a null size (TargetNullValue). The free space and used percent of that row are blank. Reads no longer turn a NULL size into 0.
  • Allocated totals, the free total, the cost share per database, and the idle-database recommendations leave out a file with no size. The utilization chart sums used space over the same files as the allocated figure, so the two stay on one footing.
  • The Darling viewer sorts by size with NULLS LAST. PostgreSQL sorts a null first under DESC.
  • The get_database_sizes MCP tool (Lite and Darling): the per-file total_size_mb, auto_growth_mb, and max_size_mb are null for that row. The database total and used figures leave the file out. The payload adds a top-level note only when a null exists, so every other server returns the same shape as before.
  • The note reads: "On Azure SQL Database Hyperscale, the transaction log lives in the log service, not in storage the database holds. The log file's size is n/a (log service), and the totals leave it out." It names no JSON field, because people read it on the web page.
  • The Darling web page: the Database Sizes table shows the read's note above the rows.
  • File-growth alerts already skip a row with no total size. So do the analysis rules that test for files of 10 GB or more. The log row drops out of them, and that code is unchanged.

Storage Growth (Lite LocalDataService.StorageGrowthSql, Darling viewer ViewerDataService.StorageGrowthSql):

  • The grid compares each database's latest size with its size 7 and 30 days ago. Rows stored before this change still hold the 1 TB log figure. When only the latest sum left the log out, the database above read as 10,240 MB now against 1,056,768 MB before. That is a drop of about 99%.
  • One predicate now leaves out any file whose row in the latest snapshot has no size. Both queries apply it in the same way to the latest sum, the 7-day sum, and the 30-day sum. It runs when the grid reads, so the old rows need no migration and no backfill.
  • A file that is gone from the latest snapshot has no row there, so the predicate does not match it. Its old size still counts, and the grid shows the shrinkage.
  • On the latest side, the predicate drops only the rows that SUM already skips. It is there so that all three sums follow one rule.
  • In the Darling viewer, the predicate binds the same latest collection_time ($2) as the latest sum, so the query still binds only literal snapshot times.

Not touched: deprecated/ and the install scripts.

Why this predicate

There were two choices. One leaves out a file whose latest row has no size. The other leaves out a LOG file on a Hyperscale database. This PR uses the first.

  • It reads only database_size_stats, the table that Storage Growth already reads, and it keys on the NULL that the collector writes for this one row.
  • database_size_stats does not store the edition. A LOG-plus-Hyperscale predicate must join server_properties and match its edition text.
  • Until the first collection after an upgrade, the latest snapshot still holds the 1 TB log row. The collector runs every 60 minutes. Under either choice, Database Sizes shows that stored figure and counts it in the totals, because that grid reads the stored rows.
  • With the latest-row predicate, Storage Growth then counts the row on both sides. Its current size matches the Database Sizes total, and its growth is the data file's own.
  • With the LOG-plus-Hyperscale predicate, Storage Growth leaves the row out at once. Its current size then does not match the Database Sizes total until the next collection.

Neither choice covers one case. A database that moved to Hyperscale in the last 30 days has old log rows with a real allocation. Both predicates leave those rows out, so the release of that log allocation does not show as shrinkage.

Other readers of file sizes

Only the two Storage Growth queries compare a past sum with the latest one. These readers were checked, and they need no change:

  • Database Sizes grid, its totals, and the utilization chart (both apps): the latest snapshot only.
  • get_database_sizes (both apps): the latest snapshot only.
  • The FinOps server inventory storage total (LocalDataService.FinOps.ServerProperties.cs, ViewerDataService.FinOps.Inventory.cs) and the idle-database list: the latest snapshot only.
  • The file-growth alert reads (LocalDataService.FileGrowth.cs, DarlingAlertReadAdapter.cs): each file is compared with its own earlier row, and a file whose current size is NULL is skipped. Both sides drop the log file together.
  • Analysis facts: the database size fact reads File I/O sizes, which File I/O on Azure SQL Database Hyperscale shows the real data file size and no size for the log file #4878 handled. The large-file growth facts read the latest row per file and need 10 GB or more, which a NULL fails. The disk-space fact reads only the volume columns.
  • No days-until-full projection reads database_size_stats. The projections that exist are for PostgreSQL targets and for the Darling store's own disk.
  • No MCP tool or web panel compares database file sizes over time. The web page's growth columns are table sizes.
  • Custom Views: the "Database file size" measure charts the stored rows for each file. It does not compare two points in time. A chart over the old rows shows the stored 1 TB for that file until those rows age out after 90 days. A Custom Views measure has no row filter of its own, so this PR does not change it.

How was this tested?

No SQL Server, Azure SQL Database, or PostgreSQL connection was used. The checks are unit tests and real DuckDB files. The Hyperscale edition string was not checked against a live Hyperscale database (see "Still to check").

Test plan

  • Both test projects build with 0 warnings and 0 errors, after merging the latest dev. That build includes Lite, the Darling service and viewer, the collectors, and the common project.
  • New tests: Lite.Tests/HyperscaleLogSizeTests.cs (16 tests), Darling/Darling.Tests/HyperscaleLogSizeTests.cs (10 tests), one upgrade test in AgedDatabaseMigrationTests, and one relaxation recorded in DuckDbSchemaEquivalenceTests.
  • Targeted run on the final head: Lite HyperscaleLogSize* and FileIoHyperscaleLogSizeTests, 18 tests, 0 failed. Darling HyperscaleLogSizeTests and FileIoHyperscaleLogSizeTests, 15 tests, 0 failed. Darling ViewerFinOpsSqlTests and DarlingMcpObjectStatsToolsSurfaceAndSqlTests, 72 tests, 0 failed.
  • Red on the old behavior. With the behavior reverted and the new tests kept, 10 Lite tests and 5 Darling tests fail. The red lines are listed below.
  • Red with the Storage Growth predicate removed from each sum in turn. The red lines are listed below.
  • Full Lite.Tests suite on the final head: 6,452 tests, 0 failed, 0 skipped.
  • Full Darling.Tests suite on the final head: 19,077 tests, 0 failed, 1,197 skipped, 1 not run. The skips are live PostgreSQL classes, which need a store. The one not run is an explicit-only test (McpSchemaCompatServiceLeakRaceTests).
  • Live PostgreSQL test classes. They skip without a store, so they did not run. DatabaseSizeLatestPlanShapeLiveTests changed only so that it compiles with a nullable size. Its seed has no NULL size, so the Storage Growth predicate leaves its results unchanged.
  • A live Hyperscale database. Nobody has run the new query against one.

Red lines (the new tests fail when the behavior is reverted):

Lite.Tests.HyperscaleLogSizeCollectorTests.AzureSqlDb_Query_NullsTheSizeOfTheHyperscaleLogRowOnly_AndKeepsTheRealSizeOtherwise [FAIL]
Lite.Tests.HyperscaleLogSizeCollectorTests.ReadAsync_NullTotalSize_IsKeptNull_AndDoesNotKillTheBatch [FAIL]
Lite.Tests.HyperscaleLogSizeCollectorTests.AzureSqlDb_Query_DetectsHyperscaleFromTheEditionOfTheCurrentDatabase [FAIL]
PerformanceMonitorLite.Tests.DuckDbSchemaEquivalenceTests.GeneratedCollectorTables_AreStorageEquivalentToPreChangeHandWritten [FAIL]
Lite.Tests.HyperscaleLogSizeReadTests.DatabaseSizesGrid_HyperscaleLogRow_IsNull_AndOutOfTheAllocatedTotals_WhileTheDataRowCounts [FAIL]
Lite.Tests.HyperscaleLogSizeReadTests.GetDatabaseSizesPayload_HyperscaleLogFile_IsNullWithTheNote_AndOutOfTheDatabaseTotal [FAIL]
PerformanceMonitorLite.Tests.AgedDatabaseMigrationTests.UpgradeFromV66_DropsTotalSizeMbNotNull_EvenWithAPreExistingIndex [FAIL]
Lite.Tests.HyperscaleLogSizeReadTests.TheGridWordsANullSize_AsTheSharedHyperscaleText [FAIL]
Lite.Tests.HyperscaleLogSizeReadTests.UtilizationChart_HyperscaleDatabase_AllocatesTheDataFileOnly_AndUsedStaysOnTheSameFooting [FAIL]
Lite.Tests.HyperscaleLogSizeReadTests.GetDatabaseSizesTool_Read_PassesTheNullSizeGrowthAndCeilingThroughAsNull [FAIL]
Darling.Tests.HyperscaleLogSizeTests.ViewerGrid_WordsANullSize_AsTheSharedHyperscaleText_AndSortsNullSizesLast [FAIL]
Darling.Tests.HyperscaleLogSizeTests.ViewerUtilizationChart_UsedIsSummedOnlyOverTheFilesWhoseSizeCounts [FAIL]
Darling.Tests.HyperscaleLogSizeTests.WebDatabaseSizesTable_RendersTheReadsNote [FAIL]
Darling.Tests.HyperscaleLogSizeTests.AzureSqlDbQuery_NullsTheSizeOfTheHyperscaleLogRowOnly_AndKeepsTheRealSizeOtherwise [FAIL]
Darling.Tests.HyperscaleLogSizeTests.GetDatabaseSizesPayload_HyperscaleLogFile_IsNullWithTheNote_AndOutOfTheDatabaseTotal [FAIL]

How the red run was made: these pieces went back to their old form. They were the collector query and read, the two DuckDB schema files, and every reader's null handling. They also included the MCP note branch, the NULLS LAST sort, the chart sum, both TargetNullValue bindings, and the web note argument. The row model stayed nullable so the test projects still compile.

Red lines for the Storage Growth predicate. Each run removed the predicate from one sum, rebuilt, and ran the Storage Growth tests:

Lite, predicate removed from the 7-day sum:
Lite.Tests.HyperscaleLogSizeReadTests.StorageGrowth_HyperscaleHistoryHoldsTheOneTerabyteLogRow_ShowsTheDataFileGrowth_NotA99PercentDrop [FAIL]
  Assert.Equal() Failure: Values differ  Expected: 140  Actual: -1046388

Lite, predicate removed from the 30-day sum:
Lite.Tests.HyperscaleLogSizeReadTests.StorageGrowth_HyperscaleHistoryHoldsTheOneTerabyteLogRow_ShowsTheDataFileGrowth_NotA99PercentDrop [FAIL]
  Assert.Equal() Failure: Values differ  Expected: 2.4  Actual: -99.0308

Lite, predicate removed from the latest sum:
Lite.Tests.HyperscaleLogSizeReadTests.StorageGrowthSql_AppliesOnePredicateToTheLatestAnd7dAnd30dSums [FAIL]
  The latest sum does not leave the log-service file out.

Darling, predicate removed from each sum in turn:
Darling.Tests.HyperscaleLogSizeTests.ViewerStorageGrowth_AppliesOnePredicateToTheLatestAnd7dAnd30dSums [FAIL]
  The latest sum does not leave the log-service file out.
  The past_7d sum does not leave the log-service file out.
  The past_30d sum does not leave the log-service file out.
  • On the latest side, the predicate drops only the rows that SUM already skips. A behavior test cannot fail there, so the SQL text test holds that side, in both apps.
  • Darling's store is PostgreSQL, and its tests need a live store. The Darling proof is the SQL text test. Lite runs the same predicate text on a real DuckDB store.
  • StorageGrowth_AFileGoneFromTheLatestSnapshot_StillCountsAsShrinkage_AndARealLogCounts and StorageGrowth_LatestSnapshotStillFromBeforeTheChange_CountsTheLogRowOnBothSides_LikeTheDatabaseSizesGrid pass on the old code too. They guard behavior the predicate must keep. A dropped file still counts, and a log row with a real size and the same file_id still counts. A latest snapshot from before the upgrade still shows the data file's growth.

Tests that cannot fail on the old behavior, and why:

  • OnPrem_Query_IsUnchanged_NoHyperscaleBranch and OnPremQuery_IsUnchanged_NoHyperscaleBranch guard behavior that must not change, so they pass before and after.
  • PostgresStore_KeepsTotalSizeNullable and GetDatabaseSizesPayload_NoLogServiceFile_HasNoNote_AndTheSameShapeAsBefore (both apps) also guard behavior that must not change.
  • Row_WithoutAnAllocation_ReadsAsNotApplicable_WithoutThrowing, TheSharedNote_NamesTheWordsTheGridShows, ViewerRow_HyperscaleLogRow_ReadsAsNotApplicable_AndOutOfTheAllocatedTotals_WhileTheDataRowCounts, and the two payload-shape tests use members that only this change adds: a nullable size, AllocatedTotalMb, FreeTotalMb, HyperscaleLogSize, and DatabaseSizesPayload. On the old code they do not compile, so they were not run red.

Still to check

  • The edition string. The DATABASEPROPERTYEX page lists these Edition values: General Purpose, Business Critical, Basic, Standard, Premium, System, FabricSQLDB, and NULL. It does not list Hyperscale. The same tier name is documented for sys.database_service_objectives.edition and for the PowerShell Edition value. That view needs the dbmanager role, so the collector cannot rely on it. The query uses the DATABASEPROPERTYEX form. If a real Hyperscale database returns a different string, the log row keeps its old size. The new tests still pass, because they pin the query text. One run of SELECT DATABASEPROPERTYEX(DB_NAME(), N'Edition') on a Hyperscale database settles it.
  • Old history. Rows stored before this change keep the old 1 TB log figure until they age out after 90 days. Storage Growth leaves them out when it reads. The latest-size reads and the allocated totals correct themselves at the next collection. Custom Views still draws them (see "Other readers of file sizes").
  • Used space on the log row. The collected value stays in the file rows of the MCP payload. It is left out of the database totals.

Checklist

  • I have read the contributing guide
  • My code builds with zero warnings (dotnet build -c Debug)
  • I have tested my changes against at least one SQL Server version
  • I have not introduced any hardcoded credentials or server names

CHANGELOG

SECTION: Fixed
ENTRY:

… leave it out of allocated totals

On Azure SQL Database Hyperscale, sys.database_files reports the LOG file at about 1 TB (1,046,528 MB) while the data file reports its real allocation. The log lives in Hyperscale's log service, so that figure is not storage the database holds or pays for (Hyperscale bills allocated data storage). Database Sizes in both apps showed the database at 1,056,768 MB allocated against 315 MB used.

The Azure SQL Database branch of the database_size_stats collector now detects Hyperscale from DATABASEPROPERTYEX(DB_NAME(), 'Edition') and stores a NULL size, growth and ceiling for the LOG row only. The data file keeps its size, and used space and the VLF count stay as collected. SQL Server, Managed Instance and non-Hyperscale Azure SQL Database are unchanged.

Lite's DuckDB store drops NOT NULL from database_size_stats.total_size_mb (schema v67); Darling's Postgres store already held it nullable. Every reader in both apps passes the NULL through instead of reading it as 0: the Database Sizes grid words it as n/a (log service) and sorts it last, the allocated and free totals, the cost share and the utilization chart leave it out, and get_database_sizes returns null with a short note in both apps. The web Database Sizes table shows that note.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 30, 2026 22:40
@erikdarlingdata
erikdarlingdata marked this pull request as draft September 30, 2026 22:50
…-day sizes too

Rows stored before the collector wrote NULL for the Hyperscale log row still hold its ~1 TB. With only the latest sum leaving the log out, a Hyperscale database read as a drop of about 99%.

- One predicate leaves out any file whose row in the latest snapshot has no size. Both apps' Storage Growth queries apply it to the latest, 7-day and 30-day sums alike. It runs at read time, so the old rows need no migration.
- A file that is gone from the latest snapshot has no row there, so it still counts on the past side, as shrinkage.
- FileIoStatsCollector.NoSizeLabel points at HyperscaleLogSize.Display, so one constant owns "n/a (log service)".
- The get_database_sizes note, which the web Database Sizes table shows, is in plain words and names no field.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 1, 2026 00:22
@erikdarlingdata
erikdarlingdata merged commit 43b0241 into dev Oct 1, 2026
18 of 20 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/hyperscale-log-size-not-allocated branch October 1, 2026 00:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant