Lite FinOps tests: CPU samples sit inside the 24-hour utilization read at any time of day - #4891
Merged
Merged
Conversation
…d at any time of day The FinOps CPU scenarios seeded their samples from TestPeriodStart, which is anchored to 04:00 UTC. Between 03:45 and 04:00 UTC every sample was more than 24 hours old, GetUtilizationEfficiencyAsync found none, and the CPU and VM right-sizing tests failed. TestDataSeeder takes an optional clock and places the FinOps CPU samples back from it (the newest 5 minutes before now). The analysis scenarios keep the 04:00 anchor. The clean FinOps server's second, now-relative CPU copy is folded into the single seed. FinOpsCpuSampleWindowTests seeds each FinOps CPU scenario at a fixed 03:50 and 12:00 UTC and checks every sample is inside the read's window, and pins that window to the read's source.
… that goes back to the anchored seed fails on any day
erikdarlingdata
marked this pull request as ready for review
October 1, 2026 04:54
erikdarlingdata
marked this pull request as draft
October 1, 2026 05:17
…reaches its standard-deviation guard Rule 14 (reserved capacity) returns no row under 24 samples. With 16 samples the clean server never reached the stddevCpu > 0 guard, and no test held it. SeedFinOpsCpuUtilizationAsync takes a sample count, and the clean server seeds 32. SeedCpuUtilizationAsync's doc says its samples are anchored for the analysis scenarios and points FinOps scenarios to SeedFinOpsCpuUtilizationAsync and the window test's list.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Six Lite
FinOpsTestsfail every day between 03:45 and 04:00 UTC and pass the rest of the day. They failed CI on #4882, which ran in that quarter hour.The FinOps CPU scenarios seeded their CPU samples from
TestDataSeeder.TestPeriodStart. That time is anchored to 04:00 UTC, because the analysis scenarios need a fixed midnight (#2177, #4385). Before 04:00 the anchor is 04:00 the day before, so the 16 samples sat at 00:00 to 03:45 the day before.GetUtilizationEfficiencyAsyncreads the last 24 hours from now (LocalDataService.FinOps.Utilization.cs,cutoff = DateTime.UtcNow.AddHours(-24)). After 03:45 it found no sample.HasCpuSamplewas false, so the CPU and VM right-sizing rules gave no advice, and the tests that expect that advice failed.What changes
Test-only. No product file changes.
TestDataSeedertakes an optional clock (Func<DateTime>). When it is left out, the seeder uses the real UTC clock. The FinOps CPU samples are placed back from that clock: 15 minutes apart, the newest 5 minutes before now (FinOpsCpuSampleTimes).SeedFinOpsCpuUtilizationAsynctakes the sample count.stddevCpu > 0guard. With 32 it does. The CPU is a flat 50, a standard deviation of 0, so that guard is what keeps rule 14 quiet.SeedCpuUtilizationAsyncworks as before. Its doc now says its samples are placed from that anchor. It sends a new FinOps scenario toSeedFinOpsCpuUtilizationAsyncand to the window test's scenario list.FinOpsCpuSampleWindowTests:GetUtilizationEfficiencyAsyncstill reads the last 24 hours. It covers thecutoffline, the parameter that carries it, and thecpu_statsfiltercollection_time >= $2.The FinOps reads the seeder feeds
GetUtilizationEfficiencyAsync,cpu_stats(rules 2 and 12 need itsHasCpuSample)GetUtilizationEfficiencyAsync,grantsmemory_grant_stats. Only analysis scenarios do.GetUtilizationEfficiencyAsync,mem_latestandserver_infov_memory_statsGetIdleDatabasesAsync,v_query_statsGetDatabaseSizeLatestAsynccollection_timeGetHighImpactQueriesAsynchoursBackfrom nowhoursBackthat reachesTestPeriodStart(#4558), with a pin at 03:59.The anchored samples are never more than about 28 hours old, so every 7-day window covers them.
The other 24-hour FinOps reads get no data from
TestDataSeeder. They areGetServerMetricsAsync,GetTempdbSummaryAsync,GetMemoryGrantEfficiencyAsyncand the workload reads. Their tests (FinOpsFleetReadParityTests,FinOpsServerInventoryTestsand theAzureSqlDatabase*Tests) seed fromDateTime.UtcNow.Darling
No change needed. The viewer's FinOps reads use the same rolling windows:
ViewerDataService.FinOps.Utilization.csreads the last 24 hours. But Darling.Tests has no anchored seeder, and its FinOps tests seed from now.ServerInventoryFleetMetricsLiveTestsseeds CPU at 1, 2 and 3 hours ago.FinOpsMergedReadsLiveTestsandViewerFinOpsIntervalHonestLiveTestsseed at 1 hour ago.ViewerFinOpsRecommendationsTestsare source pins.Test plan
Total: 15, Errors: 0, Failed: 7. The messages read likeOverProvisionedEnterprise at 03:50 UTC: 0 of its 16 CPU samples are inside the 24-hour utilization readandStableCpuForReservedCapacity at 03:50 UTC: 16 of its 32 CPU samples are inside the 24-hour utilization read.FinOpsCpuSampleWindowTestsandFinOpsTeststogetherTotal: 37, Errors: 0, Failed: 0.cpu_statsfilter changed from>=to>: the source pin fails.stddevCpu > 0guard dropped:CleanServer_NoDuckDbRecommendationsfails on a reserved-capacity row,Stable CPU utilization (avg 50.0%, CV 0.00).Failed: 0:FinOpsTests(22),FactCollectorTests(42),FactCollectorMiseryTests(23),AzureSqlDatabaseMemoryScopeTests(28),AzureSqlDatabaseHostMathTests(28),AzureSqlDatabaseHardwareTests(9),FinOpsVerdictSourcePinTests(8),FinOpsFleetReadParityTests(5).Lite.TestsandDarling.Testsbuild with 0 warnings and 0 errors.Lite.Testssuite at the head:Lite.Tests Total: 6571, Errors: 0, Failed: 0, Skipped: 0, Not Run: 0, Time: 290.839s.Darling.Testssuite at the head, with the live-target variables unset:Darling.Tests Total: 19186, Errors: 0, Failed: 0, Skipped: 1210, Not Run: 1, Time: 147.389s. The skips are the live tests that need a target.CHANGELOG
SECTION: None