Stop settle() spinning while a task is pending its first run - #90
Merged
Merged
Conversation
`_driveToStableFixpoint` holds its quiet window open while a registered task has not started (`hasPendingStartWork`). When that was the only pending work, every wait in the loop resolved synchronously on idle queues, so the loop re-checked in a tight spin without giving up its thread. A task spawned without the harness executor starts on the cooperative pool; once enough tests spun they held every pool thread, the pending tasks never started, and the pool-hosted trait-cap watchdogs starved as well. Downstream this pinned ~10 cores for 23+ minutes in a single settle(). Sleep a 1 ms poll interval on the timer in that state instead, freeing the thread. SettlePendingStartSpinTests reproduces the deadlock on a single-thread task executor: it hit its 30 s trait cap before the fix and passes in under 0.1 s after it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Reported from parallel-phoenix-apple CI (#2444, m2pro, run 36601393927).
PathInteractionTests.overviewTapped()sat 23+ minutes insettle()at about 1000% CPU. 1157 of 1164 samples were in_driveToStableFixpoint → hasPendingStartWork → AnyContext.hasPendingStartTask → allChildren, and the 30 s trait cap never fired. The whole Mac was starved (load 265)._driveToStableFixpointholds its quiet window open while any registered task hasn't run yet. When that was the only pending work, with the executor, background queue and main-observation queue all idle, every wait in the loop resolved synchronously. A continuation resumed inside its own body doesn't suspend the task, so the loop re-checked in a tight spin without ever giving up its thread.A task spawned without the harness executor (e.g. from a main-queue callback) starts on the cooperative pool. Once enough tests spin at the same time they hold every pool thread (≈ 10 on that machine, matching the ~1000% CPU). The pending tasks then never start, and no settle can finish. The
_withTestTimeoutwatchdogs are task-group children on the same pool, so they starve too and never cancel the test. That's why the trait cap never fired: the second reported issue is a symptom of the first.Fix
When a task that hasn't started is the only thing keeping the model from idle, the loop sleeps a 1 ms poll interval on the timer (
_gtsSleep, a real suspension that frees the thread) before re-checking. It's a poll cadence, not a verdict; no wall-clock decision changes. Polling rather than waiting for an event is deliberate: a task spawned off the harness executor may not see this test'sModelAccess, so its start isn't reliably observable.Validation
SettlePendingStartSpinTestsreproduces the deadlock deterministically on one thread.settle()runs on a single-thread task executor, and a registered task that hasn't started needs that thread to get going. Before the fix it spun, confirmed by instrumenting the loop (pending=truewith all queues idle), until the 30 s trait cap cancelled it. After the fix it passes in about 0.08 s.scripts/test(parallel) and--no-parallelare both green: 747 + 124 + 7 + 29 tests. Wall time fell compared with earlier today on the same machine: parallel 27.6 → 15.8 s, serial 191 → 78 s.fadeParksOnFrozenClockAndStepsWithIt, trait timeout); the fix failed 0 of 6 and was faster in 4 of 6 rounds. An earlier--loop 10on the fix had 3 early iterations fail under load ~120–190. The failing tests,fadeParksOnFrozenClockAndStepsWithItamong them, also time out on main under that load, so I don't count them against the fix.🤖 Generated with Claude Code