Conversation
The error returned by spindownDisk was printed and then discarded, and SpunDown was set to true regardless. Since the spindown branch is gated on !ds.SpunDown and only real disk activity clears that flag, an idle disk whose spindown failed was never retried: the disk kept spinning while hd-idle reported it as parked. Keep SpunDown false when the command fails so the next idle period tries again. LastSpunDownAt is still recorded on failure, which rate-limits the retry to one per idle period rather than one per poll.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
spindownDisk's error is printed and then discarded, andSpunDownis set totrueeither way (hdidle.go:156-161on master):The spindown branch is gated on
!ds.SpunDown, and only observed disk activity clears that flag (hdidle.go:176). So when the command fails on an idle disk, the disk is never retried: it keeps spinning while hd-idle considers it parked.logSpinuponly writes to the logfile on the spin-up transition, which never comes, so the logfile stays empty too.The observable result is one error line at the moment of failure and then permanent silence, with the disk spinning 24/7. That is the same symptom described in #131 ("first spindown works, later ones do not"), reachable without any state-tracking regression.
This is a different root cause from #113, where the disk wakes without registering activity in
/proc/diskstatsso the wake is never observed. That one needs a way to learn the disk's real state. This one does not: the error is already in hand at line 156 and is simply dropped.Fix
Keep
SpunDownfalse when the command fails, so the next idle period tries again.LastSpunDownAtis still recorded on failure, which rate-limits the retry to one per idle period instead of one per poll, so a disk that always rejects the command does not spam the log everypoolInterval.updateStatenow takes the spindown function as a parameter, following the existingdiskHolderGetterFuncpattern indiskstats/snapshot.go, so the failure path is testable without touching hardware.Tests
hdidle_test.go(new, table-driven per AGENTS.md):TestSpindownResultDeterminesSpunDownState: success setsSpunDownandSpinDownAt; failure leaves both unset.TestFailedSpindownIsRetriedOncePerIdlePeriod: after a failure there is no retry within the idle period and exactly one retry after it.go test ./... -race -coverpasses,go vet ./...is clean, andgofmt -lis clean for both changed files. Notegofmt -lalready reportsdiskstats/snapshot.goon unmodified master (a missing space after a comma at line 143); left untouched as it is unrelated.Hardware verification
Real Seagate ST10000NM017B (10 TB SATA), Linux 6.18, both binaries built from this branch's base. Ground truth was a 64 MB
iflag=directread at a fresh offset, which reliably separates a spinning disk from a stopped one: about 0.28 s versus about 9.5 s including spin-up.Failure path, forced with
-c ataon a transport that rejects opcode 0x85,-i 20, over 75 s (about 37 polls):spunDown=truespunDown=falseThree attempts over 75 s at
-i 20is one per idle period, as intended.Success path,
-c scsi, two full cycles:spindown, thenspinupdetected on the waking read, thenspindownagain, with the 64 MB probe measuring 9.51 s and 9.68 s at each stop and the logfile recording bothrunning:/stopped:pairs. No behavior change versus master on this path.One note on the setup, in case it is useful: the probe reads had to go through a partition (
/dev/sde2) rather than the whole disk (/dev/sde), because master computes per-disk activity from the sum of the partition counters, so whole-disk I/O does not register. v1.22 behaves differently here (it carries 675d30f, which master does not). Not related to this change, just why the test reads target a partition.