Darling: a runtime update retries a briefly locked folder before it defers, and a start after an interrupted update puts the store's own runtime back - #4934
Merged
Conversation
… defers the update
… seconds before it defers the update
… the deferral pins don't sleep
… briefly locked folder too
…update puts back the runtime that last opened the store
…e revert's move aside is pinned
… update with no server running
…s moved aside only when it was
…e that cannot load the store's TimescaleDB
erikdarlingdata
marked this pull request as ready for review
October 2, 2026 05:02
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
When a package ships a new PostgreSQL runtime, the update clears the last update's rescued runtime (
pg-runtime-prev) and then renames the live runtime into it. Just after the store stops, an antivirus scan or the exiting server can hold a file in one of those folders for a moment. Either step gave up on the first IOException or UnauthorizedAccessException, so the host kept the old runtime until the next service start, which for a service can be weeks.Two related gaps in the same update:
pgsqland the service exited. The extract failure itself was not logged when a revert move then failed.pg_ctl.exeand extracted the package as a first run. That moved forward when the extract had failed only for a moment, but never started at all when the package itself was bad, and for a later major-version update it would have put a new-major runtime in front of an old-major store.The fix
pg-runtime-prev).DarlingStoreUpgrade.TryRestoreRescuedRuntimeAsync, called once at the top ofDarlingManagedPostgres.EnsureRuntimeAsync), then takes the normal path, which sees the old stamp and retries the update. It acts only when every one of these holds, checked cheapest first: nopgsql\bin\pg_ctl.exe; a rescued runtime of the store's own PostgreSQL major; a readable runtime stamp (the main stamp, or the legacy one only when the main stamp file is missing); no server running on the store; the rescued runtime carries every TimescaleDB version the store records; the shipped package is present and readable; and the stamp differs from the shipped package. The advance writes the stamp only after a good extract, so an interrupted update leaves the old stamp, while a finished one matches the package. A host on 3.8.0 has only the legacy stamp, which never matches a 3.9 package, so an interrupted first 3.9.0 update is covered. A finished host that later receives a newer package also passes the stamp test; the TimescaleDB check stops a rescued runtime that could not open the store, and a rescued runtime that can open it is swapped forward on the same start. A partialpgsqlis moved aside and deleted, and the rescued runtime is moved back, both under the retry. In every other case it changes nothing, and the start re-extracts the shipped package as before; so does a restore move that still fails after its retries.Why retrying is safe
A failed
Directory.Moveleaves the source in place (a same-volume rename is one operation). The clear only runs once the rescued-runtime guard has decided the folder may be cleared, and clearing is repeatable after a partial pass.What still falls back, and what it costs
pg-runtime-prevshould be, a runtime folder the service account cannot write, a move across volumes) now wait about 5.5 s before the same deferral, on each start with a pending update.pg-runtime-prevemptied and the live runtime in place, the same state the old fallback left. A stop during the revert can wait up to about 11 s.pg_ctl.exeafter a finished update (the stamp matches the package), a host with no readable stamp, a rescued runtime that lacks the store's TimescaleDB, or a server running on the store keeps today's first-run extract.Tests
Three internal seams on
DarlingStoreUpgrade(the wait, the rename and the clear), defaulting to today's calls, let the portable pins stand in for a lock, because a file lock blocks rename and delete on Windows only.FileShare.None, released after about a second) run on Windows CI and skip elsewhere.pgsqlwith an interrupted stamp gets the rescued runtime back (the Warning says whether an incomplete runtime was moved aside), and so does the 3.8.0 shape with only the legacy stamp; a rescued runtime that carries the store's TimescaleDB is restored. Nothing moves when there is no store, no rescued copy, a rescued copy of another major, a livepg_ctl.exe, no stamp, an empty main stamp, a server running on the store, a rescued runtime without the store's TimescaleDB, a missing or unreadable package, a stamp that matches the package (in any case), or both stamps with the main one matching. A restore locked twice retries and restores; one that keeps failing logs and returns false. A source pin checks the call sits before the live-runtime check inEnsureRuntimeAsync.CHANGELOG
SECTION: Fixed
ENTRY: - A Darling update that brings a new PostgreSQL runtime no longer waits for the next restart when the old runtime folder is briefly locked ([#4934]) - Just after the store stops, an antivirus scan or the exiting server can hold the runtime folder for a moment, and the update then kept the old runtime until the service next started. It now retries for a few seconds first, and only then falls back as before. A start that finds the runtime missing after an interrupted update now puts back the runtime that last opened the store, and retries the update, instead of extracting the new one as if this were a first run.
REF: [#4934]: #4934