Darling: a runtime update never clears the only runtime that opens the store, and an unfinished update says how to recover - #4935
Merged
Conversation
…e's runtime back even over a partial one, and a store with no TimescaleDB record is taken to be on 2.28.1
…nly one that opens the store, and never clears it then
…cks the start-time restore
…ine marker is not refused for a missing TimescaleDB record
…one in the start-time restore
…s none, and the restore's heal and its failed re-extract are pinned end to end
…ot open the store, and it is written before the rescue move
…says so at Error and names the file to delete
erikdarlingdata
marked this pull request as ready for review
October 2, 2026 06:31
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #4934.
What was wrong
A runtime update moves the live runtime aside into
pg-runtime-prev, extracts the new one, and (on a major change) runspg_upgrade. Until the new runtime is known to open the store, the copy inpg-runtime-previs the only one that does. Nothing recorded that. A failed extract whose revert met a lock could leave a partialpgsqlthat already heldbin\pg_ctl.exe. The next start then treated it as a normal runtime: if that partialpg_ctlanswered with the store's major, the update clearedpg-runtime-prev(deleting the only good runtime) and re-extracted, and a second failed extract left no runtime that opens the store.Separately, the restore's TimescaleDB check abstained when the store had no TimescaleDB record. Every 3.x store without a record is on TimescaleDB 2.28.1, so a rescued runtime without 2.28.1 could be put back in front of a store it cannot open.
The fix
A rescue marker.
pg-runtime-prev\rescue-in-progressexists exactly whilepg-runtime-prevholds the only runtime known to open the store. It holds the hash of the package being installed.pgsqlis known to open the store again: after a same-major swap's stamp once the new runtime answers with the store's major; when a major swap's in-place upgrade commits; whenRevertRuntimesucceeds; when the failed-extract revert has put the old runtime back; after the start-time restore has put the rescued runtime back; and at the "stamp names the package" early return when the live runtime opens the store.pg-runtime-prevunless that folder provably cannot open the store, or the update that wrote the marker provably finished. The marker is stale only when: the folder has nopg_ctl.exe; there is no store (PG_VERSIONabsent); the rescued runtime and the store both report a PostgreSQL major and they differ; the rescued runtime lacks the store's TimescaleDB; or the marker's package equals the runtime stamp (the stamp is written only after a good extract, so that update finished). A stale marker is removed, with the reason logged, and the update proceeds. Anything else defers, including a version probe that timed out or gave no answer, so a good rescued runtime behind a failed probe is never deleted. The deferral repeats on every start and names the marker file to delete; it is logged at Error when the live runtime cannot answer with the store's major, and at Warning otherwise.DarlingManagedPostgres.InterruptedRuntimeUpdateHint), the start failure is logged at Error and its message adds: "An earlier runtime update did not finish; the runtime that last opened the store is at \pgsql. Delete to let the next start clear it and re-extract.", with the full paths. It reads the marker's existence only and changes no state.pgsql\bin\pg_ctl.exeis present (that runtime is no proof of a finished update). A store with no stamp file, or a stamp cut short to zero length, still restores; the swap then re-runs in the same start and writes the stamp. The running-server, same-major and TimescaleDB checks still apply; a stamp that exists but cannot be read still refuses; and a stamp that names the package (a major swap waiting for its data upgrade) still refuses so that upgrade resumes. A marker that outlived a finished update counts as no marker here too.TimescaleDB with no record. Outside the marker, the restore's check takes a store with no TimescaleDB record to be on 2.28.1, the version every 3.2-3.8 release shipped; a store with no extension at all therefore keeps the first-run extract. Under the marker it abstains, because the marker proves the rescued runtime was this store's live runtime moments before. A store whose record lists no versions still does not block.
What it costs
pg_ctl.exelanded, then a rescued runtime that stops answering its version probe for good. Every start then defers and the store fails to start; the start-failure Error above names the file to delete.Pins
In
DarlingStoreUpgradeTests, using the existing seams, so they run on every OS:RevertRuntimeremoves it;pgsqlwithpg_ctl.exegets the rescued runtime back; no stamp, or an empty main stamp, still restores, and an empty stamp is restored and then swapped in the same start; an unreadable stamp, a running server and a stamp naming the package each move nothing; no TimescaleDB record with a rescued runtime lacking 2.28.1 still restores;pgsqlwithpg_ctl.exemoves nothing; no record and a rescued runtime without 2.28.1 moves nothing, with 2.28.1 it restores; the 3.8.0-to-3.9.0 legacy-stamp case still restores;DarlingManagedPostgresTests: the start-failure hint names the marker and the rescued runtime when the marker exists and is null without it; a source pin checks the start-failure path adds it and logs at Error before it throws.The new pins that change behaviour fail on the code before this change and pass with it.
CHANGELOG
None: this extends #4934's Fixed entry.