Skip to content

Darling: a runtime update retries a briefly locked folder before it defers, and a start after an interrupted update puts the store's own runtime back - #4934

Merged
erikdarlingdata merged 12 commits into
devfrom
fix/runtime-rescue-retry
Oct 2, 2026
Merged

erikdarlingdata merged 12 commits into
devfrom
fix/runtime-rescue-retry

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

What was wrong

When a package ships a new PostgreSQL runtime, the update clears the last update's rescued runtime (pg-runtime-prev) and then renames the live runtime into it. Just after the store stops, an antivirus scan or the exiting server can hold a file in one of those folders for a moment. Either step gave up on the first IOException or UnauthorizedAccessException, so the host kept the old runtime until the next service start, which for a service can be weeks.

Two related gaps in the same update:

  • If the new runtime failed to extract, the moves that put the old runtime back had no retry either, so the same short lock left no runtime at pgsql and the service exited. The extract failure itself was not logged when a revert move then failed.
  • If an update was interrupted between moving the old runtime aside and finishing the new one's extract (the process killed, or a failed extract whose revert met a lock past its retries), the next start found no pg_ctl.exe and extracted the package as a first run. That moved forward when the extract had failed only for a moment, but never started at all when the package itself was bad, and for a later major-version update it would have put a new-major runtime in front of an old-major store.

The fix

  • The rescue's two steps (the clear and the rename) run under a bounded retry: 5 attempts, waiting 0.5, 1, 2 and 2 seconds (about 5.5 s at most). Only IOException and UnauthorizedAccessException are retried. Each retry logs one Information line. After the last attempt the existing handling runs unchanged: the same Warning, the live runtime untouched, the stamp not written, the update tried again on the next start. The wait honours the start's cancellation token.
  • The extract-failure revert logs why the extract failed, then runs both of its moves under the same retry. These waits ignore cancellation: the revert must finish to leave a bootable runtime, and the original exception is rethrown (a revert move that still fails throws in its place, as before, with the runtime still in pg-runtime-prev).
  • A start that finds the runtime missing after an interrupted update puts the rescued runtime back first (DarlingStoreUpgrade.TryRestoreRescuedRuntimeAsync, called once at the top of DarlingManagedPostgres.EnsureRuntimeAsync), then takes the normal path, which sees the old stamp and retries the update. It acts only when every one of these holds, checked cheapest first: no pgsql\bin\pg_ctl.exe; a rescued runtime of the store's own PostgreSQL major; a readable runtime stamp (the main stamp, or the legacy one only when the main stamp file is missing); no server running on the store; the rescued runtime carries every TimescaleDB version the store records; the shipped package is present and readable; and the stamp differs from the shipped package. The advance writes the stamp only after a good extract, so an interrupted update leaves the old stamp, while a finished one matches the package. A host on 3.8.0 has only the legacy stamp, which never matches a 3.9 package, so an interrupted first 3.9.0 update is covered. A finished host that later receives a newer package also passes the stamp test; the TimescaleDB check stops a rescued runtime that could not open the store, and a rescued runtime that can open it is swapped forward on the same start. A partial pgsql is moved aside and deleted, and the rescued runtime is moved back, both under the retry. In every other case it changes nothing, and the start re-extracts the shipped package as before; so does a restore move that still fails after its retries.

Why retrying is safe

A failed Directory.Move leaves the source in place (a same-volume rename is one operation). The clear only runs once the rescued-runtime guard has decided the folder may be cleared, and clearing is repeatable after a partial pass.

What still falls back, and what it costs

  • A lock that outlasts the budget behaves exactly as before.
  • Cases that can never succeed (a stray file where pg-runtime-prev should be, a runtime folder the service account cannot write, a move across volumes) now wait about 5.5 s before the same deferral, on each start with a pending update.
  • A stop during the rename's wait leaves pg-runtime-prev emptied and the live runtime in place, the same state the old fallback left. A stop during the revert can wait up to about 11 s.
  • No service start deadline applies: the service reports Running before this bootstrap runs.
  • The restore's same-major test runs the rescued binary's version probe, which can take up to the tool timeout (about five minutes) on a hung binary.
  • A host that lost pg_ctl.exe after a finished update (the stamp matches the package), a host with no readable stamp, a rescued runtime that lacks the store's TimescaleDB, or a server running on the store keeps today's first-run extract.
  • Not covered: a major upgrade whose revert of the new runtime also failed leaves a stamp that matches the package, so that host re-extracts the new major as before.
  • A package that is itself bad still fails every start, but now with the store's own runtime in place rather than none.

Tests

Three internal seams on DarlingStoreUpgrade (the wait, the rename and the clear), defaulting to today's calls, let the portable pins stand in for a lock, because a file lock blocks rename and delete on Windows only.

  • The rescue: a rename or a clear that fails twice then works swaps the runtime, with two retry lines and no Warning; one that always fails defers, with the delays exactly 0.5/1/2/2 s; cancellation during a delay propagates and leaves the runtime in place. Two real-lock pins (FileShare.None, released after about a second) run on Windows CI and skip elsewhere.
  • The revert: a failed extract whose restore move, or whose move-aside, is locked twice retries and puts the old runtime back, with the extract failure logged and the original exception rethrown.
  • The restore: an empty, partial or absent pgsql with an interrupted stamp gets the rescued runtime back (the Warning says whether an incomplete runtime was moved aside), and so does the 3.8.0 shape with only the legacy stamp; a rescued runtime that carries the store's TimescaleDB is restored. Nothing moves when there is no store, no rescued copy, a rescued copy of another major, a live pg_ctl.exe, no stamp, an empty main stamp, a server running on the store, a rescued runtime without the store's TimescaleDB, a missing or unreadable package, a stamp that matches the package (in any case), or both stamps with the main one matching. A restore locked twice retries and restores; one that keeps failing logs and returns false. A source pin checks the call sits before the live-runtime check in EnsureRuntimeAsync.
  • Each new gate's pin fails with that gate removed.

CHANGELOG

SECTION: Fixed
ENTRY: - A Darling update that brings a new PostgreSQL runtime no longer waits for the next restart when the old runtime folder is briefly locked ([#4934]) - Just after the store stops, an antivirus scan or the exiting server can hold the runtime folder for a moment, and the update then kept the old runtime until the service next started. It now retries for a few seconds first, and only then falls back as before. A start that finds the runtime missing after an interrupted update now puts back the runtime that last opened the store, and retries the update, instead of extracting the new one as if this were a first run.
REF: [#4934]: #4934

@erikdarlingdata erikdarlingdata changed the title Darling: the runtime rescue retries a briefly locked folder for a few seconds before it defers the update Darling: a runtime update retries a briefly locked folder before it defers, and a start after an interrupted update puts the store's own runtime back Oct 2, 2026
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 2, 2026 05:02
@erikdarlingdata
erikdarlingdata merged commit 1eccade into dev Oct 2, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/runtime-rescue-retry branch October 2, 2026 05:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant