Skip to content

Lite reopens its database after a fatal DuckDB error and repairs its indexes at every open - #4930

Merged
erikdarlingdata merged 10 commits into
devfrom
fix/lite-duckdb-reopen-after-fatal
Oct 2, 2026
Merged

erikdarlingdata merged 10 commits into
devfrom
fix/lite-duckdb-reopen-after-fatal

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

What users see

A Lite install that monitored 9 servers hit a corrupted-index error in its local DuckDB database, less than a second after two out-of-memory errors. That error is FATAL in DuckDB: it invalidates the whole database, and every later statement on every connection fails with "database has been invalidated because of a previous fatal error". Collection and alerting stopped for all 9 servers, and nothing on screen said so. Only a restart of Lite brought them back. The out-of-memory errors are fixed in #4924.

With this change:

  • Lite closes every connection after a fatal error and opens the same file again. The file keeps every committed row.
  • While Lite reopens the file, collection pauses. The reopen waits for collections that are already running to finish, and no new one starts. One log line says that collection is paused, and one says when it resumes. Then collection and alerting carry on.
  • The status bar says that the local database failed and that collection and alerts are stopped while Lite reopens it. The Collection Health tab says the same, and its header reads "Collection Health (stopped)". After a reopen, the status bar shows one line with the time of the reopen.
  • If the reopen fails 3 times, or the database fails again after 5 reopens in one run of Lite, Lite stops reopening it. The status bar then says that collection and alerts are stopped until Lite is restarted.
  • At every open, Lite repairs the damage that duckdb#26106 does to its indexes after a crash or forced close. That damage makes a later delete fail with a FATAL error, which stops everything again.

Cause

There are two parts.

  1. Lite had no path back from a fatal error. Its database instance stays alive while any connection to it is open. Lite keeps one connection open for its whole life (the sentinel). So the invalidated instance lived until the process ended.
  2. duckdb#26106 (open upstream, labelled "reproduced", on 1.5.4 and 1.5.5). After an unclean close, the next open replays the WAL. The next automatic or shutdown checkpoint then loses the replayed rows from the ART indexes on their table. The rows stay in the table, but index lookups miss them, and a later DELETE over them fails with "FATAL Error: ... Failed to delete all rows from index". Lite deletes over indexed rows in its archive step and in its duplicate cleanups.

We do not know what corrupted the index in the incident. Part 1 makes Lite recover from any fatal error. Part 2 removes one known cause.

Changes

Reopen after a fatal error

  • DuckDbInitializer.Recovery.cs (new): ReportFailure(exception) checks DuckDBException.ErrorType == Fatal, through inner and aggregate exceptions. It checks the type, never the message text. Out-of-memory and every other error are ignored.
  • The first fatal error sets the state to Reopening and starts one reopen in the background. It takes no database lock, so a caller that holds the read lock can call it. A second report while a reopen runs starts nothing.
  • Each attempt waits 2 s, 15 s or 60 s. Then it waits until no collection is running, and it keeps new ones from starting until it ends (see "Every connection to the file" below). Then it runs InitializeAsync, the same steps as a start. That closes the sentinel, opens the same file, runs the open steps, and opens a new sentinel. It never takes the reset path, which deletes files.
  • Up to 3 attempts. A wait for running collections is not an attempt and does not count. When all 3 attempts fail, the state is Failed, and nothing starts another reopen until Lite restarts. A reopen that works resets the count, so a later fatal error gets 3 attempts again.
  • At most 5 reopens in one run of Lite (MaxReopenCycles). A fatal error after the fifth sets the state to Failed instead of starting a sixth. So a database that keeps failing stops, instead of reopening forever.
  • If Lite closes while an attempt opens the file, the attempt closes the sentinel its open made, so no connection is left open.
  • Logs: ERROR for the fatal error, INFO for the reopen, and WARNING for each failed attempt. ERROR when all 3 attempts fail, or when a fatal error comes after the fifth reopen. One INFO line when an attempt has to wait for running collections, and a WARNING after each 3 minutes of that wait.
  • Three places report a fatal error:
    • The catch around each collector run.
    • The catch around the collection_log write that ends every collector run. It sees a fatal error even when the collector absorbed its own failed write.
    • The trim timer's memory read on the sentinel, every 60 s. This one covers a fatal error that no collector run sees, for example when every server is paused.
  • The status bar (MainWindow) and the Collection Health tab (ServerTab) read the state from memory, never from the database.

Collection pauses while the database is down

RemoteCollectorService.LocalDatabaseIsDown() reads the state from memory. While the state is Reopening or Failed:

  • A collector run returns before it starts. It records no result and writes no collection_log row, so it stays due and runs in the first sweep after the reopen. Every way into a collector run reaches this check.
  • A sweep (RunDueCollectorsAsync, RunAllCollectorsForServerAsync) returns right after it registers with the collection gate. A sweep that was already running skips its closing CHECKPOINT.
  • The Query Store backfill skips each server.
  • The collection loop skips its whole cycle, so archival, retention, the cleanups, the pileup sweep and analysis wait too.
  • One WARNING line says that collection is paused, with the status line, when the state turns down. One INFO line says that collection resumes when it turns healthy. A skipped run logs nothing more. The read of the state, the compare and the line are one step under a lock, so the lines come in order.

An Azure SQL Database collector reads each database in a loop. That loop now ends at the first fatal error, instead of reading every remaining database and failing each write. The enumerated per-item loop does the same.

Every connection to the file

The first version of this change said that the reopen's write lock waits until every caller closes its connection. That was wrong. A collector run opens its connection to the file without the lock, because it keeps the connection open while it reads from the monitored server. So the write lock did not wait for it, and an attempt opened the file while a collector still held the invalidated instance. DuckDB.NET then hands the attempt that same instance, so the attempt fails and spends one of its 3 tries. P1 shows this on the first version.

A reopen works only if no connection to the invalidated instance is open when the attempt opens the file. Three rules now give that:

  1. Collections register with the collection gate (CollectionResetGate, [BUG] Lite - Table ".main.*" could not be found #2594). These are the scheduled sweep, the sweep when a server tab opens, and the tab's refresh. An attempt takes the gate first. The gate waits until every registered collection ends, which is the point right after a sweep's last collector. It also keeps a new one from starting until the attempt ends. The size-triggered archive and reset already takes the same gate before it deletes the file.
  2. Every other connection to the file is opened under the shared read or write lock. The Query Store backfill now does that too. The attempt's InitializeAsync takes the write lock. That waits for those connections to close. It keeps new ones out until the open is done, the index rebuild included.
  3. While the database is down, collections, the backfill and the loop's housekeeping skip. So nothing new opens a connection to it. A collection that was already running ends at its first failed write.

With no collection running, the gate is free at once, so the attempt runs right after its wait. That covers no servers, every server offline, and collection paused. No collection round is needed to start the attempt.

A search of every place that opens a connection to the file found these:

Where the connection is opened Lock before this update What keeps it closed during an attempt
RunCollectorDefinitionAsync: the single read, the enumerated per-item loop and the Azure SQL Database per-database loop (3 opens) none It runs only inside a registered collection, so the gate waits for it. A new run skips while the database is down.
Query Store backfill: the slice's write batch, its candidate-database read, its earliest-collected-time read and its collector-state delete none Now the read lock, or the write lock for the delete, as SaveCollectorStateAsync already takes. The backfill also skips while the database is down.
Query Store backfill: the orphan and foreign-row prunes none They run only inside collector runs, so the gate covers them.
DuckDbInitializer: the sentinel, the open steps and the archive views read or write lock The write lock waits. InitializeAsync closes the sentinel itself.
RemoteCollectorService: the collection_log writes, collector state, the closing CHECKPOINT and its other opens (9 in all) read or write lock The write lock waits.
Analysis: the fact collectors, drill-downs, anomaly and baseline reads, and FindingStore read or write lock The write lock waits.
The alert and mute stores, Query Store slice repair, LocalDataService (every grid and chart), DeltaCalculator and ArchiveService read or write lock The write lock waits.
DataImportService none It opens a monitor.duckdb in the folder the user imports from, never Lite's own file.
ParquetCompaction none It uses an in-memory DuckDB database.

A connection that leaks anyway keeps the invalidated instance alive. Then each attempt fails, and after 3 the state is Failed. A pin covers that case.

Index repair at every open (duckdb#26106)

DuckDbInitializer.ReplayedIndexes.cs (new) runs in InitializeCoreAsync right after the open. It runs before the migrations, before any DELETE, and before the schema's index statements.

  1. An explicit CHECKPOINT. This is the upstream workaround: a checkpoint right after the replay writes the replayed rows' index entries correctly. If it fails with an ordinary error, Lite logs an ERROR, does not rebuild, and opens.
  2. Each explicit index in duckdb_indexes() is dropped and created again from its own sql. A rebuilt index holds every row of its table, so this repairs an index that an earlier open already damaged. If a step fails with an ordinary error, Lite logs an ERROR and opens.
  3. One INFO line gives the count and the elapsed time.

A FATAL error in either step is different. It invalidates the database, so the open cannot carry on. The step throws an error that names it, for example "The CHECKPOINT Lite runs right after opening ... failed with a fatal error, so the database cannot be used". A start then fails with that message instead of the next statement's "database has been invalidated", and a reopen counts it as a failed attempt. The first version of this change logged a fatal CHECKPOINT and carried on.

Both steps run at every open, a clean one included. Damage from an earlier session cannot be told from a healthy index without reading it.

The cost has a bound. When the file reaches 512 MB, the collection loop archives every table to parquet and resets the file. So the file stays near that size or below it. On the 402 MB copy of the incident's store below, the rebuild took 622 to 637 ms for 45 indexes. So the rebuild stays at about a second at most.

Each DROP and each CREATE commits on its own. A transaction that drops and re-creates an index loaded from the file fails at COMMIT when the commit is written to the WAL. DuckDB.NET dies with a native access violation that no catch can stop. Lite's checkpoint_threshold=1GB sends every such commit to the WAL, so each DROP and CREATE commits on its own. The first version of this change used one transaction, and its own tests crashed the test host on a plain reopen. A separate probe, run once per case, gave these results:

Case One transaction Separate commits
Clean file, CHECKPOINT first crash at COMMIT works
Clean file, no CHECKPOINT crash at COMMIT not run
Clean file, index used by a lookup first crash at COMMIT not run
Clean file, one transaction for each index crash at COMMIT not run
A write, then CHECKPOINT, then the rebuild works not run
WAL replay, then CHECKPOINT works works
WAL replay, then a shutdown checkpoint (damaged index) not run works, and deletes work after it
WAL replay, no CHECKPOINT crash at COMMIT works

The Python client of DuckDB 1.5.5 fails at the same COMMIT with an internal error. So the failure is in DuckDB, not in the .NET binding.

With separate commits, an index is missing between its DROP and its CREATE. Nothing else uses the file meanwhile. The open holds the write lock, which every connection outside a collection takes. At startup no collection is running yet, and a reopen also holds the collection gate. If a CREATE fails, that index stays dropped:

  • The schema's CREATE INDEX IF NOT EXISTS statements run later in the same open, so they try again for any index that Lite declares.
  • An index that Lite no longer declares stays dropped. That costs speed only. A missing index cannot be inconsistent, and Lite has no unique index.
  • The ERROR line gives the full CREATE INDEX statement of the dropped index, so someone can restore it by hand.

The incident's store has 45 explicit indexes, and Lite declares all 45: 40 generated for collectors, 2 hand-written, and 3 for analysis. All migration DROP INDEX statements use IF EXISTS, so a missing index cannot break a migration.

A declared index that cannot be created

Before this change, the schema's index statements always found their index in place and did nothing. Now an index that the repair dropped and could not create again reaches them. If the cause lasts, for example an index too large to build within the 1 GB memory limit, the schema statement fails the same way. It used to throw, so Lite could not start, and every later start and every reopen attempt failed at the same statement.

Both index loops, the schema's and the analysis schema's, now use the pattern of the #4727 missing-column heal. On an existing file, a declared index that cannot be created logs one ERROR with its statement, and the start carries on without it. The next start tries again. A fresh file still throws, because a declared index that cannot be built on an empty table is a bug for the tests to catch. A FATAL error still throws too, because the database is invalidated.

An existing file means that the file was there before the open. Lite checks that with File.Exists before it opens the file. The first version used the schema version instead. But a failed read of the version reads as 0. So an existing file whose version read failed was treated as a fresh file, and its first index that failed to build stopped the start.

Remove both steps when Lite ships a DuckDB release that fixes duckdb#26106 for both the shutdown checkpoint and the automatic checkpoint. The 2.0 nightly fixes only the shutdown path. The code comment says the same.

Other

  • Schema.cs: two doc comments gave stale counts ("36 collector tables", "34 generated collector indexes"). The catalog now has 42 collectors and 40 generated indexes. The comments now say "one per catalog collector" so that they cannot go out of date again. No code changed.
  • LocalDatabaseHealth.LocalTime has a block body. TsqlConventionGuardTests.TheMemberScan_ReadsEveryDeclarationWhole reads a multi-line expression-bodied member short of its end, and it failed on the first version of this method. With a block body the scan reads the method whole, so KnownTruncatedRanges does not change.

On a copy of the incident's store

The copy was 402 MB with a 1.5 MB WAL, 62 tables, about 996,000 rows and 45 explicit indexes. It used Lite's memory_limit=1GB and checkpoint_threshold=1GB and the separate commits. The copies were deleted afterwards. These are timings and counts only.

The timing run opened one copy twice:

Step Open after the crash (WAL replay) Next open (no WAL)
Open 673 ms 78 ms
CHECKPOINT 1,275 ms 0 ms
Rebuild of 45 indexes 622 ms (slowest: 203 ms, on a table of about 582,000 rows) 637 ms

The peak working set of the process over both opens was 352 MB. So the repair adds about 0.6 s to a normal start and about 1.9 s to the first start after a crash.

Full deletes of every indexed table, each case on a fresh copy:

Open before the deletes Result
The open and close that every earlier build does FATAL on the third indexed table: "Only deleted 205 out of 210 rows"
This change's open 44 of 44 tables, 989,294 rows deleted, no error
CHECKPOINT only, no rebuild 44 of 44 tables, no error

On this store, the CHECKPOINT alone was enough. The rebuild is for a store that an earlier build already opened and closed after a crash.

Changes after the second check

A second check of this change found nothing that blocks a merge, and two small gaps in what the log says. Both are closed here. Each has a pin, and a mutation that breaks the pin.

  • A long wait for collections now warns. An attempt waits for running collections with no time limit. An attempt fails while a collection holds its connection. A collection in a remote read with a long command timeout keeps the database down for that whole read. Before, one INFO line said so. Now, once the wait passes the collection gate's 3-minute drain timeout, a WARNING gives the minutes waited and the count of collections still running. It repeats after each further 3 minutes. The attempt keeps waiting, the wait still does not count as an attempt, and closing Lite still ends it. The interval is ReopenGateWarningInterval, which a test can shorten. The gate's DrainTimeout is now internal, so both use one value.
  • The paused and resumes lines stay in order. LocalDatabaseIsDown read the state, then swapped it into the last logged state. In a race, one caller read Reopening just before a reopen ended. Another caller then logged "resumes". The first caller swapped Reopening back in and logged "paused", and the next caller logged "resumes" a second time. P6 shows this without the lock. Now the read, the compare and the line run under one lock. A re-check after the swap is not enough. With it, the stale caller still puts Reopening back, only without a line, and the next caller still logs "resumes" again.

The status text and the CHANGELOG entry do not change. Only log lines do.

Pins

LocalDatabaseReopenTests forces a real fatal error. It sets DuckDB's debug_checkpoint_abort, and then a CHECKPOINT fails with FATAL and invalidates the instance.

Pin What it shows
R1 AFatalErrorInACollectorRun_ReopensTheDatabase_AndTheNextRunWrites A collector run on the invalidated database starts the reopen through its collection_log write. The state is Healthy with a reopen time. debug_checkpoint_abort reads NONE, so the instance is new. The next run writes its collection_log row, and a table committed before the error is still there.
R2 WhenEveryReopenFails_LiteStopsAfterThreeAttempts_AndSaysCollectionIsStopped A connection that stays open makes every attempt fail. Exactly 3 attempts run, the state is Failed, and the status line says collection is stopped until Lite restarts. A later fatal error starts nothing.
R3 AnErrorThatIsNotFatal_StartsNoReopen_AndAWrappedFatalErrorIsStillSeen A non-fatal DuckDB error and a non-DuckDB error start nothing. A fatal error inside another exception or an aggregate is still seen.
R4 TheStatusLine_SaysWhatHappenedAndWhen The exact status lines, with their times.
R5 TheSentinelProbe_StartsTheReopen_WhenNoCollectorReportsTheError The trim timer's sentinel read starts the reopen when no collector runs.
P1 AFatalErrorWhileACollectorReads_TheReopenWaitsForThatRun_ThenSucceeds An Azure SQL Database collector run registers the way every collection does. It holds its connection to the file while the test holds its read of the first database open. The fatal error comes then. For 1 s no attempt runs and the state stays Reopening. When the read ends, the run's write fails and the run ends. The reopen then works on its first attempt: the state is Healthy, not Failed, and the row committed before the error is still there.
P2 TheQueryStoreBackfill_SkipsWhileTheDatabaseIsDown_AndSaysSoOnce While a running collection keeps the database down, two backfill ticks read nothing from the file and log no failed read. Exactly one line says that collection is paused. The reopen works once the collection ends.
P3 AFatalErrorDuringAPerDatabaseRun_EndsTheLoop_SoNoLaterDatabaseIsRead The fatal error comes while the first of 3 databases is read. The run ends with the fatal error, and only the first database was read.
P4 ACollectorRun_SkipsWhileTheDatabaseIsDown_AndOneLineSaysSoEachWay While the database is down, two collector runs record no result and try no collection_log write, and exactly one line says that collection is paused. After the reopen the next run records its result and writes its row, and exactly one line says that collection resumes.
F1 WithNoCollectionRunning_TheReopenStillRuns (3 cases) With no servers, with every server offline, and with collection paused, no collection round ends to start an attempt. The reopen still works on its first attempt in each case.
H3 AfterFiveReopensInOneRun_TheNextFatalError_LeavesTheDatabaseFailed 5 fatal errors in a row each reopen the database. The sixth starts no reopen: the state is Failed, and the status line says Lite stopped reopening it.
H4 ADisposeWhileAnAttemptRuns_LeavesNoSentinelOpen Lite closes while an attempt opens the file. After the attempt, no sentinel is open, and the state is not Healthy.
P5 AReopenThatWaitsLongForACollection_WarnsWithTheRunningCount_AndKeepsWaiting A collection stays registered, and the warning interval is cut to 50 ms. The log gets at least 2 WARNING lines that give 1 collection still running. Meanwhile no attempt runs and the state stays Reopening. When the collection ends, the reopen works on its first attempt.
P6 ACallerHoldingAReadFromBeforeTheReopen_CannotLogItBackwards_SoResumesIsLoggedOnce Caller A reads the state while the database is down and stops at a test hook after the read. The reopen ends. Caller B cannot finish while A waits. Then A is released and caller C runs. The log has exactly one "paused" line and then exactly one "resumes" line.

ReplayedIndexRebuildTests makes a real unclean close. The rows are committed to the WAL, and the last connection closes with disable_checkpoint_on_shutdown.

Pin What it shows
I1 RowsReplayedFromTheWal_StayInTheirIndex_ThroughLitesOpenAndAShutdown Lite opens the file after the unclean close and closes it normally. An index lookup finds the replayed rows, and a delete over them works.
I2 AnIndexAnEarlierShutdownDamaged_IsRepairedWhenLiteOpens Another open replays the WAL and closes first, which damages the index. Lite's open repairs it, and the delete works.
I3 TheRebuild_RunsAtEveryOpen_AndLeavesEveryDefinitionAsItWas Two opens in a row. Each logs the rebuild of every index and no error, and every index definition is the same as before.
I4 TheCheckpoint_GatesTheRebuild_AndNoIndexIsDroppedAndCreatedInOneTransaction Source order: the CHECKPOINT comes before any DROP, and a CHECKPOINT that fails with an ordinary error returns first. The rebuild never opens a transaction. The call sits in InitializeCoreAsync before the migrations and before the schema's index statements.
I5 ADeclaredIndexThatCannotBeBuilt_IsLogged_AndTheOpenCarriesOn On an existing file, one declared index in each loop cannot be created: its column's type is changed to one that an index cannot hold. The open completes, with exactly one ERROR for each index, and a collector's rows still land.
H5 AnExistingFileWhoseSchemaVersionCannotBeRead_StillOpens_PastAnIndexThatCannotBeBuilt The same kind of index that cannot be built, on an existing file whose schema_version table is gone, so its version reads as 0. The open still completes, with exactly one ERROR naming that index, and 500 new rows land after it.
H2 AFatalCheckpointAtTheOpen_StopsTheOpen_WithAnErrorThatNamesTheCheckpoint The CHECKPOINT right after the open fails with FATAL. The open throws, its message names the CHECKPOINT, and the error is still seen as fatal.

Red proofs

On the base commit, with only ReplayedIndexRebuildTests.cs copied in, I1 to I4 fail. I1 and I2 fail with "FATAL Error: Invalid Input Error: Failed to delete all rows from index. Only deleted 0 out of 500 rows." I3 finds no rebuild line, and I4 finds no repair file. I5 came later, so mutations prove it.

The pins this update adds were run against the production code of the first version, with this update's tests copied in. P1 fails because all 3 attempts are spent within the first second, while the collector still holds its connection. P3 fails because the loop reads all 3 databases. P2 fails because the backfill reads the file and logs failed reads of its collector state and its candidate databases. H5 fails because the start stops at "Invalid type for index key".

F1 passes there, as expected: with no collection running, the first version had nothing to wait for. R2 and R4 fail there only on the new status text. H2, H3 and H4 need test hooks that this update adds, and P4 came later, so mutations prove them.

Each mutation below was built and run against the pins it must break, on this update's code. Then the file was restored. Every mutation broke the pins it must break.

Mutation Pins that failed
A fatal error is not recognised R1, R2, R3, R5, and every other pin in LocalDatabaseReopenTests but R4
No report from the collection_log write R1
No report from the sentinel read R5
The reopen leaves the old sentinel open R1
2 more attempts R2
No Failed state R2
Failed does not stop new reopens R2
Any DuckDB error starts a reopen R3
The status line has no time R4
The rebuild is skipped I2, I3
The CHECKPOINT is removed I4
The rebuild runs in one transaction (only I4 was run, because the other pins would crash the host) I4
The call moves after the schema's index statements I4
The schema index loop throws again, as before this change I5
The analysis index loop throws again I5
A failed declared index logs twice I5
An attempt does not wait for running collections P1, P2, P4
The per-database loop carries on after a fatal error P3
The backfill does not skip P2
A collector run does not skip P4
Every skipped run logs a line P2, P4
No limit on reopens in one run H3
A close during an attempt is not seen H4
A fatal CHECKPOINT at the open is logged and the open carries on H2
An existing file is told by its schema version H5
An attempt waits for a collection that never comes F1, all 3 cases
A long wait for collections never warns P5
The paused and resumes compare runs without the lock (the test hook kept) P6

Tests

  • The two pin classes: 23 of 23 pass. That is LocalDatabaseReopenTests 16, counting F1's 3 cases, and ReplayedIndexRebuildTests 7.
  • Lite.Tests full suite at the head: 7,142 tests, 0 failed, 0 skipped.
  • Darling.Tests full suite at the head: 19,773 tests, 0 failed, 1,298 skipped (the tests that need a live PostgreSQL), 1 not run.
  • Build: 0 warnings, 0 errors.

Not pinned: the enumerated per-item loop's rethrow of a fatal error. No test hook drives that loop without a SQL Server. It follows the same rule as the per-database loop, which P3 pins.

Noticed, not changed: Schema.CreateServerTagMapIndex is defined but never run, so idx_server_tag_map_tag is never created.

CHANGELOG

The changelog line is written from this entry at release, so CHANGELOG.md is not edited here.

SECTION: Fixed
ENTRY: After a crash or forced close, Lite could later stop collecting and alerting for every server until it was restarted; it now repairs its indexes at start, and after a fatal error it pauses collection while it reopens its database, then carries on ([#4930], [#4929]).
REF: [#4930]: #4930
REF: [#4929]: #4929

…indexes at every open

A fatal DuckDB error invalidates the whole database: every later statement fails until every
connection closes and the file opens again, so collection and alerting stopped for every server
until Lite was restarted. Lite now reopens the same file under the write lock (3 attempts with
back-off), says so in the status bar and on Collection Health, and logs the fatal error and the
reopen.

At every open, an explicit CHECKPOINT and then a rebuild of each explicit index (each DROP and
CREATE committed on its own) keep rows that WAL replay restored in their indexes (duckdb#26106),
so a later delete over them no longer fails with a FATAL error.
The index repair at the open can drop an index that then cannot be built again, for example one
too large to build within the memory limit. The schema's CREATE INDEX IF NOT EXISTS statements then
met the same failure and stopped the start, and every start and reopen attempt after it. On an
existing file, a declared index that cannot be created now logs one ERROR and the start carries on,
so the next start tries again (the same pattern as the missing-column heal). A fresh file still
fails, because there a declared index that cannot be built is a bug.

Also: the comment no longer names a server removal as a delete over indexed rows (it deletes none).
TsqlConventionGuardTests.TheMemberScan_ReadsEveryDeclarationWhole reads a multi-line expression-bodied member short of its end. The same expression in a block body reads whole, so no entry is added to KnownTruncatedRanges.
…es collection while it is down

A reopen attempt takes the collection gate the size-triggered reset uses, so it runs after every registered collection has ended and none starts until it is done. Collector runs, the Query Store backfill and the loop's housekeeping skip while the database is down, a per-database loop ends at the first fatal error, and the backfill's connections take the database lock. A fatal error in the open's CHECKPOINT, index rebuild or declared index statements stops the open with an error that names the step. One run reopens at most five times. A dispose during an attempt closes the sentinel the attempt opened. An existing file is told by whether it was there before the open.
…d and resumes lines are written in order

Past the collection gate's 3-minute drain timeout, the wait for running collections logs a Warning with their count, and again after each further 3 minutes. It keeps waiting. The health read, the compare and the paused or resumes line in LocalDatabaseIsDown are one step under a lock, so a caller holding a read from before a reopen cannot log the change backwards.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review October 2, 2026 03:04
@erikdarlingdata
erikdarlingdata merged commit ca9a9a8 into dev Oct 2, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/lite-duckdb-reopen-after-fatal branch October 2, 2026 03:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant