From d1505bb45d4461579e7228fbfcaabff7874bc5bb Mon Sep 17 00:00:00 2001 From: Ben Rusholme Date: Sun, 4 Oct 2026 22:40:51 -0700 Subject: [PATCH] runs: a reclaimed unit's further attempts go to the reclaim queue One sentence under "Attempts and retries" for the behaviour rapid's rebuild gained in #225 (d6901dc5): once a unit loses an attempt to a Spot reclaim, every further attempt goes to RAPIDPIPE_BATCH_RECLAIM_QUEUE when it is set and differs from the job queue; otherwise the job queue. Co-Authored-By: Claude Fable 5.1 --- system/runs.md | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/system/runs.md b/system/runs.md index f02b257..49753ca 100644 --- a/system/runs.md +++ b/system/runs.md @@ -206,7 +206,12 @@ terminal unit or replace its selection. The run records a maximum attempts per unit, counting the first try and every retry. The launcher retries; a Batch job runs its container once. An attempt ending `transient` returns the unit to ready for -submission as a new attempt. Only code 75 and the approved +submission as a new attempt. Once one of a unit's attempts is lost to +a Spot reclaim, in the run or in a run it was seeded from, every +further attempt of that unit is submitted to the reclaim queue +(`RAPIDPIPE_BATCH_RECLAIM_QUEUE`) when that is set and differs from +the job queue; otherwise it goes to the job queue like any other +retry. Only code 75 and the approved infrastructure failures listed in the [stage contract](stage-contract), "Exit codes", are recorded `transient`. Exhausting the allowance fails the unit. `lost` means the scheduler lost the job: unresolved