Bring cloud sync and the web beta into main - #234
SunkenInTime wants to merge 350 commits into
Conversation
…tics Add stable cloud beta automation hooks
Clarify share link disable behavior
Refresh the Icarus Online beta readiness gate
Give cloud library states an obvious next step
Refactor cloud sync around server-side page boundaries
With icarus-cloud merged, main is the one branch. Deploy Web now runs on pushes to main (and refuses any other ref), CI no longer watches icarus-cloud, and the contract gate records its merge base with origin/main, so it keeps working once icarus-cloud is gone. The release doc says so, and the typed-wrapper ADR no longer reads as if clearing an outbox were acceptable: cloud has had real users since 2026-09-27. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Live sync compares whole payloads, so this pins the property the cloud upgrade rests on: an upgraded row and the canvas encode identically. Without the version 2 stamp on what the client writes, opening the page would author a rewrite of every Paranoia on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The runner closed every new launch as soon as it found an Icarus window, but app_links only forwards a lone URI with a scheme. Double-clicking an .ica file while Icarus was open did nothing, and a launch with its own --hive-store-dir could not start beside another Icarus. Only an icarus:// link is handed to the running window natively now; every other launch reaches Dart's per-store single-instance check, as it did in 4.6.3. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#222 moved Settings.versionNumber to 104 for the Paranoia migration and left pubspec at 4.6.3+103. bump_version.ps1 refuses to run on a mismatch, so Release Desktop would have stopped before building. Both read 104 now; the release bump takes them to 105 together. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Important Review skippedToo many files! This PR contains 732 files, which is 632 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. Usage-priced reviews support at most 300 files. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: ⛔ Files ignored due to path filters (73)
📒 Files selected for processing (732)
You can disable this status message by setting the
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…opened Marking ops present when the queue activated a strategy missed two paths. A trip to a local strategy and back redraws the canvas from the server, but the queue never switches, so work sent before the trip landed as the canvas's own. And a successor promoted in the background while the user opened its strategy was stored after activation looked, so its new op ID went unmarked. The queue now remembers the reverse: op IDs it made from the canvas's desired ops, plus the ops it resends in their place. Live sync tells it to forget them whenever it drops its bases and overlays (a strategy opened or left, local ones included, or a reset), and ops asked for before that are not counted. Every other ack is marked restored, so a background promotion needs no bookkeeping at all. Remembered IDs are pruned once their records are gone, inside the queue's serialized writes. The new regressions drive the real queue: a cloud, local, cloud round trip through the page session with a drop over the landed edit, and a gated store that opens the strategy mid-promotion. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cloud bootstrap ran inline before runApp: a throwing outbox open, outbox prepare, Convex init or Supabase init left the app unshown, so a signed-out desktop user could not reach a healthy local library. startCloud now runs those steps, never throws, and on the first failure leaves the cloud off for the run: auth sees no session and every sign-in answers with the reason, the outbox stores read and write nothing, and a toast says so once the app is up. The release-build check for ICARUS_CLOUD_ENVIRONMENT still stops startup. The outbox boxes now open with Hive's crash recovery off. Recovery truncates a box at its first bad frame, silently deleting the unsent work after it; a corrupt outbox now stays on disk byte for byte and the cloud stays off. A signed-out startup also no longer clears Convex auth it never set, which opened the Convex socket for users who never asked for the cloud. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The save button awaited forceSaveNow and picked its toast from media state alone. In cloud mode a failed durable outbox write is recorded on the queue and returned normally, so the toast said Saved for work that was not even on this device; a local save that threw left no toast, and one that skipped a strategy missing from the library said Saved. The toast now reads where the work is after the save. Local: Saved only if the library was written after the press. Cloud: not saved when an outbox write failed, otherwise the sync button's own states and words (synced, syncing, offline, needs attention). Failures use the destructive toast. A failed outbox write also reads as unverified in the sync popover instead of merely unsent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The common bad frame in a box is the last one, torn by a crash or kill in the middle of a write. Recovery trims only that write, which never completed and was never reported saved. With recovery off, that ordinary crash left cloud sync off for good with no way back, worse than the rare mid-file corruption it guarded against. The startup test now checks that a torn final frame is trimmed and the ops before it survive. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Unsent changes from one account could land on the server as another account signed in on the same device. The queue checked the account before claiming a batch, then awaited the claim's durable writes and sent without checking again, so a sign-out and sign-in in that window sent the first account's batch as the second. The transport also keeps a pending request across a sign-out and resends it after the next sign-in. Either way, if the new account could edit the strategy and the revision still matched, the work landed as theirs. applyBatch now takes an optional accountSubject. When it is given and is not the signed-in identity's subject, the server refuses the whole batch with FORBIDDEN before any op, so nothing is applied or recorded and the same op ids land later for the right account. The queue always sends it, taken from the outbox records' account, never from whoever is signed in at send time. Older clients send no subject and keep their behavior. After the claim, the queue checks the account again. If it changed, signed out, or has an auth incident, the claimed records go back to their own account's queue with the same op ids and no attempt counted, to send when that account is signed in again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two fakes written alongside the account binding didn't know about it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Drawings wrote a Color into field 2 for 3.2.3 and the ARGB value into a mirror field. 4.6.3 writes the ARGB value into field 2 and reads either, so nothing was lost, but the rollback test's byte-for-byte claim compared this build's drawing writers against themselves. Drawings now write 4.6.3's record too, the rollback registry swaps in 4.6.3's drawing adapters, and the color test is the one 4.6.3 ships. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Live sync passes recovered work it keeps through the same desired-op map as new edits. If that work landed while the page write waited its turn, the queue found no record, wrote it again as the canvas's, and its replayed ack moved the held item's base: a drop then replaced the recovered edit. A desired op the queue already held when asked is now only a request to keep it: never marked as the canvas's, and not written again once it has landed or been replaced. An ack is now judged when it is published rather than before its records are stored, so a canvas drawn fresh in between gets it as restored instead of moving its new base back. The discard paths and attention retries prune remembered canvas op IDs too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| export async function assetReferencesReady(ctx: AnyCtx): Promise<boolean> { | ||
| return ( | ||
| (await isBackfilled(ctx, "elements")) && | ||
| (await isBackfilled(ctx, "lineups")) | ||
| ); |
There was a problem hiding this comment.
Existing content has no asset-reference rows, but this readiness check cannot pass until both tables are backfilled. The backfill schedules further pages only after its first invocation, and neither production deployment path starts it. Existing images can be omitted from the image list, reference queries cannot answer, and image reclaim remains paused. This violates the repository directive to backfill existing production rows when changing the schema; it must be resolved before merging.
Context Used: AGENTS.md (source)
Knowledge Base Used: Strategy persistence and integrity
Artifacts
Executed Convex persistence test source
- This is the exact authored test source run against mock persisted rows for both conditions; it shows how the comparison was made.
Backfill scheduling and deploy-path inspection
- The captured repository search and numbered cron and workflow excerpts show continuations but no initial backfill trigger in the inspected paths.
Persistence check without an initial backfill
- The passing test ran maintenance and reclaim against old-style rows without seeding the backfill; markers and references stayed empty and reclaim checked nothing.
Persistence check after explicit backfill invocation
- The passing test explicitly ran both backfill tables and reclaim; markers and references appeared, the old image became queryable, and an orphan was marked deleted.
| export function assertSupportedCloudProtocol(clientProtocolVersion: number): void { | ||
| if (clientProtocolVersion !== CURRENT_CLOUD_PROTOCOL_VERSION) { | ||
| throw clientUpgradeRequiredError(); | ||
| } |
There was a problem hiding this comment.
Keep shipped clients connected
The previously shipped client sends protocol 3, but this check accepts only protocol 4. Deploying the backend first rejects cloud writes from the still-live web build, while installed desktop builds remain affected until updated. This violates the repository directive to keep server changes compatible with the live web build and shipped desktop builds, gating new strictness behind client opt-in; it must be resolved before merging.
Context Used: AGENTS.md (source)
Artifacts
Executed cloud service-contract test source
- This authored Vitest source calls the real Convex handlers with v3 and v4 requests; it is the source used for both captures.
Cloud mutation results before the protocol change
- The contract test ran against revision 6239072 and captured three successful v3 writes and three rejected v4 writes; v3 remained usable.
Cloud mutation results after the protocol change
- The same test ran against revision 7083fe9 and captured three v3 upgrade-required errors and three successful v4 writes; v3 cloud writes are blocked.
Previous client version and production deploy order
- Executed source-inspection commands captured the prior v3 client constant and backend-before-web workflow; an existing v3 client precedes the v4 rollout.
| const olderActiveAssets = await ctx.db | ||
| .query("imageAssets") | ||
| .withIndex("by_strategyId_and_publicId_and_uploadStatus", (q) => | ||
| q | ||
| .eq("strategyId", strategy._id) | ||
| .eq("publicId", args.assetPublicId) | ||
| .eq("uploadStatus", "active"), | ||
| ) | ||
| .take(20); | ||
| let replaced = 0; | ||
| for (const olderAsset of olderActiveAssets) { | ||
| if (olderAsset._id === asset._id) { | ||
| continue; | ||
| } | ||
| await markImageAssetDeleted(ctx, olderAsset, now); |
There was a problem hiding this comment.
Older upload replaces newer image
When two attempts share an image ID and the newer upload completes first, completing the older attempt makes its stale bytes visible and marks the newer image deleted. A later sweep removes the newer upload's stored object. Completion must recognize a superseded attempt rather than treating every other active asset as older; this image-loss path must be fixed before merging.
Knowledge Base Used: Strategy persistence and integrity
Artifacts
Executed Convex upload-order test source
- The authored test invokes the real mutation and deletion sweep with a mock DB and mocked R2 DELETE, showing exactly how both states were exercised.
Before older completion: newer image active
- The passing test captured DB rows and the image URL after completing the newer attempt first, showing the newer object is active.
After older completion: newer image deleted
- The passing test captured the reversed active image and a later mocked R2 DELETE for the newer object, confirming the loss path.
A 13 MB .mp4 and bun init's index.ts ("Hello via Bun!") and bun.lock
came in with icarus-cloud. Nothing uses them; CI and the Convex tooling
run on npm.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Protocol 4 was meant to stop protocol 3 clients from reading as well as writing, but only the writes checked it. The page snapshot, the shell, the full snapshot and the element and lineup lists took no protocol, so a web tab opened before the deploy kept receiving Paranoia rows at payload version 2, read them as the old size, and authored version 1 writes against them. Those writes were refused and kept, and after a reload they replayed at protocol 4 and landed: the double move the bump was for. Each of those seven queries now requires clientProtocolVersion and refuses anything but the current one with CLIENT_UPGRADE_REQUIRED, before reading anything. A client that sends no protocol fails argument validation, which is how the live web build fails these reads; it already fails safe on a refused read and keeps its unsent edits. Folder and strategy listings carry no payloads and are left alone. The client sends its protocol on every one of these reads, and the contract, generated client and gauntlet follow. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When the server refused a build's protocol, the app showed the generic "Retry sync", and the queue counted each refusal as the op's failure: after eight it paused the work, and a reload did not resume it. A refusal of the build is never the user's fault or the op's, and no retry can fix it. The first call refused with CLIENT_UPGRADE_REQUIRED, from the op queue, the media queue, the editor's reads, or the account setup, now marks the build as refused for the rest of the run. Sync shows attention, and its popover says "Icarus was updated. Reload to keep syncing." with a Reload button on the web, or "Icarus was updated. Install the update to keep syncing." on desktop with the waiting update's button, the one the strip's update icon opens. The library's cloud error says the same and offers the same button in place of Retry. The queue holds its work while the build is refused, sends nothing, and counts no failure, so nothing pauses or needs attention for it. Work an older build paused or counted for the same refusal is put back in the queue, uncounted, when the app starts. Tests cover ten refused starts in a row and a reload over work the protocol 3 web build paused: nothing is paused or dropped, and the resumed work is sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review of the upgrade face found it could throw work away, lie in the library, strand desktop users, and match refusals too loosely. Reload no longer discards work the outbox could not keep. When an outbox write failed or a record cannot be read, that failure shows, with its own recovery, in place of the upgrade message: it may be the only copy. Reload itself commits drafts, saves the open strategy into the outbox, and reloads only once both outboxes say they hold everything; otherwise it keeps the tab and says so. The library's status card no longer says work is syncing while the server refuses this build. It says the work is kept on this device and waiting for a newer Icarus, with the same way out. A refusal now runs the update check again. On desktop with no update yet, Icarus says so and that the work is safe, and offers Check again, which looks for the update and asks the server, through health:ping with the client's protocol, whether it accepts this build again. If a rolled-back deploy does, the refusal clears: the queue sends what it held, the editor reads again, and a refused account setup runs again. health:ping only checks a protocol it is given, so the connection health check and older clients are unaffected. A refusal is recognized by its typed error code. Kept text counts only when it is exactly what this build writes or what a refused send left in the outbox, here and in the protocol 3 web build; text that merely mentions the code is not a refusal and is not resumed at startup. Disposing the auth provider now makes any setup still running stale, so it finishes without touching other providers. Reading the refusal flag from a disposed container was what failed two auth tests after they had completed. Inline button icons are 16px. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Reload button checked its own idea of "saved" before reloading, and review found two ways it could still throw work away: changes live sync withholds without marking the outbox unreliable, and an edit whose outbox write starts after the check. Leaving a strategy for Home has the same problem and its guard already solves it: it commits drafts, saves, waits for sync, and when it can't confirm the work is kept on this device it asks, offering Leave anyway only for work already in the outbox. Reload now goes through that guard, and the bespoke check is gone. Also: a Convex setup queued behind a running one no longer runs after the auth provider is disposed, and the desktop copy promises only that saved work stays on this device. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The exit guard Home and Reload share decided whether "Leave anyway" was safe before showing its dialog, and honoured that answer afterwards even if an outbox write had failed in between. It also left without asking once sync settled, even if a draft or edit had appeared during the wait. Leave anyway now leaves only if the work is still all in the durable outboxes when the choice is made, and the quiet exit requires nothing unstaged. Edits live sync withholds because a lineup lost its origin or landing still don't stop an exit: they're flagged for attention already, and blocking on them would trap the user on that strategy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A cold Windows runner's first launch of 4.6.3 took longer than 12 seconds to show its window, failing the gate before the candidate was even installed. The same gate runs before every desktop release, so a flake there blocks shipping. Each launch still has to stay alive the whole time and respond. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A newer push cancelled a running Deploy Web, possibly after it had deployed Convex but before Pages published the matching web build. If the replacement then failed early, production served the new server to the old web build. Runs now wait their turn; the queued one publishes the newest build when it runs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Thanks. Going through the four:
|
The account row read only the auth status, so a refused build showed "Cloud connection needs attention" with a Reconnect button that can't help. It now says cloud sync needs a newer Icarus and offers the same Reload, Update or Check again as the sync button. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Its fixtures were stamped when the file loaded, while the dialog reads the clock when it builds, so a slow full run drifted across an "hours ago" boundary and failed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This merges
icarus-cloudintomain, somainbecomes the only branch. It brings accounts, Convex sync, sharing, presence, page trash and the web beta, plus every fix found while getting it ready. Signed-out desktop users keep local mode exactly as 4.6.3 has it.Merge with a merge commit, not a squash, so
icarus-cloud's history stays intact. Merging deploys. Deploy Web now runs frommain: production Convex first, then the web beta.Old web tabs reload after this deploy (decided)
Protocol 4 refuses web tabs opened before the deploy (cbbdadf, 75b58e1). Cloud Paranoias are now stored at payload version 2. A protocol-3 tab would read one at the old size and, on its next edit to that page, write it back so new clients correct it a second time. Protocol 4 stops old tabs reading or writing until they reload. Their unsent edits stay on the device and send after the reload. A reload during the few minutes between the server deploy and the web deploy still gets the old build.
This breaks AGENTS.md's rule that a server change keeps working with the live web build. Dara approved it on 2026-10-01: an update means a reload, as in other apps. From this build on, a refusal says so on screen (6104d28): a reload prompt on web, an update prompt on desktop, and nothing discarded.
What changed beyond the merge
Bringing main across
StrategyMigrator. The folder grid and empty-state changes are ported to the cloud library. Defense resize was already on cloud.+104to matchSettings.versionNumber.bump_version.ps1refuses a mismatch, so Release Desktop would have failed at its first step.Rollback to 4.6.3 keeps the library whole
Sync integrity
applyBatchcarries the originating account, and the queue rechecks the account after claiming.Desktop
.icawhile Icarus is open opens it again.Wiring and cleanup
main. Web deploys now queue instead of cancelling each other, so a cancelled run can't leave production on a new server with the old web build.bun initstub that came in with icarus-cloud.Verification
windows-build, which installs the signed 4.6.3 release, upgrades to this build and rolls back..icaimports into the running window; folders line up with the strategy columns at two widths; the editor opens.Known and left
applyBatch(pages:add, folder ops, media) aren't bound to an account yet.TODO.mdat the root is July's cloud UX work list. Delete it if it's done.🤖 Generated with Claude Code
Not ready to merge: three outstanding issues still require fixes.
Findings
Summary
This PR brings cloud sync and the web beta into main. The latest change stabilizes a recently deleted dialog test. No new findings were identified, but three previously reported issues remain outstanding.
Reviews (6) · Last reviewed commit: "Pin the clock in the Recently deleted li..."