You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/ARCHITECTURE.md
+15-13Lines changed: 15 additions & 13 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -302,9 +302,9 @@ looks the sandbox up in the Redis routing record to find the owning node, and re
302
302
that node's orchestrator proxy on :5007 by default. `ORCHESTRATOR_PROXY_PORT` selects a
303
303
different downstream port when the node proxy listens elsewhere. If the sandbox is not in the record (paused), it calls
304
304
the API's `ResumeSandbox` gRPC and retries — paused sandboxes wake transparently on traffic.
305
-
By default the record is the API-owned `sandbox:catalog:{id}`; the `orchestrator-routing-prioritized`
306
-
flag switches the read to the orchestrator-owned`sandbox:routing:{id}` (see "Sandbox routing
307
-
records" below).
305
+
By default the record is the orchestrator-owned `sandbox:routing:{id}`; turning the
306
+
`orchestrator-routing-prioritized`flag off switches the read back to the API-owned
307
+
`sandbox:catalog:{id}` (see "Sandbox routing records" below).
308
308
309
309
### Dashboard API (`packages/dashboard-api`)
310
310
@@ -461,12 +461,12 @@ today. Both have the same JSON shape (`sandbox_catalog.SandboxInfo` in
461
461
462
462
| Record | Key | Writer | Written | Deleted |
463
463
|---|---|---|---|---|
464
-
| API-owned (default) |`sandbox:catalog:{sandboxID}`| API (cloud) or the cluster edge from gRPC metadata (BYOC) | after `Create` returns | before `Pause`/`Kill` is sent to the node |
465
-
| Orchestrator-owned (v1, flag-gated) |`sandbox:routing:{sandboxID}`| orchestrator, `packages/orchestrator/pkg/routing`| on `MarkRunning` (sandbox enters the live map, envd is ready) | on `MarkStopping` (kill, pause, checkpoint, crash) |
464
+
| API-owned (fallback) |`sandbox:catalog:{sandboxID}`| API (cloud) or the cluster edge from gRPC metadata (BYOC) | after `Create` returns | before `Pause`/`Kill` is sent to the node |
465
+
| Orchestrator-owned (default) |`sandbox:routing:{sandboxID}`| orchestrator, `packages/orchestrator/pkg/routing`| on `MarkRunning` (sandbox enters the live map, envd is ready) | on `MarkStopping` (kill, pause, checkpoint, crash) |
466
466
467
-
**The API-owned record is still the source of truth.**client-proxy reads `sandbox:catalog:{id}`
468
-
unless the `orchestrator-routing-prioritized` flag is on. The orchestrator-owned record is a v1
469
-
test path. It runs next to the API path and does not replace it yet.
467
+
**client-proxy reads the orchestrator-owned record by default.**It reads `sandbox:catalog:{id}`
468
+
only when the `orchestrator-routing-prioritized` flag is off. The API still writes its record next
469
+
to the orchestrator one, so the flag can be turned off without a gap.
470
470
471
471
Two feature flags in `packages/shared/pkg/featureflags` control the new path:
472
472
@@ -475,14 +475,16 @@ Two feature flags in `packages/shared/pkg/featureflags` control the new path:
475
475
(`orchestrator.routing.publish.total{result=error}`); the sandbox keeps running. Build sandboxes
476
476
are skipped. The delete is guarded by `execution_id` in a Lua script, so a stale lifecycle never
477
477
removes the record of a newer execution.
478
-
-`orchestrator-routing-prioritized` (client-proxy, **default off**): resolve the node from `sandbox:routing:{id}`
478
+
-`orchestrator-routing-prioritized` (client-proxy, **default on**): resolve the node from `sandbox:routing:{id}`
479
479
instead of `sandbox:catalog:{id}`. There is no fallback to the API-owned record on a miss. A miss
480
480
goes to the auto-resume path (`ResumeSandbox` gRPC to the API), same as today.
481
481
482
-
Rollout order: `orchestrator-routing-publish` is on by default, so every live sandbox has a record
483
-
one maximum sandbox length after the orchestrator deploy. Then turn on
484
-
`orchestrator-routing-prioritized`. To roll back, turn off `orchestrator-routing-prioritized`; the
485
-
API path is untouched. Turn off `orchestrator-routing-publish` only to stop the extra Redis write.
482
+
Deploy order: every orchestrator on the cluster must run a build with `orchestrator-routing-publish`
483
+
on for one maximum sandbox length before client-proxy is upgraded to a build with
484
+
`orchestrator-routing-prioritized` on by default. Otherwise requests for sandboxes started before
485
+
the orchestrator upgrade miss the record and go to the resume path. To roll back, turn off
486
+
`orchestrator-routing-prioritized`; the API path is untouched. Turn off
487
+
`orchestrator-routing-publish` only to stop the extra Redis write.
486
488
487
489
The TTL of both records is `sandbox_max_length_in_hours` from the write time. The record is
0 commit comments