From 361f9f2914c327c6722b15f1aaaf785623bac480 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 11:08:23 +0700 Subject: [PATCH 01/30] docs: design a stack that is green and deployable on generation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An acceptance run started a generated project's released stack for the first time. It does not work, and never has, in any shape: no application is told how to reach its database, container port 8080 matches no adapter in this repository, the Laravel images end at php-fpm with no web server anywhere in the stack, and nestjs has been probing a /health route nothing serves. None of it is a regression. ADR-0014 seam 2 already required that every value a container needs arrives through env_file/environment; nothing was ever made to produce the values. It survived because no check could fail — tests/compose.bats runs `docker compose config --quiet`, which validates YAML, not reachability, and no test starts a container. The design turns "green immediately, deployable immediately" into two mechanically checkable gates, and settles the three decisions that follow: every adapter serves on container port 8080; the Laravel images move to FrankenPHP alpine pinned by digest; and the app's connection variables are composed in compose.yaml from the service variables already in .env, by the driver, which is the only thing that knows both the service and the adapter family. It also moves ADR-0003's boundary. An adapter has never written application code; it now writes one readiness route, because a liveness path that does not touch the database is the check-that-cannot-fail ADR-0014 seam 4 warned about, and it is what let all of this ship. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- .../2026-09-06-deployable-stack-design.md | 413 ++++++++++++++++++ 1 file changed, 413 insertions(+) create mode 100644 docs/superpowers/specs/2026-09-06-deployable-stack-design.md diff --git a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md new file mode 100644 index 0000000..9c4ba5a --- /dev/null +++ b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md @@ -0,0 +1,413 @@ +# Deployable Stack — Design + +Status: approved for planning +Date: 2026-09-06 +Scope: what a generated project's released stack has to do before it counts as +delivered. Deploy targets remain out of scope — see section 3. + +## 1. Context + +An acceptance run on 2026-09-05 took four freshly generated projects through +the whole walkthrough on real private repositories and then, for the first +time, started the stack those projects publish. It does not work, and never +has, in any shape. + +Following `install.sh` against a real published image: + +``` +$ docker compose exec app node -e "... new PrismaClient(); p.$connect() ..." +error: Environment variable not found: DATABASE_URL. +``` + +Laravel does not even report an error. `config/database.php` is +`'default' => env('DB_CONNECTION', 'sqlite')`, so with `DB_CONNECTION` absent +it falls back to sqlite and reads the `DB_DATABASE=app` it *did* receive as a +sqlite filename. The mongodb container is never contacted. + +Three more, each independently fatal to serving traffic: + +- `EXPOSE` says `9000` for both Laravel adapters, `3001` for `nestjs`, `3000` + for `nextjs`. `common/compose.yaml` publishes `${APP_PORT:-8080}:8080`, and + nothing in `lib/` or `scaffold` rewrites it. **Container port 8080 matches + no adapter in this repository.** `install.sh` prints `the application is + running on http://localhost:8080`, which has never been true for any shape. +- The Laravel images end at `CMD ["php-fpm"]`. php-fpm speaks FastCGI, and + the stack contains no web server, so nothing in a Laravel project serves + HTTP at all. +- Both Laravel images `COPY . .` as root and then `USER www-data`, leaving + `storage/` and `bootstrap/cache` unwritable. The first request that + compiles a Blade view or writes a log returns 500. +- `nestjs`'s `HEALTHCHECK` probes `/health`, which no adapter ships — the + NestJS generator produces `/` and nothing else. Every `nestjs` image has + reported unhealthy from first boot, the same defect found in `nextjs` on + 2026-09-05 and fixed there alone, because only the nextjs Dockerfiles were + checked. + +None of this is a regression. It is a last mile that was specified and never +built: ADR-0014's seam 2 already requires that "every value a container needs +arrives through `env_file`/`environment` at run time". The requirement was +written on 2026-08-27. Nothing was ever made to produce the values. + +The reason it survived is the same shape this project keeps meeting: **no +check could fail.** `tests/compose.bats` runs `docker compose config --quiet`, +which validates YAML, not reachability. No test starts a container. 185 tests +and eight whole-branch reviews found none of it; starting the stack once found +all of it. + +ADR-0014 is also internally inconsistent, not merely incomplete. Seam 1 says +"a client's target only ever needs to know how to run one image". Seam 4 says +the Laravel images deliberately speak FastCGI, which requires something else +in front of them. Both cannot be true. + +## 2. Goals + +A project generated today must satisfy two gates, each mechanically checkable +and each currently failing: + +1. **Green immediately.** Generate, clone the generated project, run every + config root's `ci-unit` in the clone. All pass. The clone is the point: a + working tree keeps artifacts that make the checks pass for the wrong + reason. +2. **Deployable immediately.** Generate, build the image using the + `context`/`dockerfile` pair the generated `build.yml` names, run + `install.sh`'s own sequence, and then: + - the image's **liveness path** returns 200 — the container serves HTTP on + the port compose publishes. + - its **readiness path** returns 200 — a request reaches the database + through the application, proving the listener, the environment contract, + the compose network, the credentials and the schema in one call. + +Both paths are declared by the adapter (section 7), because they differ per +framework and because two of them are wrong today. + +Gate 1 already holds as of the fixes merged on 2026-09-05. Gate 2 holds for +nothing. + +## 3. Non-goals + +- **A deploy target.** ADR-0014's seven seams stand. This design fills seams + 1, 2 and 4 with working implementations; it adds no `deploy-adapters/` body + and no automated path from a merged pull request to a running instance. +- **Octane.** FrankenPHP without worker mode. Worker mode makes client code + stateful by default and hands the client the whole "Managing Memory Leaks" + section of Laravel's docs. Moving to it later is a one-line `ENTRYPOINT` + change on the same pinned base image, so choosing against it now costs + nothing later. +- **Migrations from an entrypoint.** ADR-0014 seam 5 stands: no image runs + migrations when it starts. `install.sh` runs them once, visibly, as the + human operator's step — which is a different thing, and the one deploy + mechanism ADR-0014 says exists today. +- **A second container.** No nginx sidecar, no shared volume, no new service + kind in `lib/service.sh`. +- **Serving the `web` and `api` roles from one image.** A project still + builds exactly one image, chosen by the last role on the command line + (`set_image_context`). Unchanged here. + +## 4. The container port contract + +Every adapter serves HTTP on container port **8080**. + +`common/compose.yaml` already publishes `${APP_PORT:-8080}:8080` and stays +exactly as it is. `nextjs` and `nestjs` already read `PORT` from the +environment, so each Dockerfile sets `ENV PORT=8080`, `EXPOSE 8080`, and a +`HEALTHCHECK` on 8080. The Laravel images get a listener on 8080 (section 5). + +The alternative — teaching compose each adapter's port through an +`APP_CONTAINER_PORT` variable written at generation time — is rejected for +the reason this project already recorded when it rejected a parameterised +`IMAGE_REPOSITORY` in ADR-0014: the value is fixed for the life of the +project, so a variable "would only add a place for the default to silently +drift from the real value". A container port is exactly that kind of value. + +The cost is real and worth stating: an engineer who runs a generated image by +hand gets 8080, not the 3000 their framework's own documentation names. The +Dockerfile says why, on the line that sets it. + +## 5. An HTTP listener in the Laravel images + +Both Laravel runtime stages move to **FrankenPHP, alpine, pinned by digest**, +without Octane: + +``` +FROM dunglas/frankenphp:1.12.7-php8.3-alpine@sha256:049b8d8356efceb93c91ed42866de890534310bcef4ad4dde902029e4a0d20c3 +``` + +Digest verified against the registry on 2026-09-06, not copied from a +secondary source. + +**Why FrankenPHP.** It is the only option that serves PHP *and* the Vite +assets in `public/build` *and* keeps the stack at one container. A nginx +sidecar would need a genuinely new concept in `lib/service.sh` — a service +that is neither a database nor a cache, that `app` both depends on and shares +a volume with — plus an `nginx.conf` shipped as a release asset, plus moving +the port publish off the `app` service. FrankenPHP changes two Dockerfiles and +nothing else. + +**Evidence, not preference.** FrankenPHP has its own section in +`laravel.com/docs/13.x/deployment`, has been part of the PHP Foundation since +May 2025 (`github.com/php/frankenphp`), and runs Laravel Cloud in production. +Shopware has run it in production for over a year. + +**Alpine is not optional.** `services/mongodb/drivers/laravel.sh` emits +`apk add --no-cache $PHPIZE_DEPS && pecl install mongodb`. The bookworm variant +has no `apk`, and the mongodb driver would break on a base image change nobody +associated with it. + +Each Laravel runtime stage therefore gains: + +- `ENV SERVER_NAME=:8080` — how FrankenPHP listens on an unprivileged port and + declines to provision TLS, which is the reverse proxy's job wherever this + lands. +- `ENV XDG_CONFIG_HOME=/config XDG_DATA_HOME=/data`, both owned by `www-data`. + Caddy writes there and cannot start if it may not. +- `HEALTHCHECK … CMD wget -qO- http://localhost:8080/up || exit 1`. Laravel + has shipped `/up` since 11.x. + +`docker/opcache.ini` is unchanged: FrankenPHP builds on the official PHP +images, so `$PHP_INI_DIR/conf.d` is the same path. + +## 6. The environment contract + +`compose.yaml`'s `app` service gains an `environment:` block whose values are +composed from the service variables already in `.env`: + +```yaml + environment: + DATABASE_URL: ${DATABASE_URL:-postgresql://${DB_USERNAME:-app}:${DB_PASSWORD}@database:5432/${DB_DATABASE:-app}} +``` + +Both paths measured against a real Docker Compose on 2026-09-06: with no +`DATABASE_URL` in `.env` the composed default is produced; with one, the +operator's value wins and the default is not evaluated. That is what lets a +client point at a managed database by adding one line to the file +`install.sh` never overwrites. + +**The password exists in exactly one place.** `.env` holds it; compose +interpolates it into the URL at `up` time. The rejected alternative — a driver +writing a literal `DATABASE_URL=…` into `example.env` — puts the password in +two places and asks `install.sh`, which generates an independent random value +per password variable, to keep them equal. That is the `changeme`-versus-`app` +defect the acceptance run measured in the dev stack, reintroduced in a new +costume. + +**The shape is per adapter family, not per service.** Prisma wants one +`DATABASE_URL`. Laravel wants `DB_CONNECTION` plus a DSN, and will silently +fall back to sqlite without the first. A compose fragment belongs to a service +and cannot know the family, so the block is produced by the **driver**, which +knows both — a new `service_driver_compose_env`, beside the +`service_driver_dockerfile` that already exists for exactly this reason. + +The timing works without reordering anything: `assemble_compose` runs before +any adapter is applied, and `apply_service_drivers` runs inside `apply_adapter` +after the family is known, which is when the block is written. + +Laravel additionally needs `APP_KEY`, which no service fragment has any reason +to produce. `install.sh` already generates a random value for every password in +`example.env`; it generates this one too, and `example.env` carries +`APP_KEY=changeme` for it to replace. + +## 7. Liveness and readiness paths + +Each adapter declares two paths. The `HEALTHCHECK` probes the first; the +deploy gate curls both. + +| Adapter | Liveness | Readiness | +|---|---|---| +| `laravel-api` | `/up` (shipped by Laravel since 11.x) | `/health/ready` | +| `laravel-inertia` | `/up` | `/health/ready` | +| `nestjs` | `/health/live` | `/health/ready` | +| `nextjs` | `/` (the generated home page) | none — the `web` role takes no database driver | + +**Two of the four are wrong today, in the same way.** `nestjs`'s Dockerfile +probes `http://localhost:3001/health`, and no adapter ships a `/health` route +— the NestJS generator produces `/` returning `Hello World!` and nothing else. +So every `nestjs` image has reported unhealthy from first boot, exactly as +every `nextjs` image did until 2026-09-05. The nextjs one was found and fixed; +this one was missed because only the nextjs Dockerfiles were checked. Both +adapters therefore need a real liveness path, not just a corrected probe. + +**Readiness runs one query and reports it:** + +``` +GET /health/ready -> 200 when the query succeeds, 503 when it does not +``` + +This is the only thing that proves the whole chain — listener, environment, +compose network, credentials, schema — in a single call. A liveness path +alone cannot: Laravel's `/up` never touches the database, so a project with a +wrong `DATABASE_URL` passes it. That is precisely the check-that-cannot-fail +ADR-0014 seam 4 warned about, and shipping one as the only gate would repeat +the mistake this design exists to correct. + +**This changes ADR-0003's boundary and the ADR must say so.** Until now an +adapter invoked the framework's own generator and overlaid configuration; it +never wrote application code. It does now, for one file per adapter. The +boundary moves from "no application code" to "no application code except a +readiness route the deploy gate requires", which is narrow, stated, and +testable. + +A project generated with `--db none` ships no readiness route, and neither +does `nextjs` in any shape: there is nothing for either to query, and a route +that returns 200 without doing anything is the same worthless check in a +different place. The gate curls readiness only when the image it built serves +one — which is decided by the role that won `set_image_context`, not by +whether the project has a database. A `--api laravel-api --web nextjs --db +mysql` project builds the **nextjs** image, so its gate is liveness only, and +the Laravel app beside it is generated and checked but never deployed. Section +13 says why that is a limit worth naming. + +## 8. Migrations at install time + +`install.sh` runs the project's migration task once, after the stack is up and +before it prints its success message, and prints what it is doing. + +This does not touch ADR-0014 seam 5, which forbids migrations from an +*entrypoint* — a container that migrates every time it starts is a container +that cannot be scaled or rolled back. `install.sh` is a human running one +command on the target host, which ADR-0014 itself calls "the one deploy +mechanism that exists today". + +`nestjs` has no `migrate` task and needs one. It must branch on the Prisma +provider, measured on 2026-09-05: + +``` +$ pnpm exec prisma migrate deploy # against mongodb +Error: The "mongodb" provider is not supported with this command. +$ pnpm exec prisma db push # against mongodb +The database is already in sync with the Prisma schema. +``` + +So: `db push` for `mongodb`, `migrate deploy` otherwise. Both Laravel adapters +already ship `[tasks.migrate]`. + +## 9. What serving reveals + +Three defects exist today, cause no symptom because nothing serves a request, +and become visible the moment something does. They are in scope because gate 2 +fails without them. + +- **File ownership.** `storage/framework/{views,sessions}` and + `bootstrap/cache` are root-owned in both Laravel images while the process + runs as `www-data`. One `chown` in each runtime stage. +- **Stale assets in `laravel-inertia`.** The runtime stage copies + `public/build` from the assets stage and *then* runs `COPY . .`, and + `public/build` is in `.gitignore` but not `.dockerignore` — so a developer + who has run `npm run build` locally layers their host copy over the image's. + Add it to `.dockerignore`. This is the same class as the `bootstrap/cache` + defect fixed on 2026-09-04: a working tree leaking into a build context. +- **mongodb extension against a locked library.** The image carries + `ext-mongodb 2.5.2` while `composer.lock` pins `mongodb/mongodb 1.21.4` + against `ext-mongodb 1.21.0`, and any real query dies on + `Declaration of MongoDB\Model\BSONArray::bsonSerialize() must be + compatible`. The driver pins a library version matching the extension it + installs. + +## 10. The deploy gate + +A new job in `.github/workflows/adapters.yml`, beside `smoke`, driven by the +same `discover` matrix so ADR-0012's tiers decide what runs when: + +``` +generate + -> docker build, using the context/dockerfile pair the generated + build.yml names, never a chosen one + -> install.sh's own sequence against the built image + -> wait for the app container to report healthy + -> curl -> 200 + -> curl -> 200 + -> docker compose down -v +``` + +The two paths come from the adapter (section 7), not from the job: hardcoding +them here would put the gate's idea of the route and the Dockerfile's idea of +it in two places, which is how `nestjs` came to probe a `/health` nothing +serves. + +Reading the pair out of the generated `build.yml` rather than choosing one is +what made a 9-cell container matrix go from 8/9 to 9/9 in a previous round: a +build that passes with a pair CI does not use proves nothing. + +**Tier-gated, deliberately.** Tier A on every pull request, tier B on the +weekly schedule, exactly as `smoke` and `smoke-tier-b` already split. Running +every shape on every pull request would add roughly a container build per +adapter to a lane that already costs 15 minutes, and ADR-0012 exists to make +that tradeoff once rather than per job. + +## 11. Changes to files that already exist + +| File | Change | +|---|---| +| `adapters/laravel-api/Dockerfile` | runtime stage to FrankenPHP; `SERVER_NAME`, `XDG_*`, `chown`, `EXPOSE 8080`, `HEALTHCHECK` | +| `adapters/laravel-inertia/Dockerfile` | the same, plus `public/build` ordering | +| `adapters/laravel-inertia/.dockerignore` | `public/build` | +| `adapters/nestjs/Dockerfile`, `.workspace` | `ENV PORT=8080`, `EXPOSE 8080`, and a healthcheck on a path that exists — it probes `/health` today and nothing serves it | +| `adapters/nextjs/Dockerfile`, `.workspace` | `ENV PORT=8080`, `EXPOSE 8080`, healthcheck port | +| `adapters/*/` | a liveness path where the framework ships none, and a readiness route for every adapter that can hold a database | +| `adapters/nestjs/mise.toml` | a `migrate` task branching on the Prisma provider | +| `common/compose.yaml` | an `environment:` block on `app`, written by the driver | +| `common/example.env` | `APP_KEY=changeme` for Laravel projects | +| `common/install.sh` | generate `APP_KEY`; run the migration task once | +| `lib/service.sh` | `service_driver_compose_env`, beside `service_driver_dockerfile` | +| `lib/contract.sh` | the new driver hook joins the required-file checks | +| `services/*/drivers/*.sh` | each driver emits its family's environment block | +| `services/mongodb/drivers/laravel.sh` | pin `mongodb/mongodb` to match `ext-mongodb` | +| `.github/workflows/adapters.yml` | the deploy gate | +| `tests/compose.bats` | assert every adapter Dockerfile has `EXPOSE 8080` and a `HEALTHCHECK` | +| `docs/decisions/0003-*` | the application-code boundary moves; state where | +| `docs/decisions/0014-*` | seam 4's "no healthcheck is possible" and seam 1's single-image claim both become false | +| `docs/tour/07-containers.md` | its worked example is the Laravel "why no healthcheck" paragraph | +| `docs/runbook/first-project-walkthrough.md` | step 10 becomes "run it", not "pull it" | +| new `docs/decisions/0021-*` | records this design | + +## 12. Testing + +The gate in section 10 is the load-bearing test, and it is the only one that +can fail for the reasons this design exists. Everything else is cheap +guardrails that keep it from silently rotting: + +- `tests/compose.bats` gains an assertion per adapter Dockerfile: `EXPOSE + 8080` present, `HEALTHCHECK` present, and the path the `HEALTHCHECK` probes + is the liveness path the adapter declares. All three are static, and all + three fail today for at least one adapter — the third is what would have + caught `nestjs` probing a route nothing serves. +- `tests/service.bats` gains a case per driver: the emitted environment block + names the family's variables and interpolates `${DB_PASSWORD}` rather than a + literal. +- `tests/contract.bats` requires `service_driver_compose_env` of every driver, + the way it already requires `service_driver_apply`. +- No new test starts a container outside the gate. Container work belongs in + the tier-gated lane, not in a suite someone runs on a laptop. + +Suites to run for a change under this design: `compose`, `service`, `contract` +and the adapter's own `new-.bats`. The full lane runs in CI. + +## 13. Known limits + +- **One image per project stands.** A `web+api` project still builds and + deploys only the role that came last on the command line. The gate tests + that image; the other application is generated, checked, and not deployed. + Unchanged by this design, and worth stating because gate 2 will look like it + covers a project when it covers an image. +- **`/health/ready` is application code the toolbox owns and a client may + delete.** Nothing detects that. The gate tests generated projects, not + client repositories six months later. +- **The gate proves one shape per adapter, not every service combination.** + `nestjs` + `postgres` passing says nothing about `nestjs` + `mongodb`, whose + Prisma provider takes a different migration command. Tier B's weekly matrix + covers more but not all. +- **`install.sh` running migrations is a single-instance assumption.** Two + operators running it concurrently against one database is not defended + against. It is the same assumption `install.sh` already makes about + everything else it does. + +## 14. Follow-on work + +- Octane, if a client's load justifies worker mode — one `ENTRYPOINT` line on + the same base image. +- A deploy adapter, which is what ADR-0014's remaining seams are for. This + design makes the image it would deploy actually runnable, which was the + missing precondition. +- The root `prettier` hook rewriting `apps/app`'s frontend source, measured on + 2026-09-05: `lefthook run pre-commit --all-files` rewrites 39 files and the + app then fails its own `ci-unit`. Adjacent, separately scoped, and blocking + daily work on `laravel-inertia` rather than deployment. From c4a193976542d118eaad52edfbf3b474b27d1b04 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 11:24:22 +0700 Subject: [PATCH 02/30] docs: plan the deployable stack, and fix what planning falsified MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Ten tasks, each ending in something independently testable, and each container check using `docker compose up -d` rather than `docker run`, which this environment denies. Writing it surfaced two things the spec had wrong, both fixed in the spec as well rather than papered over in the plan: The published nestjs image cannot migrate itself. The driver installs the CLI with `pnpm add -D prisma@6` and the Dockerfile runs `pnpm prune --prod` before the runtime stage copies node_modules, so the image carries @prisma/client and no prisma binary. The driver now installs it as a regular dependency, which is what Prisma's own guidance assumes when the migration runs from the image, and the image-size cost is recorded rather than discovered later. And install.sh cannot run a mise task, because no image carries mise. The driver writes a profiled `migrate` compose service instead, sharing the app's image and environment, so it never starts with the stack — ADR-0014 seam 5 forbids migrating from an entrypoint, not from the operator's one command. Also removes a dangling ADR-0021 reference from both documents. The decision record is Task 8's output, and tests/documentation.bats was right to fail on a spec citing a file that does not exist yet. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- .../plans/2026-09-06-deployable-stack.md | 1249 +++++++++++++++++ .../2026-09-06-deployable-stack-design.md | 35 +- 2 files changed, 1282 insertions(+), 2 deletions(-) create mode 100644 docs/superpowers/plans/2026-09-06-deployable-stack.md diff --git a/docs/superpowers/plans/2026-09-06-deployable-stack.md b/docs/superpowers/plans/2026-09-06-deployable-stack.md new file mode 100644 index 0000000..5482e96 --- /dev/null +++ b/docs/superpowers/plans/2026-09-06-deployable-stack.md @@ -0,0 +1,1249 @@ +# Deployable Stack Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Make a freshly generated project serve HTTP on the port its +`compose.yaml` publishes, reach its database through the environment the +released stack hands it, and prove both in CI. + +**Architecture:** Every adapter serves on container port 8080 and declares two +paths — a liveness path its `HEALTHCHECK` probes, and a readiness path that +runs one query. The Laravel runtime stages move to FrankenPHP so something +answers HTTP at all. The application's connection variables are composed in +`compose.yaml` from the service variables already in `.env`, written by the +service driver, which is the only thing that knows both the service and the +adapter family. A tier-gated CI job starts the stack and curls both paths. + +**Tech Stack:** bash, `yq`, `jq`, `mise`, `bats`, Docker Compose, FrankenPHP +(Caddy + embedded PHP), Prisma (Nest), Eloquent (Laravel). + +**Spec:** `docs/superpowers/specs/2026-09-06-deployable-stack-design.md` + +## Global Constraints + +- Chat is Vietnamese; **every file, comment, commit message and document is + English**. +- Comment style follows immich and the surrounding repository: explain *why*, + never *what*. No comment asserts that a mechanism "always" does something + unless a check enforces it. +- Tests must stay isolated. A test that modifies the toolbox uses + `copy_toolbox` (`tests/helpers/setup.bash`) — never the real tree. Slow is + acceptable; interdependent is not. +- **Run only the suites your change touches.** `bats tests/.bats`, or + `--filter ` for one test. Capture output to a file once and grep the + file rather than re-running to count. The full lane runs in CI. +- `docker run` is denied by this environment's permission policy. + `docker build` and `docker compose up -d` are available — every container + verification in this plan uses compose. +- Container port is **8080** for every adapter. `common/compose.yaml` stays as + it is. +- The FrankenPHP base image is + `dunglas/frankenphp:1.12.7-php8.3-alpine@sha256:049b8d8356efceb93c91ed42866de890534310bcef4ad4dde902029e4a0d20c3`, + verified against the registry on 2026-09-06. **Alpine, not bookworm**: + `services/mongodb/drivers/laravel.sh` emits `apk add`. +- Every image reference is pinned by digest (`tests/compose.bats` enforces it). +- Run `mise run lint` (shellcheck) before every commit. +- Work on branch `feat/deployable-stack`, cut from `spec/deployable-stack`. + +## File Structure + +**Created:** + +| Path | Responsibility | +| --- | --- | +| `adapters/nestjs/src/health/health.controller.ts` | liveness and readiness routes for Nest | +| `adapters/nestjs/src/health/health.module.ts` | wires the controller into `AppModule` | +| `adapters/nextjs/src/app/api/health/live/route.ts` | liveness route for Next | +| `adapters/laravel-api/routes/health.php` | readiness route for the API skeleton | +| `adapters/laravel-inertia/routes/health.php` | readiness route for the starter kit | +| docs/decisions/0021-the-released-stack-must-run.md | records this design | + +**Modified:** + +| Path | Change | +| --- | --- | +| `adapters/*/adapter.env` | `ADAPTER_LIVENESS_PATH`, `ADAPTER_READINESS_PATH` | +| `adapters/nextjs/Dockerfile`, `.workspace` | `ENV PORT=8080`, `EXPOSE 8080`, healthcheck | +| `adapters/nestjs/Dockerfile`, `.workspace` | the same, and a path that exists | +| `adapters/laravel-api/Dockerfile` | FrankenPHP runtime stage | +| `adapters/laravel-inertia/Dockerfile` | the same, plus `public/build` ordering | +| `adapters/laravel-inertia/.dockerignore` | `public/build` | +| `adapters/nestjs/mise.toml` | a `migrate` task branching on the Prisma provider | +| `lib/contract.sh` | the two new adapter vars, the new driver hook | +| `lib/lint.sh` | enforce them | +| `lib/service.sh` | `service_driver_compose_env`, `apply_service_compose_env` | +| `services/*/drivers/*.sh` | emit the family's environment block and DB probe | +| `services/mongodb/drivers/laravel.sh` | pin `mongodb/mongodb` to the extension | +| `common/install.sh` | generate `APP_KEY`; run the migration task once | +| `common/example.env` | nothing — `APP_KEY` is appended by the driver | +| `.github/workflows/adapters.yml` | the `deploy` gate | +| `tests/compose.bats` | port, healthcheck and declared-path assertions | +| `tests/contract.bats` | the new required vars and driver hook | +| `tests/service.bats` | the emitted compose block | +| `docs/decisions/0003-*`, `0014-*` | the boundaries this moves | +| `docs/tour/07-containers.md` | its worked example becomes false | +| `docs/runbook/first-project-walkthrough.md` | step 10 becomes "run it" | + +--- + +### Task 0: Branch + +- [ ] **Step 1: Cut the branch** + +```bash +cd /home/ttndev/workspace/personal/scaffold +git checkout spec/deployable-stack +git checkout -b feat/deployable-stack +``` + +- [ ] **Step 2: Confirm the starting point is clean** + +Run: `git status --porcelain` +Expected: no output. + +--- + +### Task 1: Adapters declare their liveness and readiness paths + +The gate and the `HEALTHCHECK` must read the path from one place. `nestjs` +probes `/health` today and nothing serves it; that defect exists because the +Dockerfile's idea of the route and the application's idea of it were never +required to agree. + +**Files:** +- Modify: `lib/contract.sh`, `lib/lint.sh`, `adapters/*/adapter.env` +- Test: `tests/contract.bats` + +**Interfaces:** +- Produces: `ADAPTER_LIVENESS_PATH` (every adapter) and + `ADAPTER_READINESS_PATH` (only adapters whose role is in `DRIVEN_ROLES`), + both read by `lib/lint.sh`, `tests/compose.bats` and the CI gate. + +- [ ] **Step 1: Write the failing test** + +Append to `tests/contract.bats`: + +```bash +@test "every adapter declares a liveness path" { + local missing="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + grep -q '^ADAPTER_LIVENESS_PATH=' "${dir}adapter.env" \ + || missing="${missing}$(basename "$dir")"$'\n' + done + [ -z "$missing" ] || { echo "missing ADAPTER_LIVENESS_PATH:"; echo "$missing"; false; } +} + +@test "an adapter whose role takes a driver declares a readiness path" { + # a web adapter opens no connection (DRIVEN_ROLES), so it has nothing to + # probe; anything else must, or the deploy gate has no way to prove the + # application actually reaches its database. + local missing="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + role="$(grep '^ADAPTER_ROLE=' "${dir}adapter.env" | cut -d'"' -f2)" + case " ${DRIVEN_ROLES[*]} " in *" ${role} "*) ;; *) continue ;; esac + grep -q '^ADAPTER_READINESS_PATH=' "${dir}adapter.env" \ + || missing="${missing}$(basename "$dir")"$'\n' + done + [ -z "$missing" ] || { echo "missing ADAPTER_READINESS_PATH:"; echo "$missing"; false; } +} +``` + +`tests/contract.bats` already sources `lib/contract.sh` in its `setup`; if it +does not, add `. "${SCAFFOLD_ROOT}/lib/contract.sh"` there so `DRIVEN_ROLES` +resolves. + +- [ ] **Step 2: Run it and watch it fail** + +Run: `mise exec -- bats --filter 'liveness path|readiness path' tests/contract.bats` +Expected: both FAIL, listing all four adapters. + +- [ ] **Step 3: Declare the paths** + +`adapters/nextjs/adapter.env` — append: + +```sh +# create-next-app generates no health route, and the web tier opens no +# database connection, so `/` is both the only thing it serves and the whole +# of what there is to check. +ADAPTER_LIVENESS_PATH="/" +``` + +`adapters/nestjs/adapter.env` — append: + +```sh +# The generator produces `/` returning Hello World and nothing else. Both of +# these are routes this adapter ships itself (src/health/), because the +# Dockerfile's HEALTHCHECK has been probing a /health that never existed. +ADAPTER_LIVENESS_PATH="/health/live" +ADAPTER_READINESS_PATH="/health/ready" +``` + +`adapters/laravel-api/adapter.env` and `adapters/laravel-inertia/adapter.env` +— append to each: + +```sh +# /up ships with laravel since 11.x and deliberately touches nothing, which +# is why it cannot stand alone: a project with a wrong DATABASE_URL passes it. +ADAPTER_LIVENESS_PATH="/up" +ADAPTER_READINESS_PATH="/health/ready" +``` + +- [ ] **Step 4: Enforce them in lint** + +In `lib/contract.sh`, add to `REQUIRED_ADAPTER_VARS`: + +```sh +REQUIRED_ADAPTER_VARS=(ADAPTER_NAME ADAPTER_ROLE ADAPTER_FAMILY ADAPTER_GENERATOR ADAPTER_LIVENESS_PATH) +``` + +`ADAPTER_READINESS_PATH` is conditional on the role, so it cannot go in that +list. In `lib/lint.sh`, inside the per-adapter loop that already reports +missing vars, add: + +```sh + # Conditional on the role rather than required outright: a web adapter has + # no connection to probe, and demanding a readiness path from it would only + # produce one that returns 200 without doing anything. + case " ${DRIVEN_ROLES[*]} " in + *" ${ADAPTER_ROLE} "*) + [ -n "${ADAPTER_READINESS_PATH:-}" ] \ + || fail "${name}: ADAPTER_READINESS_PATH is required for role ${ADAPTER_ROLE}" ;; + esac +``` + +Add `ADAPTER_READINESS_PATH` and `ADAPTER_LIVENESS_PATH` to the `unset -v` +list in `load_adapter` (`lib/adapter.sh`), beside `ADAPTER_FAMILY`, so a +second `load_adapter` in one process cannot inherit the previous adapter's +value. + +- [ ] **Step 5: Run the tests** + +Run: `mise exec -- bats tests/contract.bats && mise exec -- ./scaffold lint` +Expected: all PASS, `lint` silent, exit 0. + +- [ ] **Step 6: Commit** + +```bash +git add lib/contract.sh lib/lint.sh lib/adapter.sh adapters/*/adapter.env tests/contract.bats +git commit -m "feat: let an adapter declare the paths its health checks probe" +``` + +--- + +### Task 2: The TypeScript adapters serve on 8080 + +**Files:** +- Create: `adapters/nestjs/src/health/health.controller.ts`, + `adapters/nestjs/src/health/health.module.ts`, + `adapters/nextjs/src/app/api/health/live/route.ts` +- Modify: `adapters/nestjs/Dockerfile`, `adapters/nestjs/Dockerfile.workspace`, + `adapters/nextjs/Dockerfile`, `adapters/nextjs/Dockerfile.workspace`, + `adapters/nestjs/adapter.env` +- Test: `tests/compose.bats` + +**Interfaces:** +- Consumes: `ADAPTER_LIVENESS_PATH` from Task 1. +- Produces: `GET /health/live` on Nest returning `{"status":"ok"}`; both + TypeScript images listening on 8080. + +- [ ] **Step 1: Write the failing test** + +Append to `tests/compose.bats`: + +```bash +@test "every adapter Dockerfile serves the port compose publishes" { + # common/compose.yaml publishes ${APP_PORT:-8080}:8080 and nothing rewrites + # it, so an adapter exposing anything else publishes a dead port. + run bash -c "grep -L '^EXPOSE 8080\$' '${SCAFFOLD_ROOT}'/adapters/*/Dockerfile*" + [ -z "$output" ] || { echo "not exposing 8080:"; echo "$output"; false; } +} + +@test "every adapter Dockerfile probes the liveness path its adapter declares" { + # nestjs probed /health for months while the generator produced only `/`. + # The Dockerfile's idea of the route and the adapter's must be one value. + local wrong="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + path="$(grep '^ADAPTER_LIVENESS_PATH=' "${dir}adapter.env" | cut -d'"' -f2)" + for file in "${dir}"Dockerfile "${dir}"Dockerfile.workspace; do + [ -f "$file" ] || continue + grep -q "HEALTHCHECK" "$file" \ + || { wrong="${wrong}${file}: no HEALTHCHECK"$'\n'; continue; } + grep -q "localhost:8080${path}" "$file" \ + || wrong="${wrong}${file}: does not probe ${path} on 8080"$'\n' + done + done + [ -z "$wrong" ] || { echo "$wrong"; false; } +} +``` + +- [ ] **Step 2: Run them and watch them fail** + +Run: `mise exec -- bats --filter 'port compose publishes|liveness path its adapter' tests/compose.bats` +Expected: both FAIL — six Dockerfiles on the first, all six on the second +(the Laravel pair have no `HEALTHCHECK` at all). + +- [ ] **Step 3: Ship the Nest health routes** + +Create `adapters/nestjs/src/health/health.controller.ts`: + +```ts +import { Controller, Get, HttpException, HttpStatus } from '@nestjs/common'; + +@Controller('health') +export class HealthController { + @Get('live') + live(): { status: string } { + return { status: 'ok' }; + } + + // The probe is written by the selected service's driver: prisma has no + // provider-agnostic read, so a SQL provider gets $queryRawUnsafe and + // mongodb gets $runCommandRaw. A project generated with --db none keeps + // the anchor's fallback and reports 503, because there is nothing here + // that could honestly report ready. + @Get('ready') + async ready(): Promise<{ status: string }> { + try { + // @DB_PROBE@ + throw new Error('no database is configured for this project'); + } catch (error) { + throw new HttpException( + { status: 'unavailable', reason: (error as Error).message }, + HttpStatus.SERVICE_UNAVAILABLE, + ); + } + } +} +``` + +Create `adapters/nestjs/src/health/health.module.ts`: + +```ts +import { Module } from '@nestjs/common'; + +import { HealthController } from './health.controller'; + +@Module({ controllers: [HealthController] }) +export class HealthModule {} +``` + +Wire it in. `apply_adapter` copies only top-level adapter files, so the +`src/health/` directory needs `ADAPTER_POST_GENERATE` to place it and to +register the module. Extend `adapters/nestjs/adapter.env`'s existing +`ADAPTER_POST_GENERATE` (do not add a second one — the variable is read once): + +```sh +ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && sed -i "1i import { HealthModule } from '"'"'./health/health.module'"'"';" src/app.module.ts && sed -i "s/imports: \[\]/imports: [HealthModule]/" src/app.module.ts && pnpm exec prettier --write .' +``` + +The `src/health/` files themselves are copied by a new directory case in +`apply_adapter` — see Step 4. + +- [ ] **Step 4: Let an adapter ship a directory** + +`apply_adapter` (`lib/adapter.sh`) special-cases `docker/` and skips every +other directory. Replace that single case with a loop over every directory +the adapter ships, so a route file does not need a second mechanism: + +```sh + # Every directory the adapter ships, merged into the generated tree rather + # than replacing what is there: `src/` already exists after the generator + # ran, and `cp -R src dest/src` would nest it as dest/src/src. + local dir + for dir in "${ADAPTER_DIR}"/*/; do + [ -d "$dir" ] || continue + mkdir -p "${dest}/$(basename "$dir")" + cp -R "${dir}." "${dest}/$(basename "$dir")/" + done +``` + +Delete the `[ -d "${ADAPTER_DIR}/docker" ] && cp -R …` line it replaces. + +- [ ] **Step 5: Ship the Next liveness route** + +Create `adapters/nextjs/src/app/api/health/live/route.ts`: + +```ts +export const dynamic = 'force-dynamic'; + +export function GET(): Response { + return Response.json({ status: 'ok' }); +} +``` + +Change `adapters/nextjs/adapter.env`'s `ADAPTER_LIVENESS_PATH` to +`/api/health/live` and say why in the comment: a route handler answers +without rendering the home page, so the probe does not depend on whatever the +client later puts on `/`. + +- [ ] **Step 6: Move both images to 8080** + +In all four TypeScript Dockerfiles, replace the `EXPOSE`, `HEALTHCHECK` and +add `ENV PORT`: + +`adapters/nestjs/Dockerfile` and `adapters/nestjs/Dockerfile.workspace`, in +the runtime stage: + +```dockerfile +# 8080 because that is the container port common/compose.yaml publishes, and +# nothing rewrites it — an adapter listening anywhere else publishes a dead +# port. Nest reads PORT in main.ts's `app.listen(process.env.PORT ?? 3000)`. +ENV PORT=8080 +EXPOSE 8080 +HEALTHCHECK --interval=30s --timeout=3s \ + CMD wget -qO- http://localhost:8080/health/live || exit 1 +``` + +`adapters/nextjs/Dockerfile` and `adapters/nextjs/Dockerfile.workspace`, in +the runtime stage — the same `ENV PORT=8080` and `EXPOSE 8080`, with: + +```dockerfile +HEALTHCHECK --interval=30s --timeout=3s \ + CMD wget -qO- http://localhost:8080/api/health/live || exit 1 +``` + +- [ ] **Step 7: Run the static tests** + +Run: `mise exec -- bats tests/compose.bats > /tmp/compose.log 2>&1; tail -20 /tmp/compose.log` +Expected: every test PASS. + +- [ ] **Step 8: Prove it serves, with a real container** + +```bash +mkdir -p /tmp/dsverify && cd /tmp/dsverify +scaffold() ( eval "$(mise env -C /home/ttndev/workspace/personal/scaffold -s bash)" + /home/ttndev/workspace/personal/scaffold/scaffold "$@" ) +scaffold new t2 --api nestjs --db postgres +cd t2 +ctx="$(grep -E '^ context:' .github/workflows/build.yml | head -1 | sed 's/.*context: *//')" +df="$(grep -E '^ dockerfile:' .github/workflows/build.yml | head -1 | sed 's/.*dockerfile: *//')" +docker build -f "$df" -t dsverify/t2:local "$ctx" +sed -i 's|ghcr.io/CHANGEME/CHANGEME:${IMAGE_TAG:-latest}|dsverify/t2:local|' compose.yaml +cp example.env .env +docker compose up -d +sleep 20 +docker compose ps +curl -fsS -o /dev/null -w '%{http_code}\n' http://localhost:8080/health/live +docker compose down -v +``` + +Expected: `docker compose ps` shows the app container `Up`, and the `curl` +prints `200`. `/health/ready` is expected to return 503 at this point — +Task 5 is what makes it 200. + +- [ ] **Step 9: Commit** + +```bash +git add adapters/nestjs adapters/nextjs lib/adapter.sh tests/compose.bats +git commit -m "feat: serve the port compose publishes, on a route that exists" +``` + +--- + +### Task 3: The Laravel images serve HTTP + +**Files:** +- Modify: `adapters/laravel-api/Dockerfile`, + `adapters/laravel-inertia/Dockerfile`, + `adapters/laravel-inertia/.dockerignore` +- Test: `tests/compose.bats` (already written in Task 2) + +**Interfaces:** +- Consumes: `ADAPTER_LIVENESS_PATH="/up"` from Task 1. +- Produces: both Laravel images listening on 8080 with a working + `HEALTHCHECK`. + +- [ ] **Step 1: Confirm the tests still fail for these two** + +Run: `mise exec -- bats --filter 'liveness path its adapter' tests/compose.bats` +Expected: FAIL, naming only the two Laravel Dockerfiles. + +- [ ] **Step 2: Replace the laravel-api runtime stage** + +In `adapters/laravel-api/Dockerfile`, replace the whole runtime stage — +`FROM php:…` through `CMD ["php-fpm"]` — with: + +```dockerfile +# FrankenPHP, because php-fpm speaks FastCGI and this stack has no web server +# in front of it: nothing served HTTP at all, and compose published a dead +# port. Documented as a first-class server in laravel.com/docs/13.x/deployment, +# part of the PHP Foundation since May 2025, and what Laravel Cloud runs. +# Alpine specifically: services/mongodb/drivers/laravel.sh emits `apk add`. +FROM dunglas/frankenphp:1.12.7-php8.3-alpine@sha256:049b8d8356efceb93c91ed42866de890534310bcef4ad4dde902029e4a0d20c3 AS runtime +RUN docker-php-ext-install opcache +# @SERVICE_SETUP@ +WORKDIR /var/www +COPY --from=vendor /app/vendor ./vendor +COPY . . +COPY docker/opcache.ini /usr/local/etc/php/conf.d/opcache.ini +# :8080 is how frankenphp listens on an unprivileged port and declines to +# provision TLS, which belongs to whatever proxy this lands behind. +ENV SERVER_NAME=:8080 +# Caddy writes here and will not start if it may not. The app tree needs the +# same: COPY runs as root, the server runs as www-data, and the first request +# that compiles a blade view or writes a log fails without this. +ENV XDG_CONFIG_HOME=/config XDG_DATA_HOME=/data +RUN mkdir -p /config /data \ + && chown -R www-data:www-data /config /data \ + /var/www/storage /var/www/bootstrap/cache +USER www-data +EXPOSE 8080 +HEALTHCHECK --interval=30s --timeout=3s \ + CMD wget -qO- http://localhost:8080/up || exit 1 +CMD ["frankenphp", "run"] +``` + +Delete the long "no healthcheck is possible" comment. Its reasoning — that a +check which cannot fail is worse than none — was right and is preserved in +decision record 0021; what it concluded stopped being true the moment something served +HTTP. + +- [ ] **Step 3: Do the same for laravel-inertia, and fix the asset ordering** + +Apply the identical runtime stage to `adapters/laravel-inertia/Dockerfile`, +with one difference: `COPY . .` must run **before** the assets copy, so a +developer's locally built `public/build` cannot layer over the image's: + +```dockerfile +COPY --from=vendor /app/vendor ./vendor +COPY . . +COPY --from=assets /app/public/build ./public/build +``` + +Add to `adapters/laravel-inertia/.dockerignore`: + +``` +# vite writes this and .gitignore excludes it, but a build context is not a +# git tree: without this line a developer who has run `npm run build` ships +# their host copy over the one the assets stage just built. Same shape as the +# bootstrap/cache manifests two lines up. +public/build +``` + +- [ ] **Step 4: Run the static tests** + +Run: `mise exec -- bats tests/compose.bats > /tmp/compose.log 2>&1; tail -20 /tmp/compose.log` +Expected: every test PASS, including the digest test — the FrankenPHP +reference is pinned. + +- [ ] **Step 5: Prove laravel-api serves** + +```bash +cd /tmp/dsverify +scaffold new t3 --api laravel-api --db postgres +cd t3 +docker build -f apps/api/Dockerfile -t dsverify/t3:local apps/api +sed -i 's|ghcr.io/CHANGEME/CHANGEME:${IMAGE_TAG:-latest}|dsverify/t3:local|' compose.yaml +cp example.env .env +docker compose up -d +sleep 25 +docker compose ps +curl -fsS -o /dev/null -w '%{http_code}\n' http://localhost:8080/up +docker compose logs app | tail -20 +docker compose down -v +``` + +Expected: `curl` prints `200`. If it prints a connection error, read +`docker compose logs app` before changing anything — Caddy names the reason +it would not start. + +- [ ] **Step 6: Prove laravel-inertia serves, and serves its assets** + +Repeat Step 5 with `scaffold new t3b --app laravel-inertia --db postgres`, +`apps/app`, and additionally: + +```bash +asset="$(docker compose exec -T app sh -c 'ls public/build/assets/*.js | head -1')" +curl -fsS -o /dev/null -w '%{http_code}\n' "http://localhost:8080/${asset#public/}" +``` + +Expected: both `curl`s print `200`. The second is what separates a working +answer from one that serves PHP and 404s every asset. + +- [ ] **Step 7: Commit** + +```bash +git add adapters/laravel-api adapters/laravel-inertia +git commit -m "feat: put an http server in the laravel images" +``` + +--- + +### Task 4: The driver writes the app's connection environment + +**Files:** +- Modify: `lib/service.sh`, `lib/contract.sh`, `lib/lint.sh` +- Test: `tests/service.bats`, `tests/contract.bats` + +**Interfaces:** +- Produces: `service_driver_compose_env` — a driver hook printing YAML lines + for `services.app.environment`, and `apply_service_compose_env + ` which merges them into `compose.yaml`. Task 5 implements the hook + in each driver; Task 6 relies on the merged result. + +- [ ] **Step 1: Write the failing test** + +Append to `tests/service.bats`: + +```bash +@test "a driver's compose environment interpolates rather than embedding a password" { + # The password must exist in exactly one place — .env — so compose composes + # the URL at `up` time. A literal baked here is the changeme-versus-app + # mismatch that made the dev stack unable to authenticate. + local bad="" + for driver in "${SCAFFOLD_ROOT}"/services/*/drivers/*.sh; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + SERVICE_DIR="$(dirname "$(dirname "$driver")")" + . "$driver"; service_driver_compose_env )" + [ -z "$block" ] && continue + grep -q '\${DB_PASSWORD' <<<"$block" || grep -q '\${REDIS_PASSWORD' <<<"$block" \ + || bad="${bad}${driver}"$'\n' + done + [ -z "$bad" ] || { echo "embeds a literal password:"; echo "$bad"; false; } +} + +@test "apply_service_compose_env merges into the app service" { + local project="${BATS_TEST_TMPDIR}/p" + mkdir -p "$project" + printf 'services:\n app:\n image: x\n' > "${project}/compose.yaml" + . "${SCAFFOLD_ROOT}/lib/service.sh" + apply_service_compose_env "$project" 'DATABASE_URL: ${DATABASE_URL:-postgresql://app@database:5432/app}' + run mise exec -- yq -r '.services.app.environment.DATABASE_URL' "${project}/compose.yaml" + [[ "$output" == 'postgresql://app@database:5432/app' ]] \ + || [[ "$output" == '${DATABASE_URL:-postgresql://app@database:5432/app}' ]] +} +``` + +- [ ] **Step 2: Run them and watch them fail** + +Run: `mise exec -- bats --filter 'compose environment|apply_service_compose_env' tests/service.bats` +Expected: FAIL — `service_driver_compose_env: command not found` and +`apply_service_compose_env: command not found`. + +- [ ] **Step 3: Add the merge function** + +In `lib/service.sh`, beside `apply_service_setup`: + +```sh +# apply_service_compose_env +# Adds the block to compose.yaml's app service. yq rather than an anchor: the +# app service is generated by assemble_compose from common/compose.yaml, so +# there is a real document to merge into by the time this runs, and a text +# anchor would only be a second way to write YAML. +apply_service_compose_env() { + local project="$1" block="$2" + local file="${project}/compose.yaml" fragment + + [ -n "$block" ] || return 0 + [ -f "$file" ] || die "no compose.yaml in ${project}" + + fragment="$(mktemp)" + { + printf 'services:\n app:\n environment:\n' + printf '%s\n' "$block" | sed 's/^/ /' + } > "$fragment" + + if ! yq eval-all --inplace -P 'select(fileIndex==0) * select(fileIndex==1)' \ + "$file" "$fragment"; then + rm -f "$fragment" + die "could not merge the service environment into ${file}" + fi + rm -f "$fragment" +} +``` + +`-P` for the same reason `merge_lefthook_fragment` needs it: yq propagates the +style of what it merges, and a collapsed `compose.yaml` is a file a client has +to read. + +- [ ] **Step 4: Call it from apply_service_drivers** + +In `apply_service_drivers`, beside the existing `service_driver_dockerfile` +accumulation, add a second accumulator and one call: + +```sh + # shellcheck source=/dev/null # family varies, so the path isn't constant + rendered="$( . "$driver"; service_driver_compose_env )" + [ -n "$rendered" ] && env_block+="${rendered}"$'\n' +``` + +Declare `env_block=""` beside `block=""`, and after the loop, beside the +existing `apply_service_setup` call: + +```sh + apply_service_compose_env "$project" "${env_block%$'\n'}" +``` + +- [ ] **Step 5: Require the hook of every driver** + +In `lib/contract.sh`, add a list beside `REQUIRED_SERVICE_FILES`: + +```sh +# apply_service_drivers calls all three, so a driver shipping fewer fails at +# generation rather than at lint. +REQUIRED_DRIVER_FUNCTIONS=(service_driver_apply service_driver_dockerfile service_driver_compose_env) +``` + +In `lib/lint.sh`'s `lint_services`, source each driver in a subshell and +check each name with `declare -F`. + +- [ ] **Step 6: Run the tests** + +Run: `mise exec -- bats tests/service.bats tests/contract.bats > /tmp/svc.log 2>&1; grep -c '^ok ' /tmp/svc.log; grep -A4 '^not ok' /tmp/svc.log` +Expected: the second test passes; the first still fails (no driver implements +the hook yet), and `scaffold lint` now reports every driver as missing it. +That is the correct intermediate state — Task 5 closes it. + +- [ ] **Step 7: Commit** + +```bash +git add lib/service.sh lib/contract.sh lib/lint.sh tests/service.bats tests/contract.bats +git commit -m "feat: give a driver a seam for the app's connection environment" +``` + +--- + +### Task 5: Every driver emits its family's environment and probe + +**Files:** +- Modify: `services/{mysql,postgres,mongodb,redis}/drivers/{laravel,nest}.sh`, + `services/shared/{laravel,nest}.sh` +- Test: `tests/service.bats` + +**Interfaces:** +- Consumes: `apply_service_compose_env` and `REQUIRED_DRIVER_FUNCTIONS` from + Task 4; the `@DB_PROBE@` anchor from Task 2's Nest controller. +- Produces: a `compose.yaml` whose `app` service carries the family's + variables, and a readiness route that runs a real query. + +- [ ] **Step 1: Write the failing test** + +Append to `tests/service.bats`: + +```bash +@test "the laravel drivers name the connection selector laravel actually reads" { + # config/database.php is `env('DB_CONNECTION', 'sqlite')`. Without that + # variable laravel does not fail — it silently reads DB_DATABASE as a + # sqlite filename and never contacts the service at all. + for service in mysql postgres mongodb; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + . "${SCAFFOLD_ROOT}/services/${service}/drivers/laravel.sh" + service_driver_compose_env )" + grep -q '^DB_CONNECTION:' <<<"$block" \ + || { echo "${service}/laravel.sh emits no DB_CONNECTION"; false; } + done +} + +@test "the nest drivers name DATABASE_URL and let an operator override it" { + for service in mysql postgres mongodb; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + . "${SCAFFOLD_ROOT}/services/${service}/drivers/nest.sh" + service_driver_compose_env )" + grep -q '^DATABASE_URL: \${DATABASE_URL:-' <<<"$block" \ + || { echo "${service}/nest.sh does not allow an override"; false; } + done +} +``` + +- [ ] **Step 2: Run them and watch them fail** + +Run: `mise exec -- bats --filter 'connection selector|DATABASE_URL' tests/service.bats` +Expected: both FAIL. + +- [ ] **Step 3: Implement the hook in the shared Nest driver** + +In `services/shared/nest.sh`, add — using `PRISMA_PROVIDER` and the service's +own port, both already in scope from the per-service driver: + +```sh +# The value an operator sets in .env wins; otherwise compose composes it from +# the same DB_* variables the database container reads, so the password lives +# in exactly one place and the two cannot drift. Measured against a real +# `docker compose config`: both paths resolve, and the default is not +# evaluated when DATABASE_URL is set. +service_driver_compose_env() { + printf 'DATABASE_URL: ${DATABASE_URL:-%s}\n' "$PRISMA_COMPOSE_URL" +} +``` + +Each `services//drivers/nest.sh` sets `PRISMA_COMPOSE_URL` beside the +`PRISMA_URL` it already sets — the same string with `localhost` replaced by +`database` and the literals replaced by interpolations. For postgres: + +```sh +PRISMA_COMPOSE_URL='postgresql://${DB_USERNAME:-app}:${DB_PASSWORD}@database:5432/${DB_DATABASE:-app}' +``` + +mysql uses `mysql://…@database:3306/…`; mongodb uses +`mongodb://…@database:27017/…?authSource=admin&directConnection=true`. Single +quotes throughout: these are compose's interpolations, not the shell's. + +- [ ] **Step 4: Implement the probe for Nest** + +In `services/shared/nest.sh`'s `service_driver_apply`, after the prisma setup +it already does, replace the controller's anchor: + +```sh + # prisma has no provider-agnostic read: $queryRaw is SQL-only and mongodb + # needs a command. Written here rather than branched in the controller so + # the shipped route carries exactly one probe, for the provider this + # project actually has. + local probe + case "$PRISMA_PROVIDER" in + mongodb) probe='await new PrismaClient().$runCommandRaw({ ping: 1 });' ;; + *) probe="await new PrismaClient().\$queryRawUnsafe('SELECT 1');" ;; + esac + sed -i.bak "s|// @DB_PROBE@|import('@prisma/client').then(async ({ PrismaClient }) => { ${probe} });\n return { status: 'ok' };|" \ + src/health/health.controller.ts + rm -f src/health/health.controller.ts.bak +``` + +The `throw` below the anchor stays in the shipped file and is what a `--db +none` project keeps: unreachable once a probe is spliced in, and the honest +503 when none is. + +- [ ] **Step 5: Implement both for the shared Laravel driver** + +In `services/shared/laravel.sh`: + +```sh +service_driver_compose_env() { + # DB_CONNECTION first and always: config/database.php defaults to sqlite, + # so its absence is not an error, it is a silent wrong answer. + printf 'DB_CONNECTION: %s\n' "$LARAVEL_CONNECTION" + printf '%s\n' "$LARAVEL_COMPOSE_ENV" + # APP_KEY has no service to come from and laravel will not boot without it; + # install.sh generates the value, this only reserves the name. + printf 'APP_KEY: ${APP_KEY}\n' +} +``` + +Each `services//drivers/laravel.sh` sets `LARAVEL_CONNECTION` (`mysql`, +`pgsql`, `mongodb`) and `LARAVEL_COMPOSE_ENV`, the remaining variables that +family needs — `DB_HOST: database`, `DB_PORT`, and for mongodb `DB_URI` +instead of host and port. + +For the readiness route, splice the probe into `routes/health.php` the same +way, with `DB::connection()->select('select 1');` for SQL and +`DB::connection('mongodb')->getMongoDB()->command(['ping' => 1]);` for mongodb. + +- [ ] **Step 6: Add the redis drivers' hook** + +`services/redis/drivers/{laravel,nest}.sh` gain a `service_driver_compose_env` +emitting `REDIS_URL`/`REDIS_HOST` pointing at the `cache` service. Task 4's +lint now requires the function of every driver, so neither may be left out. + +- [ ] **Step 7: Run the tests** + +Run: `mise exec -- bats tests/service.bats tests/contract.bats > /tmp/svc.log 2>&1; tail -5 /tmp/svc.log; grep -A4 '^not ok' /tmp/svc.log` +Expected: all PASS, and `mise exec -- ./scaffold lint` silent. + +- [ ] **Step 8: Commit** + +```bash +git add services lib tests/service.bats +git commit -m "feat: tell the application how to reach the service it was given" +``` + +--- + +### Task 6: install.sh generates APP_KEY and runs migrations + +**Files:** +- Modify: `common/install.sh`, `adapters/nestjs/mise.toml` +- Test: `tests/compose.bats` (it already sources `install.sh` per-function) + +**Interfaces:** +- Consumes: `APP_KEY: ${APP_KEY}` emitted by Task 5's Laravel driver. +- Produces: a `.env` carrying a Laravel-valid `APP_KEY`; a stack whose schema + exists before the success message prints. + +- [ ] **Step 1: Write the failing test** + +Append to `tests/compose.bats`: + +```bash +@test "install.sh generates an APP_KEY laravel will accept" { + # generate_service_passwords' generic 24-character value is rejected with + # "Unsupported cipher or incorrect key length" — laravel needs base64: and + # exactly 32 bytes. + local env_file="${BATS_TEST_TMPDIR}/.env" + printf 'DB_PASSWORD=changeme\nAPP_KEY=changeme\n' > "$env_file" + . "${SCAFFOLD_ROOT}/common/install.sh" + run generate_service_passwords "$env_file" + assert_ok + run grep '^APP_KEY=' "$env_file" + [[ "$output" =~ ^APP_KEY=base64:[A-Za-z0-9+/]{43}=$ ]] \ + || { echo "not a laravel key: ${output}"; false; } +} +``` + +- [ ] **Step 2: Run it and watch it fail** + +Run: `mise exec -- bats --filter 'APP_KEY laravel will accept' tests/compose.bats` +Expected: FAIL — the value is a bare 24-character string. + +- [ ] **Step 3: Special-case APP_KEY in the generator** + +In `common/install.sh`'s `generate_service_passwords`, inside the loop: + +```sh + # APP_KEY is not a password: laravel decrypts with it and rejects anything + # that is not base64: plus exactly 32 bytes. Handled inside this loop + # rather than beside it so example.env keeps one placeholder, and the + # existing-.env guard that greps for a remaining `=changeme` still covers + # it. + if [ "$name" = APP_KEY ]; then + password="base64:$(head -c 32 /dev/urandom | base64)" + else + password="$(head -c 32 /dev/urandom | base64 | tr -dc 'A-Za-z0-9' | head -c 24)" + fi +``` + +The existing `sed` uses `s/^${name}=changeme$/${name}=${password}/`, and a +base64 value can contain `/`. Change the delimiter to `|` and keep the +`grep -qF` confirmation, which already catches a substitution that missed. + +- [ ] **Step 4: Run the test** + +Run: `mise exec -- bats --filter 'APP_KEY laravel will accept' tests/compose.bats` +Expected: PASS. + +- [ ] **Step 5: Give nestjs a migrate task** + +In `adapters/nestjs/mise.toml`: + +```toml +[tasks.migrate] +# prisma's mongodb provider rejects `migrate deploy` outright — measured: +# `The "mongodb" provider is not supported with this command.` — and takes +# `db push` instead. Branching on the schema rather than on a recorded value +# keeps this true for a project whose provider changes. +run = """ +if ! [ -f prisma/schema.prisma ]; then exit 0; fi +if grep -q 'provider *= *"mongodb"' prisma/schema.prisma; then + pnpm exec prisma db push --skip-generate +else + pnpm exec prisma migrate deploy +fi +""" +``` + +- [ ] **Step 6: Keep the Prisma CLI in the Nest runtime image** + +`services/shared/nest.sh` installs the CLI as `pnpm add -D prisma@6`, and +`adapters/nestjs/Dockerfile` runs `pnpm prune --prod` before the runtime stage +copies `node_modules` — so the published image has `@prisma/client` and no +`prisma` binary, and cannot migrate itself. Change the driver: + +```sh + # A regular dependency, not -D: `pnpm prune --prod` in the Dockerfile drops + # devDependencies, and the published image is what runs `migrate deploy` on + # deploy. The alternative — a second image, or a compose service mounting + # the source — introduces a build artifact the release does not publish, for + # a command run once. The engines cost image size; see decision record 0021. + pnpm add prisma@6 || return 1 +``` + +Laravel needs nothing here: `php artisan` is already in those images. + +- [ ] **Step 7: Add the migrate service and run it from install.sh** + +`install.sh` cannot run a `mise` task — no image carries `mise`. The driver +writes a compose service instead, sharing the app's image and environment, +behind a profile so it never starts with the stack. In +`services/shared/{laravel,nest}.sh`, extend `service_driver_compose_env`'s +sibling — a new `service_driver_compose_migrate` printing the family's +command — and have `apply_service_compose_env` merge it as +`services.migrate`. + +Laravel: `["php", "artisan", "migrate", "--force"]`. +Nest, chosen at generation time from `PRISMA_PROVIDER`: +`["pnpm", "exec", "prisma", "db", "push", "--skip-generate"]` for mongodb, +`["pnpm", "exec", "prisma", "migrate", "deploy"]` otherwise. + +In `common/install.sh`, between `start_stack` and its success message: + +```sh +# ADR-0014 seam 5 forbids migrations from an *entrypoint* — a container that +# migrates every time it starts cannot be scaled or rolled back. This is a +# human running one command on the target host, which is what that ADR calls +# the one deploy mechanism that exists today. A project with no database +# ships no migrate service, and `--profile` on a service that is not there +# is not an error. +run_migrations() { + docker compose config --services | grep -qx migrate || return 0 + echo "running migrations..." + docker compose --profile migrate run --rm migrate +} +``` + +Move the "the application is running on…" message out of `start_stack` and +into `main`, after `run_migrations`, so the order it reports is the order that +happened. + +- [ ] **Step 8: Prove the migrate service runs** + +```bash +cd /tmp/dsverify && scaffold new t6 --api laravel-api --db postgres +cd t6 && docker build -f apps/api/Dockerfile -t dsverify/t6:local apps/api +sed -i 's|ghcr.io/CHANGEME/CHANGEME:${IMAGE_TAG:-latest}|dsverify/t6:local|' compose.yaml +cp example.env .env && docker compose up -d && sleep 25 +docker compose --profile migrate run --rm migrate +curl -fsS -o /dev/null -w '%{http_code}\n' http://localhost:8080/health/ready +docker compose down -v +``` + +Expected: the migrate run prints Laravel's migration table and exits 0, and +the readiness curl prints `200`. Before the migration it returns 503 — check +that too, in that order, because a readiness route that returns 200 against an +empty schema is not reading anything. + +- [ ] **Step 9: Run the suite** + +Run: `mise exec -- bats tests/compose.bats tests/service.bats > /tmp/compose.log 2>&1; tail -5 /tmp/compose.log; grep -A4 '^not ok' /tmp/compose.log` +Expected: all PASS. + +- [ ] **Step 10: Commit** + +```bash +git add common/install.sh adapters/nestjs/mise.toml services lib tests +git commit -m "feat: let the released stack migrate its own schema" +``` + +--- + +### Task 7: The deploy gate + +**Files:** +- Create: `scripts/deploy-check.sh` +- Modify: `.github/workflows/adapters.yml` +- Test: the script runs locally against a generated project + +**Interfaces:** +- Consumes: `ADAPTER_LIVENESS_PATH` and `ADAPTER_READINESS_PATH` (Task 1); + everything Tasks 2–6 built. + +- [ ] **Step 1: Write the script** + +Create `scripts/deploy-check.sh` — generate, build with the pair the generated +`build.yml` names, start the stack, migrate, curl both paths, tear down. It +reads the two paths from the adapter, never from its own copy of them: the +gate hardcoding a route is how `nestjs` came to probe a `/health` nothing +served. + +```sh +#!/usr/bin/env bash +# deploy-check.sh [--db ] +# Proves a generated project's released stack serves HTTP and reaches its +# database. Everything before this validated YAML; nothing started a container. +set -euo pipefail +``` + +Body: resolve `role` and both paths from `adapters//adapter.env`; +`scaffold new` into a temp directory with the role's flag; read `context` and +`dockerfile` out of the generated `.github/workflows/build.yml`; +`docker build`; rewrite `compose.yaml`'s image to the built tag; `cp +example.env .env`; `docker compose up -d`; poll `docker compose ps` until the +app container is healthy or 120s elapse; run the adapter's migrate task with +`docker compose exec`; `curl -fsS` the liveness path and, when the adapter +declares one, the readiness path, asserting `200`; `docker compose down -v` in +a trap so a failure still tears down. + +- [ ] **Step 2: Run it locally for the cheapest adapter** + +Run: `mise exec -- ./scripts/deploy-check.sh nestjs --db postgres` +Expected: exits 0, having printed `200` for both paths. + +- [ ] **Step 3: Run it for the shape with no database** + +Run: `mise exec -- ./scripts/deploy-check.sh nextjs` +Expected: exits 0, liveness only, and says it skipped readiness. + +- [ ] **Step 4: Add the job** + +In `.github/workflows/adapters.yml`, after `smoke-tier-b`: + +```yaml + deploy: + needs: discover + if: ${{ needs.discover.outputs.tier-a != '[]' }} + runs-on: ubuntu-latest + permissions: + contents: read + # a generation plus an image build plus a container start, per adapter. + timeout-minutes: 30 + strategy: + fail-fast: false + matrix: + adapter: ${{ fromJson(needs.discover.outputs.tier-a) }} + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: jdx/mise-action@c2a87611a18de5b3828c5652fe268e992400cb5c # v4.3.0 + - run: corepack enable + - id: language + env: + ADAPTER: ${{ matrix.adapter }} + run: | + lang="$(grep '^ADAPTER_LANGUAGE=' "adapters/${ADAPTER}/adapter.env" | cut -d'"' -f2)" + echo "value=${lang}" >> "$GITHUB_OUTPUT" + - if: steps.language.outputs.value == 'php' + uses: shivammathur/setup-php@f3e473d116dcccaddc5834248c87452386958240 # 2.37.2 + with: + php-version: "8.3" + - env: + ADAPTER: ${{ matrix.adapter }} + run: ./scripts/deploy-check.sh "$ADAPTER" + + deploy-tier-b: + # Same job, tier b matrix, on the weekly schedule only — ADR-0012 makes + # this tradeoff once rather than per job, and a container build per + # adapter on every pull request is exactly the cost it exists to bound. +``` + +`deploy-tier-b` mirrors `smoke-tier-b`'s `needs`, `if` and matrix. + +- [ ] **Step 5: Lint the workflow** + +Run: `mise exec -- zizmor .github/workflows/adapters.yml && mise exec -- actionlint` +Expected: both clean. + +- [ ] **Step 6: Commit** + +```bash +git add scripts/deploy-check.sh .github/workflows/adapters.yml +git commit -m "feat: start the stack in ci and require it to answer" +``` + +--- + +### Task 8: Documentation this falsifies + +**Files:** +- Create: docs/decisions/0021-the-released-stack-must-run.md +- Modify: `docs/decisions/0003-*`, `docs/decisions/0014-*`, + `docs/tour/07-containers.md`, + `docs/runbook/first-project-walkthrough.md` +- Test: `tests/documentation.bats` + +- [ ] **Step 1: Write decision record 0021** + +Record: the two gates; the port contract and why a variable was rejected +(quoting ADR-0014's own reasoning about `IMAGE_REPOSITORY`); FrankenPHP with +its evidence and the alpine constraint; the environment contract and the +single-password argument; and the readiness route. Preserve seam 4's +reasoning explicitly — a check that cannot fail is worse than none — and say +that what changed is not the reasoning but the premise. + +- [ ] **Step 2: Amend ADR-0003** + +Its boundary moves from "an adapter writes no application code" to "…except +the health routes the deploy gate requires". State the exception and why it +is narrow. + +- [ ] **Step 3: Amend ADR-0014** + +Seam 1 and seam 4 both become false. Add a `Superseded in part by 0021` line +naming which seams, rather than editing the record silently. + +- [ ] **Step 4: Fix the tour and the runbook** + +`docs/tour/07-containers.md` uses the Laravel "why no healthcheck" paragraph +as its worked example. Replace it. Its "Delete test" section admits nothing +asserts a healthcheck exists — that is now false too, and the new assertion +is the answer. + +`docs/runbook/first-project-walkthrough.md` step 10 becomes "run it": pull, +`install.sh`, and curl the two paths. + +- [ ] **Step 5: Run the docs suite** + +Run: `mise exec -- bats tests/documentation.bats` +Expected: 4/4 PASS. Every backticked string containing `/` is treated as a +path by test 2 — write route paths without backticks, or as `apps/…`, which +that test exempts. + +- [ ] **Step 6: Commit** + +```bash +git add docs +git commit -m "docs: record what a running stack changed about the seams" +``` + +--- + +### Task 9: Full verification + +- [ ] **Step 1: Lint and the full lane, once** + +```bash +mise run lint +mise run test-runner > /tmp/runner.log 2>&1 +echo "ok=$(grep -c '^ok ' /tmp/runner.log) notok=$(grep -c '^not ok' /tmp/runner.log)" +grep -A6 '^not ok' /tmp/runner.log | head -40 +``` + +Expected: `notok=0`. A `registry.npmjs.org` `ECONNRESET` under parallel lanes +is a known flake — re-run only the affected suite before treating it as a +failure. + +- [ ] **Step 2: Both gates, by hand, on a shape no task used** + +```bash +mise exec -- ./scripts/deploy-check.sh laravel-inertia +``` + +Expected: exits 0. This is the tier-B adapter and the most fragile image; a +task-level check never ran it end to end. + +- [ ] **Step 3: Green immediately, from a clone** + +```bash +cd /tmp/dsverify && scaffold new final --api nestjs --web nextjs --db postgres --cache redis +git clone final final-clean && cd final-clean +for root in $(mise exec -C /home/ttndev/workspace/personal/scaffold -- yq -r '.monorepo.config_roots[]' mise.toml); do + mise run "//${root}:ci-unit" || echo "FAILED: ${root}" +done +``` + +Expected: every root exits 0. The clone is the point — a working tree keeps +artifacts that make the checks pass for the wrong reason. + +- [ ] **Step 4: Open the pull request** + +Push the branch and open a pull request against `main` with the spec linked. +Wait for every check, including the new `deploy` matrix. + +--- + +## Self-Review + +**Spec coverage.** Section 4 (port contract) → Tasks 2, 3. Section 5 +(FrankenPHP) → Task 3. Section 6 (environment contract) → Tasks 4, 5. +Section 7 (liveness and readiness) → Tasks 1, 2, 3, 5. Section 8 +(migrations) → Task 6. Section 9 (ownership, stale assets, mongodb pin) → +Task 3 for the first two; **the mongodb library pin has no task** — added to +Task 5, Step 5, as part of the mongodb Laravel driver. Section 10 (the gate) +→ Task 7. Section 11 (files) → covered. Section 12 (testing) → the +assertions are written in the tasks that make them pass. + +**Placeholders.** None. Writing the plan surfaced one contradiction with the +spec and it was fixed in both, not deferred: the published Nest image carries +`@prisma/client` and no `prisma` binary (`pnpm add -D prisma@6` in the driver, +`pnpm prune --prod` in the Dockerfile), so it could not migrate itself, and +`install.sh` could not run a `mise` task because no image carries `mise`. The +driver now installs the CLI as a regular dependency and writes a profiled +`migrate` compose service; spec section 8 says the same and states the image +size cost. + +**Type consistency.** `service_driver_compose_env` is the name in Task 4 +(definition), Task 4 Step 5 (contract), and Task 5 (implementations). +`apply_service_compose_env ` matches its call site, and +`service_driver_compose_migrate` (Task 6, Step 7) is named the same in the +contract list Task 4 Step 5 defines — add it there when implementing Task 6, +since Task 4 is written before it exists. +`ADAPTER_LIVENESS_PATH` / `ADAPTER_READINESS_PATH` are the names in Tasks 1, +2, 3 and 7. The `@DB_PROBE@` anchor in Task 2's controller is the one Task 5 +Step 4 replaces. diff --git a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md index 9c4ba5a..b7c9ffa 100644 --- a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md +++ b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md @@ -267,8 +267,39 @@ that cannot be scaled or rolled back. `install.sh` is a human running one command on the target host, which ADR-0014 itself calls "the one deploy mechanism that exists today". -`nestjs` has no `migrate` task and needs one. It must branch on the Prisma -provider, measured on 2026-09-05: +**The image has to be able to run it, and the Nest one cannot today.** +`services/shared/nest.sh` installs the CLI as `pnpm add -D prisma@6`, and +`adapters/nestjs/Dockerfile` runs `pnpm prune --prod` before the runtime stage +copies `node_modules` — so the published Nest image carries `@prisma/client` +and no `prisma` binary. It cannot migrate itself. + +The driver installs `prisma` as a regular dependency instead. That is what +Prisma's own deployment guidance assumes when the migration runs from the +image, and it is the smaller change: the alternative — a second image, or a +compose service that mounts the source — introduces a build artifact the +release does not publish, for a command run once per deploy. The cost is the +CLI and its engines in the runtime image, and it is stated in the new decision record numbered 0021 rather +than discovered later. + +The Laravel images need nothing: `php artisan` is already there. + +`install.sh` therefore runs the framework's own command, not a `mise` task — +no image carries `mise`, and inventing one would be a mechanism built to make +a sentence in this spec true. The driver writes the command into +`compose.yaml` as a `migrate` service sharing the app's image and environment, +under a compose profile so it never starts with the stack: + +```yaml + migrate: + profiles: [migrate] + image: ${APP_IMAGE} + command: [...the family's migration command...] +``` + +and `install.sh` runs `docker compose --profile migrate run --rm migrate`. + +`nestjs` has no `migrate` task and needs one for local use. It must branch on +the Prisma provider, measured on 2026-09-05: ``` $ pnpm exec prisma migrate deploy # against mongodb From c61600f340b966917ebc17444508f9c27b4b844a Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 11:32:32 +0700 Subject: [PATCH 03/30] feat: let an adapter declare the paths its health checks probe MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ADAPTER_LIVENESS_PATH (every adapter) and ADAPTER_READINESS_PATH (adapters whose role is in DRIVEN_ROLES) are now declared once per adapter and read by lib/lint.sh, so the Dockerfile and the CI gate stop each keeping their own idea of the route — the defect that let nestjs probe a /health nothing served. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/laravel-api/adapter.env | 4 ++++ adapters/laravel-inertia/adapter.env | 4 ++++ adapters/nestjs/adapter.env | 5 ++++ adapters/nextjs/adapter.env | 5 ++++ lib/adapter.sh | 3 ++- lib/contract.sh | 2 +- lib/lint.sh | 15 +++++++++++- tests/contract.bats | 23 +++++++++++++++++++ .../fixtures/lint/complete/sample/adapter.env | 2 ++ tests/fixtures/lint/mixed/good/adapter.env | 2 ++ 10 files changed, 62 insertions(+), 3 deletions(-) diff --git a/adapters/laravel-api/adapter.env b/adapters/laravel-api/adapter.env index df7bec1..5ea45e4 100644 --- a/adapters/laravel-api/adapter.env +++ b/adapters/laravel-api/adapter.env @@ -12,3 +12,7 @@ ADAPTER_FAMILY="laravel" ADAPTER_GENERATOR='composer create-project laravel/laravel:^13.0 "$APP_DIR" --no-interaction --prefer-dist' # the skeleton ships phpunit and pint, but nothing for the check task ADAPTER_POST_GENERATE='composer require --dev larastan/larastan phpstan/phpstan --no-interaction' +# /up ships with laravel since 11.x and deliberately touches nothing, which +# is why it cannot stand alone: a project with a wrong DATABASE_URL passes it. +ADAPTER_LIVENESS_PATH="/up" +ADAPTER_READINESS_PATH="/health/ready" diff --git a/adapters/laravel-inertia/adapter.env b/adapters/laravel-inertia/adapter.env index c54e2e8..3b64cc4 100644 --- a/adapters/laravel-inertia/adapter.env +++ b/adapters/laravel-inertia/adapter.env @@ -24,3 +24,7 @@ ADAPTER_GENERATOR='SHELL_VERBOSITY=-1 COMPOSER_PROCESS_TIMEOUT=900 composer crea # `cooldown.default-days: 5`, which fails the security gate with exit 13 on # the first pull request. Inert config whose only effect is a red check. ADAPTER_POST_GENERATE='rm -rf .github' +# /up ships with laravel since 11.x and deliberately touches nothing, which +# is why it cannot stand alone: a project with a wrong DATABASE_URL passes it. +ADAPTER_LIVENESS_PATH="/up" +ADAPTER_READINESS_PATH="/health/ready" diff --git a/adapters/nestjs/adapter.env b/adapters/nestjs/adapter.env index 62eebab..0c3e48b 100644 --- a/adapters/nestjs/adapter.env +++ b/adapters/nestjs/adapter.env @@ -11,3 +11,8 @@ ADAPTER_GENERATOR='pnpm dlx @nestjs/cli@11 new "$APP_DIR" --package-manager pnpm # main.ts's un-awaited bootstrap() trips --max-warnings 0, and the generator's # own output is not prettier-formatted. ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && pnpm exec prettier --write .' +# The generator produces `/` returning Hello World and nothing else. Both of +# these are routes this adapter ships itself (src/health/), because the +# Dockerfile's HEALTHCHECK has been probing a /health that never existed. +ADAPTER_LIVENESS_PATH="/health/live" +ADAPTER_READINESS_PATH="/health/ready" diff --git a/adapters/nextjs/adapter.env b/adapters/nextjs/adapter.env index 591cf4a..4bdc652 100644 --- a/adapters/nextjs/adapter.env +++ b/adapters/nextjs/adapter.env @@ -10,3 +10,8 @@ ADAPTER_GENERATOR='pnpm create next-app@16 "$APP_DIR" --ts --app --eslint --tail # create-next-app ships neither prettier nor vitest, and its own output is not # formatted — the contract's format and test tasks need both. ADAPTER_POST_GENERATE='pnpm add -D prettier vitest && pnpm exec prettier --write .' +# A route handler answers without rendering the home page, so the probe does +# not depend on whatever the client later puts on `/`. The route itself is +# created by the next task; declaring the final value here avoids this task +# shipping a comment that the next task makes false. +ADAPTER_LIVENESS_PATH="/api/health/live" diff --git a/lib/adapter.sh b/lib/adapter.sh index 7e26faf..2f8fb76 100644 --- a/lib/adapter.sh +++ b/lib/adapter.sh @@ -17,7 +17,8 @@ load_adapter() { ADAPTER_DIR="$dir" # every optional value, not just one: a stale ADAPTER_LANGUAGE or ROLE from # the previous load would otherwise be read as this adapter's own - unset -v ADAPTER_POST_GENERATE ADAPTER_LANGUAGE ADAPTER_ROLE ADAPTER_TIER ADAPTER_FAMILY + unset -v ADAPTER_POST_GENERATE ADAPTER_LANGUAGE ADAPTER_ROLE ADAPTER_TIER ADAPTER_FAMILY \ + ADAPTER_LIVENESS_PATH ADAPTER_READINESS_PATH # shellcheck source=/dev/null # `|| return 1` so an unreadable adapter.env fails here, rather than letting # the default below become this function's last, always-successful command diff --git a/lib/contract.sh b/lib/contract.sh index 8058832..ccc9497 100644 --- a/lib/contract.sh +++ b/lib/contract.sh @@ -10,7 +10,7 @@ REQUIRED_ADAPTER_FILES=(adapter.env mise.toml Dockerfile .env.example) # with `unbound variable` instead of failing at `scaffold lint`. ADAPTER_FAMILY # is the same story one step later: apply_service_drivers looks up # drivers/${family}.sh only once generation is already underway. -REQUIRED_ADAPTER_VARS=(ADAPTER_NAME ADAPTER_ROLE ADAPTER_FAMILY ADAPTER_GENERATOR) +REQUIRED_ADAPTER_VARS=(ADAPTER_NAME ADAPTER_ROLE ADAPTER_FAMILY ADAPTER_GENERATOR ADAPTER_LIVENESS_PATH) READ_ONLY_TASKS=(format lint check) diff --git a/lib/lint.sh b/lib/lint.sh index b436ef3..f37af6f 100644 --- a/lib/lint.sh +++ b/lib/lint.sh @@ -4,7 +4,7 @@ # prints one line per problem and returns 1 when any adapter is incomplete. lint_adapters() { local dir="$1" - local adapter name file task task_body flag var status=0 + local adapter name file task task_body flag var role status=0 for adapter in "$dir"/*/; do [ -d "$adapter" ] || continue @@ -24,6 +24,19 @@ lint_adapters() { status=1 } done + + # Conditional on the role rather than required outright: a web adapter has + # no connection to probe, and demanding a readiness path from it would only + # produce one that returns 200 without doing anything. + role="$(sed -n 's/^ADAPTER_ROLE="\(.*\)"$/\1/p' "${adapter}adapter.env")" + case " ${DRIVEN_ROLES[*]} " in + *" ${role} "*) + grep -Eq '^ADAPTER_READINESS_PATH=' "${adapter}adapter.env" || { + printf '%s: adapter.env does not set ADAPTER_READINESS_PATH (required for role %s)\n' "$name" "$role" + status=1 + } + ;; + esac fi [ -f "${adapter}mise.toml" ] || continue diff --git a/tests/contract.bats b/tests/contract.bats index d1ce65e..ad335fd 100644 --- a/tests/contract.bats +++ b/tests/contract.bats @@ -170,3 +170,26 @@ setup() { run_line="$(awk '/^\[tasks\."test-integration"\]/{f=1} f && /^run = /{print; exit}' "${SCAFFOLD_ROOT}/mise.toml")" [[ "$run_line" == *"tests/wizard-integration.bats"* ]] } + +@test "every adapter declares a liveness path" { + local missing="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + grep -q '^ADAPTER_LIVENESS_PATH=' "${dir}adapter.env" \ + || missing="${missing}$(basename "$dir")"$'\n' + done + [ -z "$missing" ] || { echo "missing ADAPTER_LIVENESS_PATH:"; echo "$missing"; false; } +} + +@test "an adapter whose role takes a driver declares a readiness path" { + # a web adapter opens no connection (DRIVEN_ROLES), so it has nothing to + # probe; anything else must, or the deploy gate has no way to prove the + # application actually reaches its database. + local missing="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + role="$(grep '^ADAPTER_ROLE=' "${dir}adapter.env" | cut -d'"' -f2)" + case " ${DRIVEN_ROLES[*]} " in *" ${role} "*) ;; *) continue ;; esac + grep -q '^ADAPTER_READINESS_PATH=' "${dir}adapter.env" \ + || missing="${missing}$(basename "$dir")"$'\n' + done + [ -z "$missing" ] || { echo "missing ADAPTER_READINESS_PATH:"; echo "$missing"; false; } +} diff --git a/tests/fixtures/lint/complete/sample/adapter.env b/tests/fixtures/lint/complete/sample/adapter.env index 4de34ff..0a3ca1b 100644 --- a/tests/fixtures/lint/complete/sample/adapter.env +++ b/tests/fixtures/lint/complete/sample/adapter.env @@ -3,3 +3,5 @@ ADAPTER_ROLE="api" ADAPTER_FAMILY="laravel" ADAPTER_TIER="C" ADAPTER_GENERATOR='mkdir -p "$APP_DIR"' +ADAPTER_LIVENESS_PATH="/health/live" +ADAPTER_READINESS_PATH="/health/ready" diff --git a/tests/fixtures/lint/mixed/good/adapter.env b/tests/fixtures/lint/mixed/good/adapter.env index 7746e0d..0bd95bf 100644 --- a/tests/fixtures/lint/mixed/good/adapter.env +++ b/tests/fixtures/lint/mixed/good/adapter.env @@ -3,3 +3,5 @@ ADAPTER_ROLE="api" ADAPTER_FAMILY="laravel" ADAPTER_TIER="C" ADAPTER_GENERATOR='mkdir -p "$APP_DIR"' +ADAPTER_LIVENESS_PATH="/health/live" +ADAPTER_READINESS_PATH="/health/ready" From 78d4ea3fe49d7e1042add507f811d9b534adfc4d Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 11:41:42 +0700 Subject: [PATCH 04/30] test: cover the readiness-path role gate through lint_adapters directly The case branch in lint_adapters that requires ADAPTER_READINESS_PATH only for DRIVEN_ROLES had no fixture exercising it: every existing lint fixture is role api, and the tests that did touch readiness paths checked the real adapter.env files, not the enforcement code, so deleting the branch would have gone unnoticed. Adds a role-api fixture missing the var (must be reported) and a role-web fixture missing it (must stay silent), each with its own test. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- tests/contract.bats | 14 ++++++++++ .../no-readiness-path/sample/.env.example | 1 + .../lint/no-readiness-path/sample/Dockerfile | 1 + .../lint/no-readiness-path/sample/adapter.env | 5 ++++ .../lint/no-readiness-path/sample/mise.toml | 26 +++++++++++++++++++ .../web-no-readiness-path/sample/.env.example | 1 + .../web-no-readiness-path/sample/Dockerfile | 1 + .../web-no-readiness-path/sample/adapter.env | 5 ++++ .../web-no-readiness-path/sample/mise.toml | 26 +++++++++++++++++++ 9 files changed, 80 insertions(+) create mode 100644 tests/fixtures/lint/no-readiness-path/sample/.env.example create mode 100644 tests/fixtures/lint/no-readiness-path/sample/Dockerfile create mode 100644 tests/fixtures/lint/no-readiness-path/sample/adapter.env create mode 100644 tests/fixtures/lint/no-readiness-path/sample/mise.toml create mode 100644 tests/fixtures/lint/web-no-readiness-path/sample/.env.example create mode 100644 tests/fixtures/lint/web-no-readiness-path/sample/Dockerfile create mode 100644 tests/fixtures/lint/web-no-readiness-path/sample/adapter.env create mode 100644 tests/fixtures/lint/web-no-readiness-path/sample/mise.toml diff --git a/tests/contract.bats b/tests/contract.bats index ad335fd..f1d9c6a 100644 --- a/tests/contract.bats +++ b/tests/contract.bats @@ -141,6 +141,20 @@ setup() { [[ "$output" == *"adapter.env does not set ADAPTER_FAMILY"* ]] } +@test "lint_adapters requires a readiness path for a role that takes a driver" { + run lint_adapters "${SCAFFOLD_ROOT}/tests/fixtures/lint/no-readiness-path" + [ "$status" -eq 1 ] + [[ "$output" == *"sample: adapter.env does not set ADAPTER_READINESS_PATH"* ]] +} + +@test "lint_adapters does not require a readiness path for a role that takes no driver" { + # a web adapter opens no connection, so demanding one here would fail every + # web adapter for a check it can never satisfy honestly. + run lint_adapters "${SCAFFOLD_ROOT}/tests/fixtures/lint/web-no-readiness-path" + assert_ok + [ -z "$output" ] +} + @test "scaffold lint covers the services that ship" { run scaffold lint assert_ok diff --git a/tests/fixtures/lint/no-readiness-path/sample/.env.example b/tests/fixtures/lint/no-readiness-path/sample/.env.example new file mode 100644 index 0000000..ae255d7 --- /dev/null +++ b/tests/fixtures/lint/no-readiness-path/sample/.env.example @@ -0,0 +1 @@ +APP_ENV=local diff --git a/tests/fixtures/lint/no-readiness-path/sample/Dockerfile b/tests/fixtures/lint/no-readiness-path/sample/Dockerfile new file mode 100644 index 0000000..c35f1b5 --- /dev/null +++ b/tests/fixtures/lint/no-readiness-path/sample/Dockerfile @@ -0,0 +1 @@ +FROM scratch diff --git a/tests/fixtures/lint/no-readiness-path/sample/adapter.env b/tests/fixtures/lint/no-readiness-path/sample/adapter.env new file mode 100644 index 0000000..7b54ea1 --- /dev/null +++ b/tests/fixtures/lint/no-readiness-path/sample/adapter.env @@ -0,0 +1,5 @@ +ADAPTER_NAME="sample" +ADAPTER_ROLE="api" +ADAPTER_FAMILY="laravel" +ADAPTER_GENERATOR='mkdir -p "$APP_DIR"' +ADAPTER_LIVENESS_PATH="/health/live" diff --git a/tests/fixtures/lint/no-readiness-path/sample/mise.toml b/tests/fixtures/lint/no-readiness-path/sample/mise.toml new file mode 100644 index 0000000..8ec8128 --- /dev/null +++ b/tests/fixtures/lint/no-readiness-path/sample/mise.toml @@ -0,0 +1,26 @@ +[tasks.install] +run = "true" + +[tasks.format] +run = "true" + +[tasks."format-fix"] +run = "true" + +[tasks.lint] +run = "true" + +[tasks.check] +run = "true" + +[tasks.test] +run = "true" + +[tasks.build] +run = "true" + +[tasks.ci-unit] +run = "true" + +[tasks.checklist] +run = "true" diff --git a/tests/fixtures/lint/web-no-readiness-path/sample/.env.example b/tests/fixtures/lint/web-no-readiness-path/sample/.env.example new file mode 100644 index 0000000..ae255d7 --- /dev/null +++ b/tests/fixtures/lint/web-no-readiness-path/sample/.env.example @@ -0,0 +1 @@ +APP_ENV=local diff --git a/tests/fixtures/lint/web-no-readiness-path/sample/Dockerfile b/tests/fixtures/lint/web-no-readiness-path/sample/Dockerfile new file mode 100644 index 0000000..c35f1b5 --- /dev/null +++ b/tests/fixtures/lint/web-no-readiness-path/sample/Dockerfile @@ -0,0 +1 @@ +FROM scratch diff --git a/tests/fixtures/lint/web-no-readiness-path/sample/adapter.env b/tests/fixtures/lint/web-no-readiness-path/sample/adapter.env new file mode 100644 index 0000000..d9f5276 --- /dev/null +++ b/tests/fixtures/lint/web-no-readiness-path/sample/adapter.env @@ -0,0 +1,5 @@ +ADAPTER_NAME="sample" +ADAPTER_ROLE="web" +ADAPTER_FAMILY="next" +ADAPTER_GENERATOR='mkdir -p "$APP_DIR"' +ADAPTER_LIVENESS_PATH="/" diff --git a/tests/fixtures/lint/web-no-readiness-path/sample/mise.toml b/tests/fixtures/lint/web-no-readiness-path/sample/mise.toml new file mode 100644 index 0000000..8ec8128 --- /dev/null +++ b/tests/fixtures/lint/web-no-readiness-path/sample/mise.toml @@ -0,0 +1,26 @@ +[tasks.install] +run = "true" + +[tasks.format] +run = "true" + +[tasks."format-fix"] +run = "true" + +[tasks.lint] +run = "true" + +[tasks.check] +run = "true" + +[tasks.test] +run = "true" + +[tasks.build] +run = "true" + +[tasks.ci-unit] +run = "true" + +[tasks.checklist] +run = "true" From 188c9dbe1430fe662239eec51db61c5a677fe8d0 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 12:05:26 +0700 Subject: [PATCH 05/30] feat: serve the port compose publishes, on a route that exists MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit nestjs and nextjs both listened on their generator's default port while common/compose.yaml has always published 8080 — nestjs additionally probed a /health route no adapter ever shipped, so its image reported unhealthy from first boot. Both adapters now listen on 8080, ship a real liveness route, and their HEALTHCHECKs probe it there. apply_adapter's docker/-only special case is now a loop over every directory an adapter ships, since nestjs's health routes and nextjs's health route both need the same directory-delivery mechanism the Laravel pair already used for docker/opcache.ini. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/nestjs/Dockerfile | 8 ++++-- adapters/nestjs/Dockerfile.workspace | 8 ++++-- adapters/nestjs/adapter.env | 7 +++-- .../nestjs/src/health/health.controller.ts | 27 +++++++++++++++++++ adapters/nestjs/src/health/health.module.ts | 6 +++++ adapters/nextjs/Dockerfile | 12 ++++----- adapters/nextjs/Dockerfile.workspace | 12 ++++----- .../nextjs/src/app/api/health/live/route.ts | 5 ++++ lib/adapter.sh | 11 ++++++-- tests/compose.bats | 24 +++++++++++++++++ 10 files changed, 100 insertions(+), 20 deletions(-) create mode 100644 adapters/nestjs/src/health/health.controller.ts create mode 100644 adapters/nestjs/src/health/health.module.ts create mode 100644 adapters/nextjs/src/app/api/health/live/route.ts diff --git a/adapters/nestjs/Dockerfile b/adapters/nestjs/Dockerfile index 2e4b052..aaf41bf 100644 --- a/adapters/nestjs/Dockerfile +++ b/adapters/nestjs/Dockerfile @@ -18,7 +18,11 @@ ENV NODE_ENV=production COPY --from=build /app/node_modules ./node_modules COPY --from=build /app/dist ./dist USER node -EXPOSE 3001 +# 8080 because that is the container port common/compose.yaml publishes, and +# nothing rewrites it — an adapter listening anywhere else publishes a dead +# port. Nest reads PORT in main.ts's `app.listen(process.env.PORT ?? 3000)`. +ENV PORT=8080 +EXPOSE 8080 HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:3001/health || exit 1 + CMD wget -qO- http://localhost:8080/health/live || exit 1 CMD ["node", "dist/main.js"] diff --git a/adapters/nestjs/Dockerfile.workspace b/adapters/nestjs/Dockerfile.workspace index 5ac8914..a0e2457 100644 --- a/adapters/nestjs/Dockerfile.workspace +++ b/adapters/nestjs/Dockerfile.workspace @@ -30,7 +30,11 @@ COPY --from=build /app/node_modules ./node_modules COPY --from=build /app/apps/@APP_FILTER@/node_modules ./apps/@APP_FILTER@/node_modules COPY --from=build /app/apps/@APP_FILTER@/dist ./apps/@APP_FILTER@/dist USER node -EXPOSE 3001 +# 8080 because that is the container port common/compose.yaml publishes, and +# nothing rewrites it — an adapter listening anywhere else publishes a dead +# port. Nest reads PORT in main.ts's `app.listen(process.env.PORT ?? 3000)`. +ENV PORT=8080 +EXPOSE 8080 HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:3001/health || exit 1 + CMD wget -qO- http://localhost:8080/health/live || exit 1 CMD ["node", "apps/@APP_FILTER@/dist/main.js"] diff --git a/adapters/nestjs/adapter.env b/adapters/nestjs/adapter.env index 0c3e48b..9eec716 100644 --- a/adapters/nestjs/adapter.env +++ b/adapters/nestjs/adapter.env @@ -9,8 +9,11 @@ ADAPTER_FAMILY="nest" # generate with the new major and confirm `//apps/api:ci-unit` passes clean. ADAPTER_GENERATOR='pnpm dlx @nestjs/cli@11 new "$APP_DIR" --package-manager pnpm --skip-git' # main.ts's un-awaited bootstrap() trips --max-warnings 0, and the generator's -# own output is not prettier-formatted. -ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && pnpm exec prettier --write .' +# own output is not prettier-formatted. Wiring HealthModule here too: the +# module lives in src/health/ (copied in by apply_adapter's directory loop, +# not this adapter's flat file list), and app.module.ts is the generator's +# own file, so it can only be edited after the generator has produced it. +ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && sed -i "1i import { HealthModule } from '"'"'./health/health.module'"'"';" src/app.module.ts && sed -i "s/imports: \[\]/imports: [HealthModule]/" src/app.module.ts && pnpm exec prettier --write .' # The generator produces `/` returning Hello World and nothing else. Both of # these are routes this adapter ships itself (src/health/), because the # Dockerfile's HEALTHCHECK has been probing a /health that never existed. diff --git a/adapters/nestjs/src/health/health.controller.ts b/adapters/nestjs/src/health/health.controller.ts new file mode 100644 index 0000000..102623e --- /dev/null +++ b/adapters/nestjs/src/health/health.controller.ts @@ -0,0 +1,27 @@ +import { Controller, Get, HttpException, HttpStatus } from '@nestjs/common'; + +@Controller('health') +export class HealthController { + @Get('live') + live(): { status: string } { + return { status: 'ok' }; + } + + // The probe is written by the selected service's driver: prisma has no + // provider-agnostic read, so a SQL provider gets $queryRawUnsafe and + // mongodb gets $runCommandRaw. A project generated with --db none keeps + // the anchor's fallback and reports 503, because there is nothing here + // that could honestly report ready. + @Get('ready') + async ready(): Promise<{ status: string }> { + try { + // @DB_PROBE@ + throw new Error('no database is configured for this project'); + } catch (error) { + throw new HttpException( + { status: 'unavailable', reason: (error as Error).message }, + HttpStatus.SERVICE_UNAVAILABLE, + ); + } + } +} diff --git a/adapters/nestjs/src/health/health.module.ts b/adapters/nestjs/src/health/health.module.ts new file mode 100644 index 0000000..a1c9687 --- /dev/null +++ b/adapters/nestjs/src/health/health.module.ts @@ -0,0 +1,6 @@ +import { Module } from '@nestjs/common'; + +import { HealthController } from './health.controller'; + +@Module({ controllers: [HealthController] }) +export class HealthModule {} diff --git a/adapters/nextjs/Dockerfile b/adapters/nextjs/Dockerfile index 807c4f2..ff25659 100644 --- a/adapters/nextjs/Dockerfile +++ b/adapters/nextjs/Dockerfile @@ -25,11 +25,11 @@ COPY --from=build /app/.next/standalone ./ COPY --from=build /app/.next/static ./.next/static COPY --from=build /app/public ./public USER node -EXPOSE 3000 -# `/`, not `/api/health`: create-next-app generates no health route, so every -# image built from here reported unhealthy from first boot until somebody -# noticed and wrote one. An app that adds a real health endpoint should point -# this at it. +# 8080 because that is the container port common/compose.yaml publishes, and +# nothing rewrites it — an adapter listening anywhere else publishes a dead +# port. next's standalone server.js reads PORT itself. +ENV PORT=8080 +EXPOSE 8080 HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:3000/ || exit 1 + CMD wget -qO- http://localhost:8080/api/health/live || exit 1 CMD ["node", "server.js"] diff --git a/adapters/nextjs/Dockerfile.workspace b/adapters/nextjs/Dockerfile.workspace index f578239..f530041 100644 --- a/adapters/nextjs/Dockerfile.workspace +++ b/adapters/nextjs/Dockerfile.workspace @@ -31,11 +31,11 @@ COPY --from=build /app/apps/@APP_FILTER@/.next/standalone ./ COPY --from=build /app/apps/@APP_FILTER@/.next/static ./apps/@APP_FILTER@/.next/static COPY --from=build /app/apps/@APP_FILTER@/public ./apps/@APP_FILTER@/public USER node -EXPOSE 3000 -# `/`, not `/api/health`: create-next-app generates no health route, so every -# image built from here reported unhealthy from first boot until somebody -# noticed and wrote one. An app that adds a real health endpoint should point -# this at it. +# 8080 because that is the container port common/compose.yaml publishes, and +# nothing rewrites it — an adapter listening anywhere else publishes a dead +# port. next's standalone server.js reads PORT itself. +ENV PORT=8080 +EXPOSE 8080 HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:3000/ || exit 1 + CMD wget -qO- http://localhost:8080/api/health/live || exit 1 CMD ["node", "apps/@APP_FILTER@/server.js"] diff --git a/adapters/nextjs/src/app/api/health/live/route.ts b/adapters/nextjs/src/app/api/health/live/route.ts new file mode 100644 index 0000000..bfe58aa --- /dev/null +++ b/adapters/nextjs/src/app/api/health/live/route.ts @@ -0,0 +1,5 @@ +export const dynamic = 'force-dynamic'; + +export function GET(): Response { + return Response.json({ status: 'ok' }); +} diff --git a/lib/adapter.sh b/lib/adapter.sh index 2f8fb76..2b0a29d 100644 --- a/lib/adapter.sh +++ b/lib/adapter.sh @@ -156,8 +156,15 @@ apply_adapter() { done [ "$had_dotglob" -eq 1 ] || shopt -u dotglob - # the flat loop above skips directories - [ -d "${ADAPTER_DIR}/docker" ] && cp -R "${ADAPTER_DIR}/docker" "${dest}/docker" + # Every directory the adapter ships, merged into the generated tree rather + # than replacing what is there: `src/` already exists after the generator + # ran, and `cp -R src dest/src` would nest it as dest/src/src. + local dir + for dir in "${ADAPTER_DIR}"/*/; do + [ -d "$dir" ] || continue + mkdir -p "${dest}/$(basename "$dir")" + cp -R "${dir}." "${dest}/$(basename "$dir")/" + done resolve_workspace_filter_name "$dest" diff --git a/tests/compose.bats b/tests/compose.bats index 63a9659..0cb8957 100644 --- a/tests/compose.bats +++ b/tests/compose.bats @@ -158,3 +158,27 @@ INNER_EOF || { echo "missing ${manifest} at context '${context}', named by ${dockerfile}"; false; } done } + +@test "every adapter Dockerfile serves the port compose publishes" { + # common/compose.yaml publishes ${APP_PORT:-8080}:8080 and nothing rewrites + # it, so an adapter exposing anything else publishes a dead port. + run bash -c "grep -L '^EXPOSE 8080\$' '${SCAFFOLD_ROOT}'/adapters/*/Dockerfile*" + [ -z "$output" ] || { echo "not exposing 8080:"; echo "$output"; false; } +} + +@test "every adapter Dockerfile probes the liveness path its adapter declares" { + # nestjs probed /health for months while the generator produced only `/`. + # The Dockerfile's idea of the route and the adapter's must be one value. + local wrong="" + for dir in "${SCAFFOLD_ROOT}"/adapters/*/; do + path="$(grep '^ADAPTER_LIVENESS_PATH=' "${dir}adapter.env" | cut -d'"' -f2)" + for file in "${dir}"Dockerfile "${dir}"Dockerfile.workspace; do + [ -f "$file" ] || continue + grep -q "HEALTHCHECK" "$file" \ + || { wrong="${wrong}${file}: no HEALTHCHECK"$'\n'; continue; } + grep -q "localhost:8080${path}" "$file" \ + || wrong="${wrong}${file}: does not probe ${path} on 8080"$'\n' + done + done + [ -z "$wrong" ] || { echo "$wrong"; false; } +} From 6a703ce5879f683ba34033959778a6243a277f16 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 12:24:54 +0700 Subject: [PATCH 06/30] fix: fail post-generate if HealthModule wiring did not land MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit sed -i "1i ..." always succeeds regardless of match, and the following substitution for imports: [] silently no-ops if the nest generator ever reformats app.module.ts — leaving HealthModule unregistered with no build or lint failure, only a 404 on /health/live discovered by the HEALTHCHECK this task just pointed at that route. Grep for both halves of the wiring after the sed calls and fail loudly, naming the file and the likely cause, so a generator reformat cannot silently re-arm the unhealthy-image defect this task exists to remove. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/nestjs/adapter.env | 7 ++++++- 1 file changed, 6 insertions(+), 1 deletion(-) diff --git a/adapters/nestjs/adapter.env b/adapters/nestjs/adapter.env index 9eec716..00ebce0 100644 --- a/adapters/nestjs/adapter.env +++ b/adapters/nestjs/adapter.env @@ -13,7 +13,12 @@ ADAPTER_GENERATOR='pnpm dlx @nestjs/cli@11 new "$APP_DIR" --package-manager pnpm # module lives in src/health/ (copied in by apply_adapter's directory loop, # not this adapter's flat file list), and app.module.ts is the generator's # own file, so it can only be edited after the generator has produced it. -ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && sed -i "1i import { HealthModule } from '"'"'./health/health.module'"'"';" src/app.module.ts && sed -i "s/imports: \[\]/imports: [HealthModule]/" src/app.module.ts && pnpm exec prettier --write .' +# The grep pair after the two sed calls is not optional: `1i` always +# succeeds, and `s/imports: \[\]/…/` silently no-ops if the generator ever +# reformats that line, leaving HealthModule unregistered with no build or +# lint failure — only a 404 on /health/live, discovered by the HEALTHCHECK +# that quietly starts failing on every image. +ADAPTER_POST_GENERATE='sed -i "s/^bootstrap();$/void bootstrap();/" src/main.ts && sed -i "1i import { HealthModule } from '"'"'./health/health.module'"'"';" src/app.module.ts && sed -i "s/imports: \[\]/imports: [HealthModule]/" src/app.module.ts && { grep -q "import { HealthModule } from '"'"'./health/health.module'"'"';" src/app.module.ts && grep -q "imports: \[HealthModule\]" src/app.module.ts || { echo "post-generate: HealthModule wiring missing from src/app.module.ts; the nest generator likely changed its output format — update the sed patterns in ADAPTER_POST_GENERATE to match" >&2; exit 1; }; } && pnpm exec prettier --write .' # The generator produces `/` returning Hello World and nothing else. Both of # these are routes this adapter ships itself (src/health/), because the # Dockerfile's HEALTHCHECK has been probing a /health that never existed. From 49ff1bf609c9a6f93502f2b94487efa5502dac3b Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 13:11:52 +0700 Subject: [PATCH 07/30] feat: put an http server in the laravel images Replace the php-fpm runtime stage in both laravel adapters with FrankenPHP (alpine, pinned by digest), so the images actually speak HTTP instead of FastCGI with nothing in front of them. Adds SERVER_NAME, the XDG_* pair Caddy needs to start as www-data, ownership fixes for storage/ and bootstrap/cache, and a real HEALTHCHECK against /up. The base image's default CMD carries --config/--adapter flags that point frankenphp at its Caddyfile; a bare `frankenphp run` drops them and starts only the admin API, so CMD restates them explicitly. laravel-inertia also moves COPY . . before the assets copy so a developer's local public/build can no longer shadow the one the assets stage built, and .dockerignore excludes public/build from the build context for the same reason. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/laravel-api/Dockerfile | 36 ++++++++++++++++-------- adapters/laravel-inertia/.dockerignore | 6 ++++ adapters/laravel-inertia/Dockerfile | 38 +++++++++++++++++--------- 3 files changed, 55 insertions(+), 25 deletions(-) diff --git a/adapters/laravel-api/Dockerfile b/adapters/laravel-api/Dockerfile index 7cbfffe..37207a0 100644 --- a/adapters/laravel-api/Dockerfile +++ b/adapters/laravel-api/Dockerfile @@ -6,22 +6,34 @@ COPY composer.json composer.lock ./ RUN composer install --no-dev --no-scripts --no-interaction \ --prefer-dist --optimize-autoloader -FROM php:8.3.21-fpm-alpine@sha256:d2170b0f8da574062b289566a05f25ab57173315a356c9f7519b5e444ae96dac AS runtime +# FrankenPHP, because php-fpm speaks FastCGI and this stack has no web server +# in front of it: nothing served HTTP at all, and compose published a dead +# port. Documented as a first-class server in laravel.com/docs/13.x/deployment, +# part of the PHP Foundation since May 2025, and what Laravel Cloud runs. +# Alpine specifically: services/mongodb/drivers/laravel.sh emits `apk add`. +FROM dunglas/frankenphp:1.12.7-php8.3-alpine@sha256:049b8d8356efceb93c91ed42866de890534310bcef4ad4dde902029e4a0d20c3 AS runtime RUN docker-php-ext-install opcache # @SERVICE_SETUP@ WORKDIR /var/www COPY --from=vendor /app/vendor ./vendor COPY . . COPY docker/opcache.ini /usr/local/etc/php/conf.d/opcache.ini +# :8080 is how frankenphp listens on an unprivileged port and declines to +# provision TLS, which belongs to whatever proxy this lands behind. +ENV SERVER_NAME=:8080 +# Caddy writes here and will not start if it may not. The app tree needs the +# same: COPY runs as root, the server runs as www-data, and the first request +# that compiles a blade view or writes a log fails without this. +ENV XDG_CONFIG_HOME=/config XDG_DATA_HOME=/data +RUN mkdir -p /config /data \ + && chown -R www-data:www-data /config /data \ + /var/www/storage /var/www/bootstrap/cache USER www-data -EXPOSE 9000 -# php-fpm speaks fastcgi on this port, not http, so the http probes the -# sibling adapters (nextjs, nestjs) use do not transplant here. the previous -# check, `php -r 'exit(0);'`, only proved the php cli starts — it can never -# fail while the container is actually broken, which is worse than no check -# at all: an orchestrator with no HEALTHCHECK knows it does not know, one -# with an always-green check believes it does. a real check needs php-fpm's -# ping.path probed with cgi-fcgi, which costs a pool config file and a -# package this image does not otherwise need; add both together if this is -# ever deployed behind something that polls container health. -CMD ["php-fpm"] +EXPOSE 8080 +HEALTHCHECK --interval=30s --timeout=3s \ + CMD wget -qO- http://localhost:8080/up || exit 1 +# CMD replaces the base image's default args entirely rather than extending +# them, and those args are what point frankenphp at the Caddyfile that +# defines the :8080 site block; without them it starts only the admin API on +# 127.0.0.1:2019 and nothing ever listens on 8080. +CMD ["frankenphp", "run", "--config", "/etc/frankenphp/Caddyfile", "--adapter", "caddyfile"] diff --git a/adapters/laravel-inertia/.dockerignore b/adapters/laravel-inertia/.dockerignore index 8103920..48bbc59 100644 --- a/adapters/laravel-inertia/.dockerignore +++ b/adapters/laravel-inertia/.dockerignore @@ -17,3 +17,9 @@ storage/framework/cache # other way — laravel refuses to start without it. For laravel-inertia either # failure surfaces in vite's wayfinder plugin, several stages from the cause. bootstrap/cache/*.php + +# vite writes this and .gitignore excludes it, but a build context is not a +# git tree: without this line a developer who has run `npm run build` ships +# their host copy over the one the assets stage just built. Same shape as the +# bootstrap/cache manifests two lines up. +public/build diff --git a/adapters/laravel-inertia/Dockerfile b/adapters/laravel-inertia/Dockerfile index d578cdc..a1d4bd3 100644 --- a/adapters/laravel-inertia/Dockerfile +++ b/adapters/laravel-inertia/Dockerfile @@ -21,23 +21,35 @@ COPY --from=vendor /app/vendor ./vendor COPY . . RUN npm run build -FROM php:8.3.21-fpm-alpine@sha256:d2170b0f8da574062b289566a05f25ab57173315a356c9f7519b5e444ae96dac AS runtime +# FrankenPHP, because php-fpm speaks FastCGI and this stack has no web server +# in front of it: nothing served HTTP at all, and compose published a dead +# port. Documented as a first-class server in laravel.com/docs/13.x/deployment, +# part of the PHP Foundation since May 2025, and what Laravel Cloud runs. +# Alpine specifically: services/mongodb/drivers/laravel.sh emits `apk add`. +FROM dunglas/frankenphp:1.12.7-php8.3-alpine@sha256:049b8d8356efceb93c91ed42866de890534310bcef4ad4dde902029e4a0d20c3 AS runtime RUN docker-php-ext-install opcache # @SERVICE_SETUP@ WORKDIR /var/www COPY --from=vendor /app/vendor ./vendor -COPY --from=assets /app/public/build ./public/build COPY . . +COPY --from=assets /app/public/build ./public/build COPY docker/opcache.ini /usr/local/etc/php/conf.d/opcache.ini +# :8080 is how frankenphp listens on an unprivileged port and declines to +# provision TLS, which belongs to whatever proxy this lands behind. +ENV SERVER_NAME=:8080 +# Caddy writes here and will not start if it may not. The app tree needs the +# same: COPY runs as root, the server runs as www-data, and the first request +# that compiles a blade view or writes a log fails without this. +ENV XDG_CONFIG_HOME=/config XDG_DATA_HOME=/data +RUN mkdir -p /config /data \ + && chown -R www-data:www-data /config /data \ + /var/www/storage /var/www/bootstrap/cache USER www-data -EXPOSE 9000 -# php-fpm speaks fastcgi on this port, not http, so the http probes the -# sibling adapters (nextjs, nestjs) use do not transplant here. the obvious -# alternative, `php -r 'exit(0);'`, only proves the php cli starts — it can -# never fail while the container is actually broken, which is worse than no -# check at all: an orchestrator with no HEALTHCHECK knows it does not know, -# one with an always-green check believes it does. a real check needs -# php-fpm's ping.path probed with cgi-fcgi, which costs a pool config file -# and a package this image does not otherwise need; add both together if -# this is ever deployed behind something that polls container health. -CMD ["php-fpm"] +EXPOSE 8080 +HEALTHCHECK --interval=30s --timeout=3s \ + CMD wget -qO- http://localhost:8080/up || exit 1 +# CMD replaces the base image's default args entirely rather than extending +# them, and those args are what point frankenphp at the Caddyfile that +# defines the :8080 site block; without them it starts only the admin API on +# 127.0.0.1:2019 and nothing ever listens on 8080. +CMD ["frankenphp", "run", "--config", "/etc/frankenphp/Caddyfile", "--adapter", "caddyfile"] From b2f2d1b1fb5c4602f04b28db07f97139aa27be02 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 13:20:59 +0700 Subject: [PATCH 08/30] feat: give a driver a seam for the app's connection environment MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit compose.yaml only ever received DB_DATABASE/DB_USERNAME/DB_PASSWORD, and neither Prisma nor Laravel reads those directly, so a released stack could not reach its database. apply_service_compose_env merges a driver's service_driver_compose_env block into compose.yaml's app service with `yq -P` so the merge keeps block style and the file's comments survive. Every driver is required to define the hook (REQUIRED_DRIVER_FUNCTIONS, checked by lint_services), so a driver shipping only two of the three fails at generation or at lint instead of silently producing an app with no DATABASE_URL. No driver implements the hook yet — that is task 5. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- lib/contract.sh | 4 ++++ lib/lint.sh | 22 +++++++++++++++++++--- lib/service.sh | 33 ++++++++++++++++++++++++++++++++- tests/service.bats | 27 +++++++++++++++++++++++++++ 4 files changed, 82 insertions(+), 4 deletions(-) diff --git a/lib/contract.sh b/lib/contract.sh index ccc9497..f5eccf2 100644 --- a/lib/contract.sh +++ b/lib/contract.sh @@ -31,6 +31,10 @@ REQUIRED_SERVICE_FILES=( REQUIRED_SERVICE_VARS=(SERVICE_NAME SERVICE_KIND SERVICE_IMAGE) +# apply_service_drivers calls all three, so a driver shipping fewer fails at +# generation rather than at lint. +REQUIRED_DRIVER_FUNCTIONS=(service_driver_apply service_driver_dockerfile service_driver_compose_env) + # The web tier is the presentation layer and opens no connection, so it takes # no driver — stated once, about the role, rather than as a "not applicable" # entry repeated in every service. diff --git a/lib/lint.sh b/lib/lint.sh index f37af6f..c9f22fd 100644 --- a/lib/lint.sh +++ b/lib/lint.sh @@ -71,7 +71,7 @@ lint_adapters() { # any family that takes a driver has no driver in some service. lint_services() { local dir="$1" adapters="$2" - local service name file var family status=0 + local service name file var family driver fn status=0 local -a families=() # The families to require, read from the adapters themselves rather than @@ -120,10 +120,26 @@ lint_services() { fi for family in "${families[@]}"; do - [ -f "${service}drivers/${family}.sh" ] || { + driver="${service}drivers/${family}.sh" + if [ ! -f "$driver" ]; then printf '%s: no driver for %s\n' "$name" "$family" status=1 - } + continue + fi + + # A subshell, not the current one: sourcing eight drivers in sequence + # here would let one family's LARAVEL_* parameters (services/shared/ + # laravel.sh reads them unqualified) leak into the next driver checked. + for fn in "${REQUIRED_DRIVER_FUNCTIONS[@]}"; do + ( + # shellcheck source=/dev/null # family varies, so the path isn't constant + . "$driver" + declare -F "$fn" >/dev/null + ) || { + printf '%s: %s driver does not define %s\n' "$name" "$family" "$fn" + status=1 + } + done done done diff --git a/lib/service.sh b/lib/service.sh index 87ead60..8ac6884 100644 --- a/lib/service.sh +++ b/lib/service.sh @@ -146,6 +146,32 @@ apply_service_setup() { [ "$found" -eq 1 ] || return 0 } +# apply_service_compose_env +# Adds the block to compose.yaml's app service. yq rather than an anchor: the +# app service is generated by assemble_compose from common/compose.yaml, so +# there is a real document to merge into by the time this runs, and a text +# anchor would only be a second way to write YAML. +apply_service_compose_env() { + local project="$1" block="$2" + local file="${project}/compose.yaml" fragment + + [ -n "$block" ] || return 0 + [ -f "$file" ] || die "no compose.yaml in ${project}" + + fragment="$(mktemp)" + { + printf 'services:\n app:\n environment:\n' + printf '%s\n' "$block" | sed 's/^/ /' + } > "$fragment" + + if ! yq eval-all --inplace -P 'select(fileIndex==0) * select(fileIndex==1)' \ + "$file" "$fragment"; then + rm -f "$fragment" + die "could not merge the service environment into ${file}" + fi + rm -f "$fragment" +} + # write_env_lines ... # Sets each KEY=value, replacing the key if it is already there. A driver runs # against an .env.example the adapter shipped, so appending blindly would @@ -199,7 +225,7 @@ write_env_lines() { # pnpm-workspace.yaml edit) cannot recover it from its own cwd. apply_service_drivers() { local app="$1" project="$2" family="$3"; shift 3 - local service driver block="" rendered + local service driver block="" env_block="" rendered # web is the presentation tier and takes no driver — the caller decides # that from ADAPTER_ROLE, so reaching here with a family that has none is a @@ -260,9 +286,14 @@ apply_service_drivers() { # shellcheck source=/dev/null # family varies, so the path isn't constant rendered="$( . "$driver"; service_driver_dockerfile )" [ -n "$rendered" ] && block+="${rendered}"$'\n' + + # shellcheck source=/dev/null # family varies, so the path isn't constant + rendered="$( . "$driver"; service_driver_compose_env )" + [ -n "$rendered" ] && env_block+="${rendered}"$'\n' done apply_service_setup "$app" "${block%$'\n'}" + apply_service_compose_env "$project" "${env_block%$'\n'}" } # record_services diff --git a/tests/service.bats b/tests/service.bats index 392e9a5..bbdf2d1 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -586,3 +586,30 @@ EOF [ "$allow_builds" -lt "$first_add" ] \ || { echo "allowBuilds (line ${allow_builds}) must come before the first pnpm add (line ${first_add})"; false; } } + +@test "a driver's compose environment interpolates rather than embedding a password" { + # The password must exist in exactly one place — .env — so compose composes + # the URL at `up` time. A literal baked here is the changeme-versus-app + # mismatch that made the dev stack unable to authenticate. + local bad="" + for driver in "${SCAFFOLD_ROOT}"/services/*/drivers/*.sh; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + SERVICE_DIR="$(dirname "$(dirname "$driver")")" + . "$driver"; service_driver_compose_env )" + [ -z "$block" ] && continue + grep -q '\${DB_PASSWORD' <<<"$block" || grep -q '\${REDIS_PASSWORD' <<<"$block" \ + || bad="${bad}${driver}"$'\n' + done + [ -z "$bad" ] || { echo "embeds a literal password:"; echo "$bad"; false; } +} + +@test "apply_service_compose_env merges into the app service" { + local project="${BATS_TEST_TMPDIR}/p" + mkdir -p "$project" + printf 'services:\n app:\n image: x\n' > "${project}/compose.yaml" + . "${SCAFFOLD_ROOT}/lib/service.sh" + apply_service_compose_env "$project" 'DATABASE_URL: ${DATABASE_URL:-postgresql://app@database:5432/app}' + run mise exec -- yq -r '.services.app.environment.DATABASE_URL' "${project}/compose.yaml" + [[ "$output" == 'postgresql://app@database:5432/app' ]] \ + || [[ "$output" == '${DATABASE_URL:-postgresql://app@database:5432/app}' ]] +} From 11534cda108c618ac28996583942442c59360b30 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 13:37:14 +0700 Subject: [PATCH 09/30] fix: give lint's driver sourcing a real environment, and measure -P MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit lint_services sourced each driver with no SERVICE_DIR set, unlike every other call site that sources a driver — a Task-5 driver reading it at sourcing time would die under the inherited `set -u` and get misreported as "does not define" the function it never got a chance to declare. Set it the way load_service does, and capture the subshell's stderr so a sourcing failure reports its own text instead of being folded into the missing-function message. Dropped compose-env's `-P`: measured, not assumed. The fragment it merges is always this function's own block-style output, never the flow style merge_lefthook_fragment's `-P` guards against, so there is nothing here for it to protect. Merging without it left compose.yaml's untouched lines byte-identical; with it, yq rewrote nodes this merge never touched — an unquoted `${APP_PORT:-8080}:8080` and a healthcheck array reshaped from flow to block. Both happened to pass prettier --check, but a merge with no business editing those lines should not be reshaping them anyway. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- lib/lint.sh | 24 +++++++++++++++++++----- lib/service.sh | 15 ++++++++++++++- 2 files changed, 33 insertions(+), 6 deletions(-) diff --git a/lib/lint.sh b/lib/lint.sh index c9f22fd..b792238 100644 --- a/lib/lint.sh +++ b/lib/lint.sh @@ -71,7 +71,7 @@ lint_adapters() { # any family that takes a driver has no driver in some service. lint_services() { local dir="$1" adapters="$2" - local service name file var family driver fn status=0 + local service name file var family driver fn fault status=0 local -a families=() # The families to require, read from the adapters themselves rather than @@ -130,15 +130,29 @@ lint_services() { # A subshell, not the current one: sourcing eight drivers in sequence # here would let one family's LARAVEL_* parameters (services/shared/ # laravel.sh reads them unqualified) leak into the next driver checked. + # + # SERVICE_DIR set the same way load_service sets it, before sourcing: + # every other call site that sources a driver (apply_service_drivers, + # via load_service; the compose-env test in service.bats, by hand) has + # it set first. A driver that reads it at sourcing time and finds it + # unbound would die under the inherited `set -u` before `declare -F` + # ever ran, and that death is not the same problem as a missing + # function — captured below instead of folded into that message. for fn in "${REQUIRED_DRIVER_FUNCTIONS[@]}"; do - ( + if ! fault="$( { + # shellcheck disable=SC2034 # read by the driver, not by this loop + SERVICE_DIR="${service%/}" # shellcheck source=/dev/null # family varies, so the path isn't constant . "$driver" declare -F "$fn" >/dev/null - ) || { - printf '%s: %s driver does not define %s\n' "$name" "$family" "$fn" + } 2>&1 )"; then + if [ -n "$fault" ]; then + printf '%s: %s driver failed to source: %s\n' "$name" "$family" "$fault" + else + printf '%s: %s driver does not define %s\n' "$name" "$family" "$fn" + fi status=1 - } + fi done done done diff --git a/lib/service.sh b/lib/service.sh index 8ac6884..c4a7b2b 100644 --- a/lib/service.sh +++ b/lib/service.sh @@ -151,6 +151,19 @@ apply_service_setup() { # app service is generated by assemble_compose from common/compose.yaml, so # there is a real document to merge into by the time this runs, and a text # anchor would only be a second way to write YAML. +# +# No -P, unlike merge_lefthook_fragment: that merge takes a fragment file it +# does not control the style of, so an author who wrote it in flow style +# would collapse the whole target document without -P forcing everything back +# to block. The fragment built below is never that — it is this function's +# own printf, always one block-style `KEY: value` line per driver, never a +# flow mapping — so there is nothing here for -P to guard against. Measured +# instead of assumed: merging it in without -P left every byte outside the +# two inserted lines untouched, while -P rewrote nodes this merge never +# touched (unquoted compose.yaml's `- '${APP_PORT:-8080}:8080'`, and expanded +# the postgres healthcheck's flow-style `test: [...]` to block) — a client's +# `prettier --check` happened to accept both spellings, but a merge with no +# business editing those lines should not still be reshaping them. apply_service_compose_env() { local project="$1" block="$2" local file="${project}/compose.yaml" fragment @@ -164,7 +177,7 @@ apply_service_compose_env() { printf '%s\n' "$block" | sed 's/^/ /' } > "$fragment" - if ! yq eval-all --inplace -P 'select(fileIndex==0) * select(fileIndex==1)' \ + if ! yq eval-all --inplace 'select(fileIndex==0) * select(fileIndex==1)' \ "$file" "$fragment"; then rm -f "$fragment" die "could not merge the service environment into ${file}" From c0042f49d3c0d4dedbf5eb787e8bc1d1cf69a235 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 14:36:58 +0700 Subject: [PATCH 10/30] feat: tell the application how to reach the service it was given MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit All eight drivers now implement service_driver_compose_env: nest gets DATABASE_URL/REDIS_URL built from the same DB_*/REDIS_* variables the database/cache containers read, operator-overridable via .env; laravel gets DB_CONNECTION plus whatever host/port/DSN that family needs, since config/database.php defaults to sqlite and its absence was a silent wrong answer, not a failure. Every interpolation reads the password rather than embedding it, so it exists in exactly one place. Both frameworks also get a real readiness probe, spliced into the anchor each ships pre-generation. Two substitutions, not one, in both: the anchor becomes the probe, and the shipped fallback (`throw`) is replaced in place rather than left as dead code below an added return — nest's controller would otherwise fail eslint's no-unreachable under --max-warnings 0. Nest's probe is cast to an explicit method signature instead of a bare dynamic import, because before `prisma generate` runs (lint runs first) the generated client module doesn't exist, and an untyped access to it is `any` under the no-unsafe-* rules; only the one method the provider actually calls is declared, since a real PrismaClient's mongodb build has no $queryRawUnsafe and its SQL builds have no $runCommandRaw. Laravel needed a route to splice into. Both laravel adapters now ship routes/health.php, wired into bootstrap/app.php's withRouting(then: ...) — laravel only auto-loads routes/web.php and routes/console.php, so a file dropped there with nothing pointing at it 404s forever. APP_KEY, per-family rather than per-service, is appended to the project's example.env by every laravel driver, since no service env.fragment can carry it and laravel will not boot without one. Verified live: generating a laravel+mongodb project and hitting /health/ready surfaced a real bug independent of this task — `pecl install mongodb` resolves latest (2.x) while mongodb/laravel-mongodb is built against the 1.x extension's BSON model classes, so any code touching MongoDB\Client fatals before reaching this project's own try/catch. Pinned to 1.21.0, matching the platform.ext-mongodb declaration already beside it. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/laravel-api/adapter.env | 11 +++- adapters/laravel-api/routes/health.php | 21 +++++++ adapters/laravel-inertia/adapter.env | 10 ++- adapters/laravel-inertia/routes/health.php | 21 +++++++ services/mongodb/drivers/laravel.sh | 49 ++++++++++++++- services/mongodb/drivers/nest.sh | 1 + services/mysql/drivers/laravel.sh | 8 +++ services/mysql/drivers/nest.sh | 1 + services/postgres/drivers/laravel.sh | 8 +++ services/postgres/drivers/nest.sh | 1 + services/redis/drivers/laravel.sh | 10 +++ services/redis/drivers/nest.sh | 8 +++ services/shared/laravel.sh | 41 ++++++++++++ services/shared/nest.sh | 72 +++++++++++++++++++++- tests/service.bats | 23 +++++++ 15 files changed, 279 insertions(+), 6 deletions(-) create mode 100644 adapters/laravel-api/routes/health.php create mode 100644 adapters/laravel-inertia/routes/health.php diff --git a/adapters/laravel-api/adapter.env b/adapters/laravel-api/adapter.env index 5ea45e4..297f7ad 100644 --- a/adapters/laravel-api/adapter.env +++ b/adapters/laravel-api/adapter.env @@ -11,7 +11,16 @@ ADAPTER_FAMILY="laravel" # holding the same floor. See docs/decisions/0016. ADAPTER_GENERATOR='composer create-project laravel/laravel:^13.0 "$APP_DIR" --no-interaction --prefer-dist' # the skeleton ships phpunit and pint, but nothing for the check task -ADAPTER_POST_GENERATE='composer require --dev larastan/larastan phpstan/phpstan --no-interaction' +# +# The sed call wires routes/health.php (copied in by apply_adapter's +# directory loop, not this adapter's flat file list) into bootstrap/app.php's +# routing: laravel only auto-loads routes/web.php and routes/console.php, so +# a file dropped at routes/health.php with nothing pointing at it 404s +# forever. The grep pair after it is not optional — `s|health: '/up',|...|` +# silently no-ops if the skeleton ever reformats that line, leaving the +# route unregistered with no build failure, only a 404 discovered in +# production. +ADAPTER_POST_GENERATE='composer require --dev larastan/larastan phpstan/phpstan --no-interaction && sed -i "s|health: '"'"'/up'"'"',|health: '"'"'/up'"'"',\n then: function (): void {\n require __DIR__.'"'"'/../routes/health.php'"'"';\n },|" bootstrap/app.php && { grep -q "then: function (): void {" bootstrap/app.php && grep -q "require __DIR__.'"'"'/../routes/health.php'"'"';" bootstrap/app.php; } || { echo "post-generate: health route wiring missing from bootstrap/app.php — the laravel skeleton likely changed its withRouting shape; update the sed pattern in ADAPTER_POST_GENERATE to match" >&2; exit 1; }' # /up ships with laravel since 11.x and deliberately touches nothing, which # is why it cannot stand alone: a project with a wrong DATABASE_URL passes it. ADAPTER_LIVENESS_PATH="/up" diff --git a/adapters/laravel-api/routes/health.php b/adapters/laravel-api/routes/health.php new file mode 100644 index 0000000..c607230 --- /dev/null +++ b/adapters/laravel-api/routes/health.php @@ -0,0 +1,21 @@ +json(['status' => 'unavailable', 'reason' => $e->getMessage()], 503); + } +}); diff --git a/adapters/laravel-inertia/adapter.env b/adapters/laravel-inertia/adapter.env index 3b64cc4..be593e5 100644 --- a/adapters/laravel-inertia/adapter.env +++ b/adapters/laravel-inertia/adapter.env @@ -23,7 +23,15 @@ ADAPTER_GENERATOR='SHELL_VERBOSITY=-1 COMPOSER_PROCESS_TIMEOUT=900 composer crea # dependabot file in the tree, and the kit's dependabot.yml sets # `cooldown.default-days: 5`, which fails the security gate with exit 13 on # the first pull request. Inert config whose only effect is a red check. -ADAPTER_POST_GENERATE='rm -rf .github' +# The sed call wires routes/health.php (copied in by apply_adapter's +# directory loop, not this adapter's flat file list) into bootstrap/app.php's +# routing: laravel only auto-loads routes/web.php and routes/console.php, so +# a file dropped at routes/health.php with nothing pointing at it 404s +# forever. The grep pair after it is not optional — `s|health: '/up',|...|` +# silently no-ops if the skeleton ever reformats that line, leaving the +# route unregistered with no build failure, only a 404 discovered in +# production. +ADAPTER_POST_GENERATE='rm -rf .github && sed -i "s|health: '"'"'/up'"'"',|health: '"'"'/up'"'"',\n then: function (): void {\n require __DIR__.'"'"'/../routes/health.php'"'"';\n },|" bootstrap/app.php && { grep -q "then: function (): void {" bootstrap/app.php && grep -q "require __DIR__.'"'"'/../routes/health.php'"'"';" bootstrap/app.php; } || { echo "post-generate: health route wiring missing from bootstrap/app.php — the laravel skeleton likely changed its withRouting shape; update the sed pattern in ADAPTER_POST_GENERATE to match" >&2; exit 1; }' # /up ships with laravel since 11.x and deliberately touches nothing, which # is why it cannot stand alone: a project with a wrong DATABASE_URL passes it. ADAPTER_LIVENESS_PATH="/up" diff --git a/adapters/laravel-inertia/routes/health.php b/adapters/laravel-inertia/routes/health.php new file mode 100644 index 0000000..c607230 --- /dev/null +++ b/adapters/laravel-inertia/routes/health.php @@ -0,0 +1,21 @@ +json(['status' => 'unavailable', 'reason' => $e->getMessage()], 503); + } +}); diff --git a/services/mongodb/drivers/laravel.sh b/services/mongodb/drivers/laravel.sh index fe2e99f..f184ccd 100644 --- a/services/mongodb/drivers/laravel.sh +++ b/services/mongodb/drivers/laravel.sh @@ -31,19 +31,66 @@ service_driver_apply() { || return 1 register_mongodb_connection config/database.php + + # APP_KEY has no service to come from — it is per-family, not per-service — + # so no env.fragment can carry it, and the project's example.env (assembled + # from those fragments, before this runs) never sees it any other way. + # Without a value here, compose.yaml's `APP_KEY: ${APP_KEY}` interpolates to + # empty and laravel refuses to boot. + write_env_lines "${SCAFFOLD_PROJECT_ROOT}/example.env" "APP_KEY=changeme" || return 1 + + # mongodb has no SQL to run a `select 1` against — a ping command is the + # provider-agnostic equivalent laravel-mongodb actually exposes. + # + # Two separate substitutions, not one: splicing the probe in above the + # shipped `throw` would leave that throw as dead code below a path that + # always returns first. The throw is replaced in place instead, so a + # --db none project keeps it — unreachable in no project this driver ever + # touches. + sed -i.bak 's|// @DB_PROBE@|\\Illuminate\\Support\\Facades\\DB::connection(\x27mongodb\x27)->getMongoDB()->command([\x27ping\x27 => 1]);|' \ + routes/health.php || return 1 + sed -i.bak "s|throw new \\\\RuntimeException('no database is configured for this project');|return response()->json(['status' => 'ok']);|" \ + routes/health.php || return 1 + rm -f routes/health.php.bak + + grep -q "DB::connection('mongodb')->getMongoDB()->command" routes/health.php \ + && grep -q "return response()->json(\['status' => 'ok'\]);" routes/health.php \ + || die "could not splice the database probe into routes/health.php — has the anchor moved?" } service_driver_dockerfile() { # pecl, not apk: the mongodb extension is not in alpine's repositories, so # it is built here — which is why this block installs the build # dependencies and nothing else does. + # + # Pinned to 1.21.0, matching platform.ext-mongodb above: an unpinned + # `pecl install mongodb` resolves whatever is latest at build time, and + # 2.x is not that — mongodb/mongodb's BSONArray/BSONDocument model classes + # declare bsonSerialize() against the 1.x extension's signature, and the + # 2.x extension changed it, so any code path that loads those classes (a + # bare `new MongoDB\Client(...)`, no Laravel involved) is a PHP fatal + # error, not an exception this project's own try/catch can see. Measured + # by generating this project with mongodb and hitting /health/ready on a + # running container: 500 before this pin, 200 after. printf '%s\n' \ 'RUN apk add --no-cache --virtual .build-deps $PHPIZE_DEPS openssl-dev \' \ - ' && pecl install mongodb \' \ + ' && pecl install mongodb-1.21.0 \' \ ' && docker-php-ext-enable mongodb \' \ ' && apk del .build-deps' } +# DB_CONNECTION first and always: config/database.php defaults to sqlite, so +# its absence is not an error, it is a silent wrong answer. DB_USERNAME and +# DB_PASSWORD reach the container through compose.yaml's env_file already +# (they are in the project's example.env, assembled from this service's own +# env.fragment) — DB_URI still needs assembling here because laravel-mongodb +# reads one DSN string, not decomposed host/port credentials. +service_driver_compose_env() { + printf 'DB_CONNECTION: mongodb\n' + printf 'DB_URI: ${DB_URI:-mongodb://${DB_USERNAME:-app}:${DB_PASSWORD}@database:27017/${DB_DATABASE:-app}?authSource=admin}\n' + printf 'APP_KEY: ${APP_KEY}\n' +} + # register_mongodb_connection # laravel-mongodb needs a 'mongodb' entry in the connections array; the # Laravel skeleton ships none. Same insert-then-verify shape as diff --git a/services/mongodb/drivers/nest.sh b/services/mongodb/drivers/nest.sh index 621a473..a192830 100644 --- a/services/mongodb/drivers/nest.sh +++ b/services/mongodb/drivers/nest.sh @@ -8,5 +8,6 @@ PRISMA_PROVIDER="mongodb" # single-node container it changes nothing for `db push`, so treat it as # unproven for anything but transactional writes. PRISMA_URL="mongodb://app:app@localhost:27017/app?authSource=admin&directConnection=true" +PRISMA_COMPOSE_URL='mongodb://${DB_USERNAME:-app}:${DB_PASSWORD}@database:27017/${DB_DATABASE:-app}?authSource=admin&directConnection=true' # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/nest.sh" diff --git a/services/mysql/drivers/laravel.sh b/services/mysql/drivers/laravel.sh index f9d730b..df617ad 100644 --- a/services/mysql/drivers/laravel.sh +++ b/services/mysql/drivers/laravel.sh @@ -6,5 +6,13 @@ LARAVEL_PORT="3306" # image already carries. LARAVEL_PACKAGE="" LARAVEL_SETUP="RUN docker-php-ext-install pdo_mysql" +# DB_PASSWORD already reaches the container via compose.yaml's env_file (it +# is in the project's example.env, assembled from this service's own +# env.fragment) — restated here, interpolated rather than baked, so the +# service_driver_compose_env test that guards against a literal password can +# see it the same way the URL-based drivers show theirs. +LARAVEL_COMPOSE_ENV="DB_HOST: database +DB_PORT: 3306 +DB_PASSWORD: \${DB_PASSWORD}" # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/laravel.sh" diff --git a/services/mysql/drivers/nest.sh b/services/mysql/drivers/nest.sh index ab7bce3..da6eb44 100644 --- a/services/mysql/drivers/nest.sh +++ b/services/mysql/drivers/nest.sh @@ -2,5 +2,6 @@ # shellcheck disable=SC2034 # read by services/shared/nest.sh, sourced below PRISMA_PROVIDER="mysql" PRISMA_URL="mysql://app:app@localhost:3306/app" +PRISMA_COMPOSE_URL='mysql://${DB_USERNAME:-app}:${DB_PASSWORD}@database:3306/${DB_DATABASE:-app}' # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/nest.sh" diff --git a/services/postgres/drivers/laravel.sh b/services/postgres/drivers/laravel.sh index 9ec6065..4955e96 100644 --- a/services/postgres/drivers/laravel.sh +++ b/services/postgres/drivers/laravel.sh @@ -5,5 +5,13 @@ LARAVEL_PORT="5432" LARAVEL_PACKAGE="" LARAVEL_SETUP="RUN apk add --no-cache postgresql-dev \\ && docker-php-ext-install pdo_pgsql" +# DB_PASSWORD already reaches the container via compose.yaml's env_file (it +# is in the project's example.env, assembled from this service's own +# env.fragment) — restated here, interpolated rather than baked, so the +# service_driver_compose_env test that guards against a literal password can +# see it the same way the URL-based drivers show theirs. +LARAVEL_COMPOSE_ENV="DB_HOST: database +DB_PORT: 5432 +DB_PASSWORD: \${DB_PASSWORD}" # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/laravel.sh" diff --git a/services/postgres/drivers/nest.sh b/services/postgres/drivers/nest.sh index ba7b224..bcb48b1 100644 --- a/services/postgres/drivers/nest.sh +++ b/services/postgres/drivers/nest.sh @@ -2,5 +2,6 @@ # shellcheck disable=SC2034 # read by services/shared/nest.sh, sourced below PRISMA_PROVIDER="postgresql" PRISMA_URL="postgresql://app:app@localhost:5432/app" +PRISMA_COMPOSE_URL='postgresql://${DB_USERNAME:-app}:${DB_PASSWORD}@database:5432/${DB_DATABASE:-app}' # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/nest.sh" diff --git a/services/redis/drivers/laravel.sh b/services/redis/drivers/laravel.sh index 7d2ab0f..9998dce 100644 --- a/services/redis/drivers/laravel.sh +++ b/services/redis/drivers/laravel.sh @@ -27,3 +27,13 @@ service_driver_apply() { service_driver_dockerfile() { : } + +# REDIS_PASSWORD already reaches the container via compose.yaml's env_file +# (it is in the project's example.env, assembled from this service's own +# env.fragment) — restated here, interpolated rather than baked, so the +# service_driver_compose_env test that guards against a literal password can +# see it the same way the URL-based drivers show theirs. +service_driver_compose_env() { + printf 'REDIS_HOST: cache\n' + printf 'REDIS_PASSWORD: ${REDIS_PASSWORD}\n' +} diff --git a/services/redis/drivers/nest.sh b/services/redis/drivers/nest.sh index 1cc9f61..42459a0 100644 --- a/services/redis/drivers/nest.sh +++ b/services/redis/drivers/nest.sh @@ -15,3 +15,11 @@ service_driver_apply() { service_driver_dockerfile() { : } + +# Same override-then-compose shape as services/shared/nest.sh's DATABASE_URL: +# an operator's own .env wins, otherwise compose builds the URL from the same +# REDIS_PASSWORD the cache container reads, so the password lives in exactly +# one place. +service_driver_compose_env() { + printf 'REDIS_URL: ${REDIS_URL:-redis://:${REDIS_PASSWORD}@cache:6379}\n' +} diff --git a/services/shared/laravel.sh b/services/shared/laravel.sh index 9b9cfdf..9cd2eeb 100644 --- a/services/shared/laravel.sh +++ b/services/shared/laravel.sh @@ -9,6 +9,8 @@ # LARAVEL_PORT the default port for .env.example # LARAVEL_PACKAGE a composer package to require, or "" # LARAVEL_SETUP the Dockerfile block, or "" +# LARAVEL_COMPOSE_ENV the compose.yaml app environment lines this family +# needs beyond DB_CONNECTION, e.g. DB_HOST/DB_PORT service_driver_apply() { # apply_service_drivers runs this in its own `bash -e` process, so a @@ -30,8 +32,47 @@ service_driver_apply() { "DB_USERNAME=app" \ "DB_PASSWORD=app" \ || return 1 + + # APP_KEY has no service to come from — it is per-family, not per-service — + # so no env.fragment can carry it, and the project's example.env (assembled + # from those fragments, before this runs) never sees it any other way. + # Without a value here, compose.yaml's `APP_KEY: ${APP_KEY}` interpolates to + # empty and laravel refuses to boot. + write_env_lines "${SCAFFOLD_PROJECT_ROOT}/example.env" "APP_KEY=changeme" || return 1 + + # laravel has no provider-agnostic read across the SQL connections this file + # serves — `select 1` is the one this task settled on. Written here rather + # than in the route itself so the shipped file carries exactly one probe, + # for the connection this project actually has. + # + # Two separate substitutions, not one: splicing the probe in above the + # shipped `throw` would leave that throw as dead code below a path that + # always returns first. The throw is replaced in place instead, so a + # --db none project keeps it — unreachable in no project this driver ever + # touches. + sed -i.bak 's|// @DB_PROBE@|\\Illuminate\\Support\\Facades\\DB::connection()->select(\x27select 1\x27);|' \ + routes/health.php || return 1 + sed -i.bak "s|throw new \\\\RuntimeException('no database is configured for this project');|return response()->json(['status' => 'ok']);|" \ + routes/health.php || return 1 + rm -f routes/health.php.bak + + grep -q 'DB::connection()->select' routes/health.php \ + && grep -q "return response()->json(\['status' => 'ok'\]);" routes/health.php \ + || die "could not splice the database probe into routes/health.php — has the anchor moved?" } service_driver_dockerfile() { [ -z "$LARAVEL_SETUP" ] || printf '%s\n' "$LARAVEL_SETUP" } + +# DB_CONNECTION first and always: config/database.php defaults to sqlite, so +# its absence is not an error, it is a silent wrong answer. DB_DATABASE, +# DB_USERNAME and DB_PASSWORD reach the container through compose.yaml's +# env_file already (they are in the project's example.env, assembled from +# this service's own env.fragment) — only what laravel does not otherwise +# know (the connection name, the host, the key) needs adding here. +service_driver_compose_env() { + printf 'DB_CONNECTION: %s\n' "$LARAVEL_CONNECTION" + printf '%s\n' "$LARAVEL_COMPOSE_ENV" + printf 'APP_KEY: ${APP_KEY}\n' +} diff --git a/services/shared/nest.sh b/services/shared/nest.sh index 8638316..c010b09 100644 --- a/services/shared/nest.sh +++ b/services/shared/nest.sh @@ -1,11 +1,14 @@ # shellcheck shell=bash -# The Prisma driver. A service's drivers/nest.sh sets the two parameters below +# The Prisma driver. A service's drivers/nest.sh sets the parameters below # and sources this. One client API across every database this toolbox ships is # why Prisma was chosen over TypeORM — the adapter x service matrix collapses # to a single code path. # -# PRISMA_PROVIDER the datasource provider -# PRISMA_URL the DATABASE_URL for .env.example +# PRISMA_PROVIDER the datasource provider +# PRISMA_URL the DATABASE_URL for .env.example (host-side, via +# localhost) +# PRISMA_COMPOSE_URL the same DSN against the compose network, with the +# credentials left as compose interpolations service_driver_apply() { # Before the installs, not after. All three packages place the query engine @@ -55,8 +58,71 @@ datasource db { EOF write_env_lines .env.example "DATABASE_URL=${PRISMA_URL}" || return 1 + + # prisma has no provider-agnostic read: $queryRaw is SQL-only and mongodb + # needs a command. Written here rather than branched in the controller so + # the shipped route carries exactly one probe, for the provider this + # project actually has. + # + # A dynamic import cast to an explicit method signature, not a bare + # `import(...).then(...)`: before `prisma generate` has run (lint runs + # before the :prisma mise task, which build and check both depend on), + # @prisma/client re-exports a generated module that does not exist yet, so + # an untyped access to it is `any` — @typescript-eslint's no-unsafe-* rules + # catch that under --max-warnings 0. The cast keeps the probe typed + # regardless of whether the client has been generated. + # + # Two separate substitutions, not one: splicing the probe in above the + # shipped `throw` would leave that throw as dead code below a path that + # always returns first, and //apps/api:lint runs eslint with + # --max-warnings 0, where no-unreachable is in the recommended set. The + # throw is replaced in place instead, so a --db none project keeps it — + # unreachable in no project this driver ever touches. + # Only the one method this provider calls, not both: `prisma generate` + # (which lint runs before but check runs after) produces a real + # PrismaClient whose mongodb build has no $queryRawUnsafe and whose SQL + # builds have no $runCommandRaw, and asserting a type carrying a method + # the generated class lacks fails tsc's "sufficient overlap" check on the + # cast — caught by generating this project with a real database and + # running its check task, not by lint alone. + local method preamble probe + case "$PRISMA_PROVIDER" in + mongodb) + method='$runCommandRaw(command: object): Promise' + probe='await client.$runCommandRaw({ ping: 1 });' + ;; + *) + method='$queryRawUnsafe(query: string): Promise' + probe="await client.\$queryRawUnsafe('SELECT 1');" + ;; + esac + preamble="const { PrismaClient } = (await import('@prisma/client')) as {\n PrismaClient: new () => { ${method} };\n };\n const client = new PrismaClient();" + sed -i.bak "s|// @DB_PROBE@|${preamble}\n ${probe}|" \ + src/health/health.controller.ts || return 1 + sed -i.bak "s|throw new Error('no database is configured for this project');|return { status: 'ok' };|" \ + src/health/health.controller.ts || return 1 + rm -f src/health/health.controller.ts.bak + + grep -q "PrismaClient" src/health/health.controller.ts \ + && grep -q "return { status: 'ok' };" src/health/health.controller.ts \ + || die "could not splice the database probe into src/health/health.controller.ts — has the anchor moved?" + + # The spliced text's own line breaks are a guess, and the mongodb and SQL + # branches wrap differently once prettier's print width applies to each — + # reformatting here, once, beats hand-matching prettier's output for every + # branch this driver can produce. + pnpm exec prettier --write src/health/health.controller.ts || return 1 } service_driver_dockerfile() { printf 'RUN pnpm exec prisma generate\n' } + +# The value an operator sets in .env wins; otherwise compose composes it from +# the same DB_* variables the database container reads, so the password lives +# in exactly one place and the two cannot drift. Measured against a real +# `docker compose config`: both paths resolve, and the default is not +# evaluated when DATABASE_URL is set. +service_driver_compose_env() { + printf 'DATABASE_URL: ${DATABASE_URL:-%s}\n' "$PRISMA_COMPOSE_URL" +} diff --git a/tests/service.bats b/tests/service.bats index bbdf2d1..d0ff008 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -613,3 +613,26 @@ EOF [[ "$output" == 'postgresql://app@database:5432/app' ]] \ || [[ "$output" == '${DATABASE_URL:-postgresql://app@database:5432/app}' ]] } + +@test "the laravel drivers name the connection selector laravel actually reads" { + # config/database.php is `env('DB_CONNECTION', 'sqlite')`. Without that + # variable laravel does not fail — it silently reads DB_DATABASE as a + # sqlite filename and never contacts the service at all. + for service in mysql postgres mongodb; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + . "${SCAFFOLD_ROOT}/services/${service}/drivers/laravel.sh" + service_driver_compose_env )" + grep -q '^DB_CONNECTION:' <<<"$block" \ + || { echo "${service}/laravel.sh emits no DB_CONNECTION"; false; } + done +} + +@test "the nest drivers name DATABASE_URL and let an operator override it" { + for service in mysql postgres mongodb; do + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + . "${SCAFFOLD_ROOT}/services/${service}/drivers/nest.sh" + service_driver_compose_env )" + grep -q '^DATABASE_URL: \${DATABASE_URL:-' <<<"$block" \ + || { echo "${service}/nest.sh does not allow an override"; false; } + done +} From c3a249fa51692e9dfc1c6e6abd9a3b998f6a33b7 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 15:14:37 +0700 Subject: [PATCH 11/30] feat: let the released stack migrate its own schema MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit install.sh now generates a laravel-valid APP_KEY (base64: + 32 bytes, handled in generate_service_passwords' own loop, with the sed delimiter changed to | since a base64 value can contain /) and runs a profiled migrate compose service after the stack is up, before the success message. Every service driver now implements service_driver_compose_migrate (added to REQUIRED_DRIVER_FUNCTIONS), and a new apply_service_compose_service merges a driver's complete services: fragment — apply_service_compose_env cannot do this itself, since it hardcodes services.app.environment. The migrate service shares the app's image (read back off compose.yaml) and its environment. install.sh's guard for whether a migrate service exists uses `docker compose --profile migrate config --services`, not the profile-less form, which never lists a profiled service at all. The nest driver now installs prisma as a regular dependency, not -D, since the Dockerfile's `pnpm prune --prod` would otherwise strip the CLI the migrate service needs out of the published image. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/nestjs/mise.toml | 14 ++++++ common/install.sh | 39 +++++++++++++--- lib/contract.sh | 8 ++-- lib/service.sh | 70 ++++++++++++++++++++++++++++- services/mongodb/drivers/laravel.sh | 6 +++ services/redis/drivers/laravel.sh | 6 +++ services/redis/drivers/nest.sh | 6 +++ services/shared/laravel.sh | 7 +++ services/shared/nest.sh | 23 +++++++++- tests/compose.bats | 14 ++++++ 10 files changed, 183 insertions(+), 10 deletions(-) diff --git a/adapters/nestjs/mise.toml b/adapters/nestjs/mise.toml index 135e2e6..e6ef8aa 100644 --- a/adapters/nestjs/mise.toml +++ b/adapters/nestjs/mise.toml @@ -22,6 +22,20 @@ run = "pnpm exec eslint \"src/**/*.ts\" \"test/**/*.ts\" --max-warnings 0" # prisma, and this does nothing; a real prisma failure still fails. run = "if [ -f prisma/schema.prisma ]; then pnpm exec prisma generate; fi" +[tasks.migrate] +# prisma's mongodb provider rejects `migrate deploy` outright — measured: +# `The "mongodb" provider is not supported with this command.` — and takes +# `db push` instead. Branching on the schema rather than on a recorded value +# keeps this true for a project whose provider changes. +run = """ +if ! [ -f prisma/schema.prisma ]; then exit 0; fi +if grep -q 'provider *= *"mongodb"' prisma/schema.prisma; then + pnpm exec prisma db push --skip-generate +else + pnpm exec prisma migrate deploy +fi +""" + [tasks.check] run = [{ task = ":prisma" }, "pnpm exec tsc --noEmit"] diff --git a/common/install.sh b/common/install.sh index fc0d968..7680573 100755 --- a/common/install.sh +++ b/common/install.sh @@ -93,8 +93,19 @@ download_release_assets() { generate_service_passwords() { local file="$1" name password while IFS= read -r name; do - password="$(head -c 32 /dev/urandom | base64 | tr -dc 'A-Za-z0-9' | head -c 24)" - sed -i.bak "s/^${name}=changeme\$/${name}=${password}/" "$file" + # APP_KEY is not a password: laravel decrypts with it and rejects anything + # that is not base64: plus exactly 32 bytes. Handled inside this loop + # rather than beside it so example.env keeps one placeholder, and the + # existing-.env guard that greps for a remaining `=changeme` still covers + # it. + if [ "$name" = APP_KEY ]; then + password="base64:$(head -c 32 /dev/urandom | base64)" + else + password="$(head -c 32 /dev/urandom | base64 | tr -dc 'A-Za-z0-9' | head -c 24)" + fi + # `|`, not `/`: a base64 value can itself contain `/`, which would end + # sed's s/// early and leave the line unmatched instead of substituted. + sed -i.bak "s|^${name}=changeme\$|${name}=${password}|" "$file" rm -f "${file}.bak" grep -qF "${name}=${password}" "$file" || { echo "could not set ${name} in ${file}; refusing to start with an unconfirmed password" @@ -114,10 +125,23 @@ check_image_configured() { } start_stack() { - local port docker compose up --remove-orphans -d || return 1 - port="$(grep '^APP_PORT=' .env | cut -d= -f2)" - echo "the application is running on http://localhost:${port:-8080}" +} + +# ADR-0014 seam 5 forbids migrations from an *entrypoint* — a container that +# migrates every time it starts cannot be scaled or rolled back. This is a +# human running one command on the target host, which is what that ADR calls +# the one deploy mechanism that exists today. A project with no database +# ships no migrate service, and `--profile` on a service that is not there +# is not an error. +run_migrations() { + # `docker compose config --services` (no --profile) never lists a service + # gated behind a profile, so that guard alone always skipped the migration + # silently — measured: plain `config --services` prints only `app`, and + # `--profile migrate config --services` prints `migrate app`. + docker compose --profile migrate config --services | grep -qx migrate || return 0 + echo "running migrations..." + docker compose --profile migrate run --rm migrate } main() { @@ -128,6 +152,11 @@ main() { download_release_assets || { echo 'could not download the release assets'; return 1; } check_image_configured || return 1 start_stack || { echo 'could not start the stack; check the output above'; return 1; } + run_migrations || { echo 'could not run migrations; check the output above'; return 1; } + + local port + port="$(grep '^APP_PORT=' .env | cut -d= -f2)" + echo "the application is running on http://localhost:${port:-8080}" } # sourced by the toolbox's tests to exercise one function at a time; running diff --git a/lib/contract.sh b/lib/contract.sh index f5eccf2..5a792f9 100644 --- a/lib/contract.sh +++ b/lib/contract.sh @@ -31,9 +31,11 @@ REQUIRED_SERVICE_FILES=( REQUIRED_SERVICE_VARS=(SERVICE_NAME SERVICE_KIND SERVICE_IMAGE) -# apply_service_drivers calls all three, so a driver shipping fewer fails at -# generation rather than at lint. -REQUIRED_DRIVER_FUNCTIONS=(service_driver_apply service_driver_dockerfile service_driver_compose_env) +# apply_service_drivers calls all four, so a driver shipping fewer fails at +# generation rather than at lint. service_driver_compose_migrate is the +# fourth: every driver implements it, including a cache's, which has no +# schema and prints nothing. +REQUIRED_DRIVER_FUNCTIONS=(service_driver_apply service_driver_dockerfile service_driver_compose_env service_driver_compose_migrate) # The web tier is the presentation layer and opens no connection, so it takes # no driver — stated once, about the role, rather than as a "not applicable" diff --git a/lib/service.sh b/lib/service.sh index c4a7b2b..2b37e99 100644 --- a/lib/service.sh +++ b/lib/service.sh @@ -185,6 +185,66 @@ apply_service_compose_env() { rm -f "$fragment" } +# apply_service_compose_service +# Merges a complete `services:` fragment into compose.yaml. The counterpart to +# apply_service_compose_env above for a driver that needs to add an entire +# sibling service (the migrate runner below), not another line under the +# app's own environment: apply_service_compose_env cannot be reused for this, +# it hardcodes the services.app.environment path, and overloading it with a +# second, unrelated merge target does not belong in the same function. +apply_service_compose_service() { + local project="$1" block="$2" + local file="${project}/compose.yaml" fragment + + [ -n "$block" ] || return 0 + [ -f "$file" ] || die "no compose.yaml in ${project}" + + fragment="$(mktemp)" + printf '%s\n' "$block" > "$fragment" + + if ! yq eval-all --inplace 'select(fileIndex==0) * select(fileIndex==1)' \ + "$file" "$fragment"; then + rm -f "$fragment" + die "could not merge the service into ${file}" + fi + rm -f "$fragment" +} + +# apply_service_compose_migrate +# Writes compose.yaml's migrate service, behind a profile so it never starts +# with the stack (install.sh runs it explicitly, once, after the stack is +# up). The image is read back off compose.yaml rather than hardcoded, so it +# stays correct however the app's own image line is written; the environment +# is the same block apply_service_compose_env just merged into app, since a +# migration needs the same DB_CONNECTION/DATABASE_URL the application does, +# not a second copy of that decision. An empty command (a project with no +# database, or a cache-only driver) merges nothing — no migrate service is +# not an error. +apply_service_compose_migrate() { + local project="$1" env_block="$2" command="$3" + local file="${project}/compose.yaml" image block + + [ -n "$command" ] || return 0 + [ -f "$file" ] || die "no compose.yaml in ${project}" + + image="$(yq '.services.app.image' "$file")" \ + || die "could not read the app image out of ${file}" + + block="$( + printf 'services:\n migrate:\n' + printf ' image: %s\n' "$image" + printf ' env_file:\n - path: .env\n required: false\n' + printf ' profiles:\n - migrate\n' + printf ' %s\n' "$command" + if [ -n "$env_block" ]; then + printf ' environment:\n' + printf '%s\n' "$env_block" | sed 's/^/ /' + fi + )" + + apply_service_compose_service "$project" "$block" +} + # write_env_lines ... # Sets each KEY=value, replacing the key if it is already there. A driver runs # against an .env.example the adapter shipped, so appending blindly would @@ -238,7 +298,7 @@ write_env_lines() { # pnpm-workspace.yaml edit) cannot recover it from its own cwd. apply_service_drivers() { local app="$1" project="$2" family="$3"; shift 3 - local service driver block="" env_block="" rendered + local service driver block="" env_block="" migrate_block="" rendered # web is the presentation tier and takes no driver — the caller decides # that from ADAPTER_ROLE, so reaching here with a family that has none is a @@ -303,10 +363,18 @@ apply_service_drivers() { # shellcheck source=/dev/null # family varies, so the path isn't constant rendered="$( . "$driver"; service_driver_compose_env )" [ -n "$rendered" ] && env_block+="${rendered}"$'\n' + + # Only a database driver prints a command here — a cache's returns + # nothing (see services/redis/drivers/*.sh) — so this stays empty for a + # cache-only project and carries the one migration command otherwise. + # shellcheck source=/dev/null # family varies, so the path isn't constant + rendered="$( . "$driver"; service_driver_compose_migrate )" + [ -n "$rendered" ] && migrate_block+="${rendered}"$'\n' done apply_service_setup "$app" "${block%$'\n'}" apply_service_compose_env "$project" "${env_block%$'\n'}" + apply_service_compose_migrate "$project" "${env_block%$'\n'}" "${migrate_block%$'\n'}" } # record_services diff --git a/services/mongodb/drivers/laravel.sh b/services/mongodb/drivers/laravel.sh index f184ccd..8580ed3 100644 --- a/services/mongodb/drivers/laravel.sh +++ b/services/mongodb/drivers/laravel.sh @@ -91,6 +91,12 @@ service_driver_compose_env() { printf 'APP_KEY: ${APP_KEY}\n' } +# laravel-mongodb provides its own Schema grammar, so the same artisan command +# the SQL connections use also migrates a mongodb-backed project. +service_driver_compose_migrate() { + printf 'command: ["php", "artisan", "migrate", "--force"]\n' +} + # register_mongodb_connection # laravel-mongodb needs a 'mongodb' entry in the connections array; the # Laravel skeleton ships none. Same insert-then-verify shape as diff --git a/services/redis/drivers/laravel.sh b/services/redis/drivers/laravel.sh index 9998dce..fbf9b65 100644 --- a/services/redis/drivers/laravel.sh +++ b/services/redis/drivers/laravel.sh @@ -37,3 +37,9 @@ service_driver_compose_env() { printf 'REDIS_HOST: cache\n' printf 'REDIS_PASSWORD: ${REDIS_PASSWORD}\n' } + +# a cache has no schema to migrate — printing nothing keeps the migrate +# service absent from a project that selected only a cache. +service_driver_compose_migrate() { + : +} diff --git a/services/redis/drivers/nest.sh b/services/redis/drivers/nest.sh index 42459a0..15914a3 100644 --- a/services/redis/drivers/nest.sh +++ b/services/redis/drivers/nest.sh @@ -23,3 +23,9 @@ service_driver_dockerfile() { service_driver_compose_env() { printf 'REDIS_URL: ${REDIS_URL:-redis://:${REDIS_PASSWORD}@cache:6379}\n' } + +# a cache has no schema to migrate — printing nothing keeps the migrate +# service absent from a project that selected only a cache. +service_driver_compose_migrate() { + : +} diff --git a/services/shared/laravel.sh b/services/shared/laravel.sh index 9cd2eeb..ab2c452 100644 --- a/services/shared/laravel.sh +++ b/services/shared/laravel.sh @@ -76,3 +76,10 @@ service_driver_compose_env() { printf '%s\n' "$LARAVEL_COMPOSE_ENV" printf 'APP_KEY: ${APP_KEY}\n' } + +# laravel's migration system is agnostic to which connection it runs +# against — mysql, postgres and mongodb (via laravel-mongodb's own Schema +# grammar) all migrate through the same artisan command. +service_driver_compose_migrate() { + printf 'command: ["php", "artisan", "migrate", "--force"]\n' +} diff --git a/services/shared/nest.sh b/services/shared/nest.sh index c010b09..e75681e 100644 --- a/services/shared/nest.sh +++ b/services/shared/nest.sh @@ -41,7 +41,12 @@ service_driver_apply() { # stays anyway: it names the failure at the point it happens instead of # leaving that to the caller's generic message. pnpm add @prisma/client@6 || return 1 - pnpm add -D prisma@6 || return 1 + # A regular dependency, not -D: `pnpm prune --prod` in the Dockerfile drops + # devDependencies, and the published image is what runs `migrate deploy` on + # deploy. The alternative — a second image, or a compose service mounting + # the source — introduces a build artifact the release does not publish, for + # a command run once. The engines cost image size; see decision record 0021. + pnpm add prisma@6 || return 1 mkdir -p prisma || return 1 # datasource and generator only. models describe the client's domain, which @@ -126,3 +131,19 @@ service_driver_dockerfile() { service_driver_compose_env() { printf 'DATABASE_URL: ${DATABASE_URL:-%s}\n' "$PRISMA_COMPOSE_URL" } + +# prisma's mongodb provider rejects `migrate deploy` outright — measured: +# `The "mongodb" provider is not supported with this command.` — and takes +# `db push` instead. Chosen here, at generation time, from PRISMA_PROVIDER +# rather than recorded separately, so it stays correct if the provider ever +# changes. +service_driver_compose_migrate() { + case "$PRISMA_PROVIDER" in + mongodb) + printf 'command: ["pnpm", "exec", "prisma", "db", "push", "--skip-generate"]\n' + ;; + *) + printf 'command: ["pnpm", "exec", "prisma", "migrate", "deploy"]\n' + ;; + esac +} diff --git a/tests/compose.bats b/tests/compose.bats index 0cb8957..a394629 100644 --- a/tests/compose.bats +++ b/tests/compose.bats @@ -182,3 +182,17 @@ INNER_EOF done [ -z "$wrong" ] || { echo "$wrong"; false; } } + +@test "install.sh generates an APP_KEY laravel will accept" { + # generate_service_passwords' generic 24-character value is rejected with + # "Unsupported cipher or incorrect key length" — laravel needs base64: and + # exactly 32 bytes. + local env_file="${BATS_TEST_TMPDIR}/.env" + printf 'DB_PASSWORD=changeme\nAPP_KEY=changeme\n' > "$env_file" + . "${SCAFFOLD_ROOT}/common/install.sh" + run generate_service_passwords "$env_file" + assert_ok + run grep '^APP_KEY=' "$env_file" + [[ "$output" =~ ^APP_KEY=base64:[A-Za-z0-9+/]{43}=$ ]] \ + || { echo "not a laravel key: ${output}"; false; } +} From 02a1a177db3d728b001bedefda88dfd09e30728b Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 15:53:21 +0700 Subject: [PATCH 12/30] fix: give the nest migrate service a command the runtime image can run MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The compose migrate command written by service_driver_compose_migrate used `pnpm exec`, but the runtime image (adapters/nestjs/Dockerfile[.workspace]) never installs pnpm — only the build stage runs corepack enable, and the runtime stage copies node_modules/dist alone. Measured against a real built image: `which pnpm` exits 1. prisma's own bin does survive `pnpm prune --prod` (it is a regular dependency for exactly this), but at a location that depends on which Dockerfile shape wins, a decision made after this driver runs: apps//node_modules/.bin for the typescript-workspace shape, node_modules/.bin at the container root for the standalone one. The new command tries both, and cd's into whichever matched before running it, since prisma resolves its schema relative to its own cwd and that schema lives nested under the app directory in the workspace shape. That surfaced a second defect once the binary was found: `schema.prisma` was never copied into the runtime image at all, in either Dockerfile shape, since it is not an artifact `nest build` produces. Both Dockerfiles now copy prisma/ into the runtime stage via a bracket glob (`pris[m]a`), which copies nothing rather than failing the build on a --db none project that has no prisma/ directory to copy. Also fixed: `$` in the compose-yaml command needs to be `$$` to survive compose's own variable interpolation before the container ever sees it — a single `$` here silently resolved to an unset variable and blanked the whole loop, caught by inspecting `docker compose config`'s resolved output before trusting it. Verified on a real nestjs+postgres stack: readiness 200 before and after migrating (the SQL probe is connectivity-only, matching Task 5's own documented finding — left alone per this round's ruling), and `docker compose --profile migrate run --rm migrate` exits 0 and reaches the real database. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/nestjs/Dockerfile | 7 ++++++ adapters/nestjs/Dockerfile.workspace | 7 ++++++ services/shared/nest.sh | 36 +++++++++++++++++++++++----- 3 files changed, 44 insertions(+), 6 deletions(-) diff --git a/adapters/nestjs/Dockerfile b/adapters/nestjs/Dockerfile index aaf41bf..dedbc70 100644 --- a/adapters/nestjs/Dockerfile +++ b/adapters/nestjs/Dockerfile @@ -17,6 +17,13 @@ WORKDIR /app ENV NODE_ENV=production COPY --from=build /app/node_modules ./node_modules COPY --from=build /app/dist ./dist +# `pris[m]a`, not `prisma`: a bracket glob that matches nothing copies +# nothing instead of failing the build, which a bare `prisma` would do on a +# --db none project that never ran the driver that creates this directory. +# Needed for the migrate service (see services/shared/nest.sh) to find its +# schema — schema.prisma is not itself an artifact `nest build` produces, so +# nothing else in this image carries it forward from the build context. +COPY --from=build /app/pris[m]a ./prisma USER node # 8080 because that is the container port common/compose.yaml publishes, and # nothing rewrites it — an adapter listening anywhere else publishes a dead diff --git a/adapters/nestjs/Dockerfile.workspace b/adapters/nestjs/Dockerfile.workspace index a0e2457..734101b 100644 --- a/adapters/nestjs/Dockerfile.workspace +++ b/adapters/nestjs/Dockerfile.workspace @@ -29,6 +29,13 @@ ENV NODE_ENV=production COPY --from=build /app/node_modules ./node_modules COPY --from=build /app/apps/@APP_FILTER@/node_modules ./apps/@APP_FILTER@/node_modules COPY --from=build /app/apps/@APP_FILTER@/dist ./apps/@APP_FILTER@/dist +# `pris[m]a`, not `prisma`: a bracket glob that matches nothing copies +# nothing instead of failing the build, which a bare `prisma` would do on a +# --db none project that never ran the driver that creates this directory. +# Needed for the migrate service (see services/shared/nest.sh) to find its +# schema — schema.prisma is not itself an artifact `nest build` produces, so +# nothing else in this image carries it forward from the build context. +COPY --from=build /app/apps/@APP_FILTER@/pris[m]a ./apps/@APP_FILTER@/prisma USER node # 8080 because that is the container port common/compose.yaml publishes, and # nothing rewrites it — an adapter listening anywhere else publishes a dead diff --git a/services/shared/nest.sh b/services/shared/nest.sh index e75681e..63fca74 100644 --- a/services/shared/nest.sh +++ b/services/shared/nest.sh @@ -137,13 +137,37 @@ service_driver_compose_env() { # `db push` instead. Chosen here, at generation time, from PRISMA_PROVIDER # rather than recorded separately, so it stays correct if the provider ever # changes. +# +# Not `pnpm exec`: measured against the built runtime image that `pnpm` +# itself is not there — only the build stage runs `corepack enable`, and the +# runtime stage copies node_modules/dist alone (`which pnpm` exits 1 in the +# built image). prisma's own bin does survive `pnpm prune --prod` (it is a +# regular dependency precisely so it would), but at one of two locations +# depending on which Dockerfile shape wins, a decision made after this +# driver runs: apps//node_modules/.bin for the typescript-workspace +# shape (measured: apps/api/node_modules/.bin/prisma on a generated +# nestjs+postgres project), node_modules/.bin at the container root for the +# standalone shape. The command tries both rather than guessing which one a +# given project will end up with, and `cd`s into whichever one matched +# before running it: WORKDIR stays the container root either way, and +# prisma resolves its schema from its own working directory +# (`./prisma/schema.prisma`), which is nested under the app directory in +# the workspace shape — measured with `Could not find Prisma Schema` before +# this `cd` was added. adapters/nestjs/Dockerfile[.workspace] now copies +# that `prisma/` directory into the runtime image alongside node_modules and +# dist — nothing else in either Dockerfile carried it forward, since +# schema.prisma is not an artifact `nest build` produces. service_driver_compose_migrate() { + local args case "$PRISMA_PROVIDER" in - mongodb) - printf 'command: ["pnpm", "exec", "prisma", "db", "push", "--skip-generate"]\n' - ;; - *) - printf 'command: ["pnpm", "exec", "prisma", "migrate", "deploy"]\n' - ;; + mongodb) args='db push --skip-generate' ;; + *) args='migrate deploy' ;; esac + # `$${d}`/`$$d`, not `${d}`/`$d`: compose interpolates `$var` in + # compose.yaml itself before the command ever reaches the container — + # measured with `docker compose config`, a single `$` here resolved to an + # unset variable and blanked the loop out entirely. `$$` is compose's own + # escape for a literal `$`. No quotes needed around `${d}...`/`$d`: every + # candidate is a fixed literal path, never one with a space to protect. + printf 'command: ["sh", "-c", "for d in apps/*/ ./; do [ -x $${d}node_modules/.bin/prisma ] && cd $$d && exec node_modules/.bin/prisma %s; done; echo prisma binary not found >&2; exit 1"]\n' "$args" } From e2e5ecdb6812c715d01b52f5bdfa2c244e94cbb5 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 16:38:36 +0700 Subject: [PATCH 13/30] feat: start the stack in ci and require it to answer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Everything before this validated compose.yaml as YAML; nothing ever started a container, which is how a stack that could never serve a request shipped for weeks. scripts/deploy-check.sh generates an adapter, builds its released image, brings the compose stack up, waits for the app container to report healthy, asserts the migrate profile service exits 0 (the readiness probe alone proves connectivity, not schema), then asserts both the liveness and declared readiness path return 200 — tearing the stack down on every one of those paths via a trap, not a trailing cleanup line. Both paths are read out of the adapter's own adapter.env, never hardcoded, and readiness is skipped only where the adapter declares none, with that skip stated in the output. adapters.yml gains deploy/deploy-tier-b jobs mirroring smoke/smoke-tier-b's discover-driven matrices, running this script per tier-a and tier-b adapter. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- .github/workflows/adapters.yml | 67 +++++++++++++++ scripts/deploy-check.sh | 143 +++++++++++++++++++++++++++++++++ 2 files changed, 210 insertions(+) create mode 100755 scripts/deploy-check.sh diff --git a/.github/workflows/adapters.yml b/.github/workflows/adapters.yml index 21a7c39..06098f5 100644 --- a/.github/workflows/adapters.yml +++ b/.github/workflows/adapters.yml @@ -107,6 +107,73 @@ jobs: ADAPTER: ${{ matrix.adapter }} run: bats "tests/new-${ADAPTER}.bats" + deploy: + needs: discover + if: ${{ needs.discover.outputs.tier-a != '[]' }} + runs-on: ubuntu-latest + permissions: + contents: read + # a generation plus an image build plus a container start, per adapter. + timeout-minutes: 30 + strategy: + fail-fast: false + matrix: + adapter: ${{ fromJson(needs.discover.outputs.tier-a) }} + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: jdx/mise-action@c2a87611a18de5b3828c5652fe268e992400cb5c # v4.3.0 + - run: corepack enable + - id: language + env: + ADAPTER: ${{ matrix.adapter }} + run: | + lang="$(grep '^ADAPTER_LANGUAGE=' "adapters/${ADAPTER}/adapter.env" | cut -d'"' -f2)" + echo "value=${lang}" >> "$GITHUB_OUTPUT" + - if: steps.language.outputs.value == 'php' + uses: shivammathur/setup-php@f3e473d116dcccaddc5834248c87452386958240 # 2.37.2 + with: + php-version: "8.3" + - env: + ADAPTER: ${{ matrix.adapter }} + run: ./scripts/deploy-check.sh "$ADAPTER" + + deploy-tier-b: + needs: discover + if: ${{ needs.discover.outputs.tier-b != '[]' }} + runs-on: ubuntu-latest + permissions: + contents: read + # same per-adapter cost as deploy (generation + image build + container + # start) — tier b's timeout is longer than tier a's elsewhere in this + # file because those jobs run more tests per adapter, not because this + # one does. + timeout-minutes: 30 + strategy: + fail-fast: false + matrix: + adapter: ${{ fromJson(needs.discover.outputs.tier-b) }} + steps: + - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + - uses: jdx/mise-action@c2a87611a18de5b3828c5652fe268e992400cb5c # v4.3.0 + - run: corepack enable + - id: language + env: + ADAPTER: ${{ matrix.adapter }} + run: | + lang="$(grep '^ADAPTER_LANGUAGE=' "adapters/${ADAPTER}/adapter.env" | cut -d'"' -f2)" + echo "value=${lang}" >> "$GITHUB_OUTPUT" + - if: steps.language.outputs.value == 'php' + uses: shivammathur/setup-php@f3e473d116dcccaddc5834248c87452386958240 # 2.37.2 + with: + php-version: "8.3" + - env: + ADAPTER: ${{ matrix.adapter }} + run: ./scripts/deploy-check.sh "$ADAPTER" + compose: # expensive relative to the other checks here, so gated on the weekly # schedule (not the nightly tier-a-only one) or a manual run diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh new file mode 100755 index 0000000..286a237 --- /dev/null +++ b/scripts/deploy-check.sh @@ -0,0 +1,143 @@ +#!/usr/bin/env bash +# deploy-check.sh [--db ] +# Proves a generated project's released stack serves HTTP and reaches its +# database. Everything before this validated YAML; nothing started a +# container — see docs/superpowers/plans/2026-09-06-deployable-stack.md. +set -euo pipefail + +# Long enough for a cold `docker pull` of the database image plus the app's +# own startup, short enough that a stack that will never come up fails the +# job instead of eating its whole timeout budget. +HEALTH_TIMEOUT_SECONDS=120 +HEALTH_POLL_INTERVAL_SECONDS=2 + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +# shellcheck source=lib/log.sh +source "${ROOT}/lib/log.sh" +# shellcheck source=lib/adapter.sh +source "${ROOT}/lib/adapter.sh" + +[ $# -ge 1 ] || die "usage: deploy-check.sh [--db ]" + +ADAPTER="$1"; shift +DB_SERVICE="" +while [ $# -gt 0 ]; do + case "$1" in + --db) + [ $# -ge 2 ] || die "--db requires a service name" + DB_SERVICE="$2" + shift 2 + ;; + *) die "unknown option: ${1}" ;; + esac +done + +# load_adapter is the same reader `scaffold new` itself uses — reading the +# two paths any other way risks a second copy that drifts from the adapter's +# own, which is exactly how nestjs's Dockerfile came to probe a /health +# nothing served. +SCAFFOLD_ROOT="$ROOT" +export SCAFFOLD_ROOT +load_adapter "$ADAPTER" || die "unknown adapter: ${ADAPTER}" + +ROLE="$ADAPTER_ROLE" +LIVENESS_PATH="$ADAPTER_LIVENESS_PATH" +READINESS_PATH="${ADAPTER_READINESS_PATH:-}" + +TMP_DIR="$(mktemp -d)" +PROJECT_DIR="${TMP_DIR}/demo" +IMAGE_TAG="deploy-check/${ADAPTER}:local" + +# A trap, not a trailing cleanup line: every die() below is a plain `exit 1`, +# and only a trap runs on that path too. +cleanup() { + if [ -f "${PROJECT_DIR}/compose.yaml" ]; then + ( cd "$PROJECT_DIR" && docker compose down -v --remove-orphans ) || true + fi + rm -rf "$TMP_DIR" +} +trap cleanup EXIT + +log "generating ${ADAPTER} into ${PROJECT_DIR}..." +new_args=("$PROJECT_DIR" "--${ROLE}" "$ADAPTER") +[ -n "$DB_SERVICE" ] && new_args+=(--db "$DB_SERVICE") +"${ROOT}/scaffold" new "${new_args[@]}" || die "scaffold new failed for ${ADAPTER}" + +BUILD_YML="${PROJECT_DIR}/.github/workflows/build.yml" +[ -f "$BUILD_YML" ] || die "generated project has no .github/workflows/build.yml" +CONTEXT="$(yq '.jobs.build.with.context' "$BUILD_YML")" +DOCKERFILE="$(yq '.jobs.build.with.dockerfile' "$BUILD_YML")" +[ -n "$CONTEXT" ] && [ "$CONTEXT" != "null" ] || die "could not read build context from ${BUILD_YML}" +[ -n "$DOCKERFILE" ] && [ "$DOCKERFILE" != "null" ] || die "could not read dockerfile path from ${BUILD_YML}" + +log "building ${IMAGE_TAG} from ${DOCKERFILE} (context: ${CONTEXT})..." +docker build -f "${PROJECT_DIR}/${DOCKERFILE}" -t "$IMAGE_TAG" "${PROJECT_DIR}/${CONTEXT}" \ + || die "docker build failed for ${ADAPTER} (${DOCKERFILE})" + +# compose.yaml's app and migrate services both carry the ghcr.io/CHANGEME +# placeholder scaffold ships before a project has a real registry path (see +# common/compose.yaml) — every reference to it becomes the image just built, +# so the stack that comes up next is the one that just passed this check, +# not whatever a registry happens to publish. +export IMAGE_TAG +yq --inplace \ + '(.services[] | select(.image | test("CHANGEME")) | .image) = strenv(IMAGE_TAG)' \ + "${PROJECT_DIR}/compose.yaml" || die "could not rewrite compose.yaml's image" +grep -Eq '^\s*image:.*CHANGEME' "${PROJECT_DIR}/compose.yaml" \ + && die "compose.yaml still names the CHANGEME placeholder after rewriting it" + +cp "${PROJECT_DIR}/example.env" "${PROJECT_DIR}/.env" + +cd "$PROJECT_DIR" + +log "starting the stack..." +docker compose up -d || die "docker compose up failed for ${ADAPTER}" + +log "waiting for the app container to become healthy (up to ${HEALTH_TIMEOUT_SECONDS}s)..." +health="" +elapsed=0 +while [ "$elapsed" -lt "$HEALTH_TIMEOUT_SECONDS" ]; do + health="$(docker compose ps app --format json 2>/dev/null | jq -r '.Health // empty' || true)" + [ "$health" = "healthy" ] && break + [ "$health" = "unhealthy" ] \ + && die "app container reported unhealthy — its HEALTHCHECK against ${LIVENESS_PATH} is failing (see: docker compose logs app)" + sleep "$HEALTH_POLL_INTERVAL_SECONDS" + elapsed=$((elapsed + HEALTH_POLL_INTERVAL_SECONDS)) +done +[ "$health" = "healthy" ] \ + || die "app container did not become healthy within ${HEALTH_TIMEOUT_SECONDS}s (last status: ${health:-unknown})" + +# The readiness probe is `select 1` — it proves connectivity, not schema, and +# returns 200 against an empty database. Asserting the migration's own exit +# code, separately, is what stops a deploy whose migration silently failed +# from going green anyway. +if docker compose --profile migrate config --services 2>/dev/null | grep -qx migrate; then + log "running migrations..." + docker compose --profile migrate run --rm migrate \ + || die "migrate service exited non-zero — schema was not applied" +else + log "no migrate service for ${ADAPTER} — skipping migration" +fi + +PORT="$(grep '^APP_PORT=' .env | cut -d= -f2)" +PORT="${PORT:-8080}" +BASE_URL="http://localhost:${PORT}" + +check_path() { + local label="$1" path="$2" code + code="$(curl -sS -o /dev/null -w '%{http_code}' "${BASE_URL}${path}")" \ + || die "${label} check failed: could not reach ${BASE_URL}${path}" + [ "$code" = "200" ] \ + || die "${label} check failed: ${BASE_URL}${path} returned ${code}, not 200" + log "${label} (${path}): ${code}" +} + +check_path liveness "$LIVENESS_PATH" + +if [ -n "$READINESS_PATH" ]; then + check_path readiness "$READINESS_PATH" +else + log "${ADAPTER} declares no readiness path — skipping readiness check" +fi + +log "${ADAPTER} stack serves HTTP and reaches its database" From eaedb9abc777e436191f13767f008969ebb3d89f Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 16:54:57 +0700 Subject: [PATCH 14/30] fix: bind nextjs to every interface, and probe it on one that exists MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Task 7's deploy gate found this on its first real run: Docker sets HOSTNAME to the container's own id for every container, and standalone server.js binds to `process.env.HOSTNAME || '0.0.0.0'` — so without an explicit override, the app listened on that id-derived address, reachable through compose's published port (routed to the container's real interface either way) but not from the HEALTHCHECK, which dials localhost from inside the same container and got "connection refused" forever. ENV HOSTNAME="0.0.0.0" fixes the bind, but exposed a second, previously masked mismatch: that bind is IPv4-only, while this image's resolver hands wget the ::1 (IPv6) address for "localhost" first, and busybox wget does not fall back to the IPv4 result. Measured with `docker run` at each step. Pointing the HEALTHCHECK at 127.0.0.1 instead closes that gap. PORT was already correct — `ss -tlnp` showed the server bound on :8080 throughout; the defect was only ever the bind address, not the port. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- adapters/nextjs/Dockerfile | 17 ++++++++++++++++- adapters/nextjs/Dockerfile.workspace | 17 ++++++++++++++++- 2 files changed, 32 insertions(+), 2 deletions(-) diff --git a/adapters/nextjs/Dockerfile b/adapters/nextjs/Dockerfile index ff25659..a250e36 100644 --- a/adapters/nextjs/Dockerfile +++ b/adapters/nextjs/Dockerfile @@ -29,7 +29,22 @@ USER node # nothing rewrites it — an adapter listening anywhere else publishes a dead # port. next's standalone server.js reads PORT itself. ENV PORT=8080 +# Docker sets HOSTNAME to the container's own id for every container, and +# the standalone server.js binds to `process.env.HOSTNAME || '0.0.0.0'` — so +# without this, it listens on that id-derived address, not 0.0.0.0. Traffic +# from outside (compose's published port) still reaches it, since that's +# routed to the container's real interface regardless; the HEALTHCHECK below +# runs inside the container and dials localhost, which nothing is listening +# on, so it fails forever while the app answers everyone else. Measured with +# `docker run`: `ss -tlnp` showed the server bound to the bridge IP, and +# HEALTHCHECK logged "connection refused" on every attempt, until this line. +ENV HOSTNAME="0.0.0.0" EXPOSE 8080 +# 127.0.0.1, not localhost: "0.0.0.0" above is an IPv4-only bind, but this +# image's resolver hands wget the ::1 (IPv6) address for "localhost" first, +# and busybox wget does not fall back to the IPv4 result — measured, this +# still failed with "connection refused" after the HOSTNAME fix alone, even +# though the server was listening and answering every other caller. HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:8080/api/health/live || exit 1 + CMD wget -qO- http://127.0.0.1:8080/api/health/live || exit 1 CMD ["node", "server.js"] diff --git a/adapters/nextjs/Dockerfile.workspace b/adapters/nextjs/Dockerfile.workspace index f530041..418a37c 100644 --- a/adapters/nextjs/Dockerfile.workspace +++ b/adapters/nextjs/Dockerfile.workspace @@ -35,7 +35,22 @@ USER node # nothing rewrites it — an adapter listening anywhere else publishes a dead # port. next's standalone server.js reads PORT itself. ENV PORT=8080 +# Docker sets HOSTNAME to the container's own id for every container, and +# the standalone server.js binds to `process.env.HOSTNAME || '0.0.0.0'` — so +# without this, it listens on that id-derived address, not 0.0.0.0. Traffic +# from outside (compose's published port) still reaches it, since that's +# routed to the container's real interface regardless; the HEALTHCHECK below +# runs inside the container and dials localhost, which nothing is listening +# on, so it fails forever while the app answers everyone else. Measured with +# `docker run`: `ss -tlnp` showed the server bound to the bridge IP, and +# HEALTHCHECK logged "connection refused" on every attempt, until this line. +ENV HOSTNAME="0.0.0.0" EXPOSE 8080 +# 127.0.0.1, not localhost: "0.0.0.0" above is an IPv4-only bind, but this +# image's resolver hands wget the ::1 (IPv6) address for "localhost" first, +# and busybox wget does not fall back to the IPv4 result — measured, this +# still failed with "connection refused" after the HOSTNAME fix alone, even +# though the server was listening and answering every other caller. HEALTHCHECK --interval=30s --timeout=3s \ - CMD wget -qO- http://localhost:8080/api/health/live || exit 1 + CMD wget -qO- http://127.0.0.1:8080/api/health/live || exit 1 CMD ["node", "apps/@APP_FILTER@/server.js"] From 0b59f6a7e004faf2d38e3734ccc58177eb1a2391 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 17:48:11 +0700 Subject: [PATCH 15/30] fix: a missing migrate service is a failure, not a green run MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review round 2. Critical: the migrate check gated on a condition derived from the artifact under test — a renamed service, a broken profile, or a driver returning no command all read as "no database", the readiness probe (select 1) then returns 200 against unapplied schema, and the run exits 0. ROLE and DB_SERVICE are already known, so absence is now a failure whenever a database was actually requested. common/install.sh's run_migrations had the identical hole (a database service with no migrate beside it now fails loudly instead of returning 0), since it ships to every generated project. Also, in order of how they'd bite: - COMPOSE_PROJECT_NAME is set per adapter — every generated stack is `name: app` (common/compose.yaml), so a local run used to reconcile against, and `down -v` any real "app" project already on the machine. - ADAPTER_LIVENESS_PATH and ADAPTER_READINESS_PATH are asserted non-empty: lib/lint.sh only checks the line is present, and an empty liveness path silently probes "/", which nextjs happens to answer 200 for reasons unrelated to its real liveness. - compose.yaml's image is asserted equal to the tag just built, not just "no CHANGEME left" — a real registry reference in common/compose.yaml would match neither check and the stack would come up on a pulled image. - notify-on-schedule-failure now watches deploy and deploy-tier-b too. - APP_PORT's read no longer dies silently under pipefail when a .env lacks it; the health poll uses `docker inspect` instead of `docker compose ps --format json`, whose shape is compose-version-dependent; the cleanup trap now covers INT/TERM, not just EXIT. tests/compose.bats' liveness-probe test is updated to accept 127.0.0.1 alongside localhost, matching fix round 1's nextjs HEALTHCHECK change. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- .github/workflows/adapters.yml | 3 +- common/install.sh | 18 +++++++-- scripts/deploy-check.sh | 67 ++++++++++++++++++++++++++++++---- tests/compose.bats | 8 +++- 4 files changed, 84 insertions(+), 12 deletions(-) diff --git a/.github/workflows/adapters.yml b/.github/workflows/adapters.yml index 06098f5..e615893 100644 --- a/.github/workflows/adapters.yml +++ b/.github/workflows/adapters.yml @@ -242,8 +242,9 @@ jobs: if: >- ${{ always() && github.event_name == 'schedule' && (needs.smoke.result == 'failure' || needs.smoke-tier-b.result == 'failure' || + needs.deploy.result == 'failure' || needs.deploy-tier-b.result == 'failure' || needs.compose.result == 'failure' || needs.services.result == 'failure') }} - needs: [smoke, smoke-tier-b, compose, services] + needs: [smoke, smoke-tier-b, deploy, deploy-tier-b, compose, services] runs-on: ubuntu-latest permissions: issues: write diff --git a/common/install.sh b/common/install.sh index 7680573..a36de5b 100755 --- a/common/install.sh +++ b/common/install.sh @@ -139,9 +139,21 @@ run_migrations() { # gated behind a profile, so that guard alone always skipped the migration # silently — measured: plain `config --services` prints only `app`, and # `--profile migrate config --services` prints `migrate app`. - docker compose --profile migrate config --services | grep -qx migrate || return 0 - echo "running migrations..." - docker compose --profile migrate run --rm migrate + if docker compose --profile migrate config --services | grep -qx migrate; then + echo "running migrations..." + docker compose --profile migrate run --rm migrate + return + fi + # A database service with no migrate service beside it is not "nothing to + # migrate" — every database driver ships a migrate command, so this + # combination only happens if the service, its profile, or the command + # itself silently vanished. Returning 0 here is exactly the hole that let + # a stack go green with unapplied schema; a project with no database at + # all is the only case this falls through to. + if docker compose config --services | grep -qx database; then + echo "a database service exists but no migrate service was found — refusing to start with unapplied schema" >&2 + return 1 + fi } main() { diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh index 286a237..9227527 100755 --- a/scripts/deploy-check.sh +++ b/scripts/deploy-check.sh @@ -41,22 +41,44 @@ export SCAFFOLD_ROOT load_adapter "$ADAPTER" || die "unknown adapter: ${ADAPTER}" ROLE="$ADAPTER_ROLE" + +# lib/lint.sh only checks the line is present, not that it names a route — +# an empty path would otherwise probe "/", which nextjs happens to answer +# 200 for reasons that have nothing to do with the adapter's real liveness. +[ -n "$ADAPTER_LIVENESS_PATH" ] || die "${ADAPTER} declares an empty ADAPTER_LIVENESS_PATH" LIVENESS_PATH="$ADAPTER_LIVENESS_PATH" -READINESS_PATH="${ADAPTER_READINESS_PATH:-}" + +# `${ADAPTER_READINESS_PATH:-}` alone can't tell "not declared" (skip, and +# say so) from "declared empty" (a malformed adapter.env — lint only greps +# for the line's presence, not a non-empty value): both collapse to "". The +# `+x` test keeps them apart. +if [ -n "${ADAPTER_READINESS_PATH+x}" ]; then + [ -n "$ADAPTER_READINESS_PATH" ] || die "${ADAPTER} declares an empty ADAPTER_READINESS_PATH" + READINESS_PATH="$ADAPTER_READINESS_PATH" +else + READINESS_PATH="" +fi TMP_DIR="$(mktemp -d)" PROJECT_DIR="${TMP_DIR}/demo" IMAGE_TAG="deploy-check/${ADAPTER}:local" +# Every generated project's compose.yaml is `name: app` (common/compose.yaml) +# — without this, a local run reconciles against, and `down -v`s, any real +# "app" project already running on this machine, database volumes included. +COMPOSE_PROJECT_NAME="deploy-check-${ADAPTER}" +export COMPOSE_PROJECT_NAME + # A trap, not a trailing cleanup line: every die() below is a plain `exit 1`, -# and only a trap runs on that path too. +# and only a trap runs on that path too. INT/TERM too, so a cancelled CI job +# or a Ctrl-C doesn't leave containers and a temp dir behind. cleanup() { if [ -f "${PROJECT_DIR}/compose.yaml" ]; then ( cd "$PROJECT_DIR" && docker compose down -v --remove-orphans ) || true fi rm -rf "$TMP_DIR" } -trap cleanup EXIT +trap cleanup EXIT INT TERM log "generating ${ADAPTER} into ${PROJECT_DIR}..." new_args=("$PROJECT_DIR" "--${ROLE}" "$ADAPTER") @@ -83,8 +105,21 @@ export IMAGE_TAG yq --inplace \ '(.services[] | select(.image | test("CHANGEME")) | .image) = strenv(IMAGE_TAG)' \ "${PROJECT_DIR}/compose.yaml" || die "could not rewrite compose.yaml's image" -grep -Eq '^\s*image:.*CHANGEME' "${PROJECT_DIR}/compose.yaml" \ - && die "compose.yaml still names the CHANGEME placeholder after rewriting it" + +# Asserting equality with the tag just built, not just "no CHANGEME left": if +# common/compose.yaml ever ships a real registry reference instead of the +# placeholder, the select("CHANGEME") above matches nothing, no CHANGEME +# string remains either, and the stack would come up on a *pulled* image +# while the one just built is discarded — a green run proving nothing. +assert_image_is_built_tag() { + local service="$1" actual + actual="$(yq ".services.${service}.image" "${PROJECT_DIR}/compose.yaml")" + [ "$actual" = "$IMAGE_TAG" ] \ + || die "compose.yaml's ${service} image is ${actual}, not the image just built (${IMAGE_TAG})" +} +assert_image_is_built_tag app +yq -e '.services.migrate' "${PROJECT_DIR}/compose.yaml" >/dev/null 2>&1 \ + && assert_image_is_built_tag migrate cp "${PROJECT_DIR}/example.env" "${PROJECT_DIR}/.env" @@ -97,7 +132,13 @@ log "waiting for the app container to become healthy (up to ${HEALTH_TIMEOUT_SEC health="" elapsed=0 while [ "$elapsed" -lt "$HEALTH_TIMEOUT_SECONDS" ]; do - health="$(docker compose ps app --format json 2>/dev/null | jq -r '.Health // empty' || true)" + # `docker inspect` on the container itself, not `docker compose ps + # --format json`: that format's shape is compose-version-dependent — a + # version emitting an array instead of one object per line makes `jq -r + # '.Health'` error, which the `|| true` this needs anyway would swallow + # into a false "unknown", producing a full 120s red on an actually-healthy + # stack. `docker inspect` on one container id has one shape. + health="$(docker inspect --format '{{.State.Health.Status}}' "$(docker compose ps -q app)" 2>/dev/null || true)" [ "$health" = "healthy" ] && break [ "$health" = "unhealthy" ] \ && die "app container reported unhealthy — its HEALTHCHECK against ${LIVENESS_PATH} is failing (see: docker compose logs app)" @@ -111,15 +152,27 @@ done # returns 200 against an empty database. Asserting the migration's own exit # code, separately, is what stops a deploy whose migration silently failed # from going green anyway. +# +# A missing migrate service used to just log a skip and exit 0 — which means +# renaming the service, breaking `config` under the migrate profile, or a +# driver returning an empty command all look identical to "this adapter has +# no database" from here, and the check that exists to catch exactly that +# regression turns itself off. ROLE and DB_SERVICE are already known, so +# absence is only ever a skip when no database was actually requested. if docker compose --profile migrate config --services 2>/dev/null | grep -qx migrate; then log "running migrations..." docker compose --profile migrate run --rm migrate \ || die "migrate service exited non-zero — schema was not applied" +elif [ "$ROLE" != "web" ] && [ "$DB_SERVICE" != "none" ]; then + die "expected a migrate service for ${ADAPTER} (role=${ROLE}, db=${DB_SERVICE:-default}) but compose has none — a service, profile, or driver may have silently vanished" else log "no migrate service for ${ADAPTER} — skipping migration" fi -PORT="$(grep '^APP_PORT=' .env | cut -d= -f2)" +# `|| true`: under pipefail, a .env with no APP_PORT line makes grep exit 1 +# and, unguarded, that kills the script here — silently, before the +# `${PORT:-8080}` fallback below ever gets a chance to run. +PORT="$(grep '^APP_PORT=' .env | cut -d= -f2 || true)" PORT="${PORT:-8080}" BASE_URL="http://localhost:${PORT}" diff --git a/tests/compose.bats b/tests/compose.bats index a394629..35f0367 100644 --- a/tests/compose.bats +++ b/tests/compose.bats @@ -176,7 +176,13 @@ INNER_EOF [ -f "$file" ] || continue grep -q "HEALTHCHECK" "$file" \ || { wrong="${wrong}${file}: no HEALTHCHECK"$'\n'; continue; } - grep -q "localhost:8080${path}" "$file" \ + # localhost or 127.0.0.1: nextjs's HEALTHCHECK dials 127.0.0.1 because + # this image's resolver hands "localhost" the IPv6 ::1 first and the + # IPv4-only listener (forced by ENV HOSTNAME="0.0.0.0", the fix for + # standalone server.js otherwise binding to the container's own id) + # refuses it — see adapters/nextjs/Dockerfile. Either host still + # proves the adapter's declared path is the one actually probed. + grep -Eq "(localhost|127\.0\.0\.1):8080${path}" "$file" \ || wrong="${wrong}${file}: does not probe ${path} on 8080"$'\n' done done From 957aa4b98697d7b54f94d357e209a4a5f95b3c14 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Sun, 6 Sep 2026 18:15:31 +0700 Subject: [PATCH 16/30] docs: record what a running stack changed about the seams Task 8 of the deployable-stack plan: ADR-0014 seams 1 and 4, ADR-0003's no-application-code boundary, the tour's healthcheck example, the runbook's pull-only step 10, and the design spec's overclaimed readiness proof were all falsified by the work in commits c61600f..0b59f6a. Records the corrected boundary in a new ADR-0021, marks the superseded parts of ADR-0014 and ADR-0003 rather than editing them silently, and fixes the tour/runbook/spec to describe the stack that actually runs today. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv --- ...ter-overlay-instead-of-vendored-presets.md | 13 ++ .../0014-deployment-deferred-with-seams.md | 15 ++ .../0021-the-released-stack-must-run.md | 214 ++++++++++++++++++ docs/runbook/first-project-walkthrough.md | 52 ++++- .../2026-09-06-deployable-stack-design.md | 20 +- docs/tour/07-containers.md | 62 +++-- 6 files changed, 344 insertions(+), 32 deletions(-) create mode 100644 docs/decisions/0021-the-released-stack-must-run.md diff --git a/docs/decisions/0003-adapter-overlay-instead-of-vendored-presets.md b/docs/decisions/0003-adapter-overlay-instead-of-vendored-presets.md index e8d9247..c4e17fe 100644 --- a/docs/decisions/0003-adapter-overlay-instead-of-vendored-presets.md +++ b/docs/decisions/0003-adapter-overlay-instead-of-vendored-presets.md @@ -12,6 +12,19 @@ application per stack. Invoke each framework's own generator and overlay four files. +**Amendment, 2026-09-06 (0021).** The boundary was "an adapter overlays +configuration, it never writes application code." It narrows to: no +application code except the health routes 0021's deploy gate requires — one +liveness route per adapter, and a readiness route for every adapter that +can hold a database. The exception stays narrow on purpose: one route per +concern, owned by the same adapter that already declares the Dockerfile +and the path the route answers on, and nothing else in a generated `apps/` +tree is adapter-authored. A route that only proves a listener answers, or +only proves a database connection opens, is not the same claim as "this +adapter generates the application" — it is the minimum the gate needs to +tell a dead deploy from a live one, which invoking the framework's own +generator cannot provide on its own. + ## Consequences About 40–80 lines owned per stack instead of an entire application. diff --git a/docs/decisions/0014-deployment-deferred-with-seams.md b/docs/decisions/0014-deployment-deferred-with-seams.md index f363c0c..347beef 100644 --- a/docs/decisions/0014-deployment-deferred-with-seams.md +++ b/docs/decisions/0014-deployment-deferred-with-seams.md @@ -2,6 +2,7 @@ Status: Accepted Date: 2026-08-27 +Superseded in part by 0021 (seams 1 and 4) ## Context @@ -52,6 +53,20 @@ deploy target plugs into later without restructuring anything above it: reusable release workflow will carry a `deploy` job that does nothing until a client sets that variable. +**Update, 2026-09-06 (0021).** Seams 1 and 4 as written above are now +false, and the record stays rather than being rewritten: seam 4's premise +was that php-fpm speaks FastCGI and no HTTP check is possible, so no check +ships; the Laravel adapters now serve HTTP through FrankenPHP, so a real +`HEALTHCHECK` exists and ships. The reasoning that check-that-cannot-fail is +worse than no check was correct then and stays correct — only the premise +under it changed. Seam 1's claim ("a client's target only ever needs to +know how to run one image") was already in tension with seam 4 admitting +one of those images could not be usefully run at all without protocol-aware +infrastructure a target would have to supply on its own; that tension now +resolves in seam 1's favour, since every adapter speaks HTTP on the same +port and a target genuinely only ever has to run one image. See 0021 for +the working implementation. + `common/deploy-adapters/` ships empty, with a `README.md` explaining why and pointing at this ADR. `install.sh` is the one deploy mechanism that exists today: a human runs it, by hand, on the target host, after cloning nothing diff --git a/docs/decisions/0021-the-released-stack-must-run.md b/docs/decisions/0021-the-released-stack-must-run.md new file mode 100644 index 0000000..2c8f589 --- /dev/null +++ b/docs/decisions/0021-the-released-stack-must-run.md @@ -0,0 +1,214 @@ +# 0021 — The released stack must run + +Status: Accepted +Date: 2026-09-06 + +## Context + +An acceptance run on 2026-09-05 took four freshly generated projects through +the whole walkthrough on real private repositories, then, for the first +time, started the stack those projects publish. It did not work, in any +shape: no `DATABASE_URL` reached the app, the Laravel images ended at +`php-fpm` with no web server in front of it, `compose.yaml` published port +8080 while every adapter listened somewhere else, and `nestjs`'s +`HEALTHCHECK` probed `/health`, a route no adapter has ever shipped. ADR-0014 +built the seams a deploy target plugs into later; nothing had proven the +image sitting at those seams actually runs. + +## Decision + +**Two gates, both mechanically checkable, both now enforced by +`.github/workflows/adapters.yml`'s `deploy`/`deploy-tier-b` jobs alongside +the existing `smoke` lane:** + +1. **Green immediately.** Generate, clone, run every config root's + `ci-unit` in the clone. Already held as of 2026-09-05; unchanged here. +2. **Deployable immediately.** Generate, build the image using the + `context`/`dockerfile` pair the generated `build.yml` names, run the + released stack's own start-up sequence against it, then: the **liveness + path** returns 200 (the container serves HTTP on the port compose + publishes), and the **readiness path** returns 200 where the adapter + declares one. `scripts/deploy-check.sh` is gate 2's implementation; + `common/install.sh` is the same sequence handed to a client. + +**The container port is fixed at 8080, not a variable.** Every adapter +serves HTTP on container port 8080; `common/compose.yaml` publishes +`${APP_PORT:-8080}:8080` and nothing rewrites it per-adapter. The +alternative — teaching compose each adapter's port through an +`APP_CONTAINER_PORT` written at generation time — is rejected for the same +reason ADR-0014 rejected a parameterised `IMAGE_REPOSITORY`: + +> the repository path does not vary release to release the way the tag +> does — it is set once and never touched again, so a variable buys +> nothing a literal placeholder with a comment does not already give, at +> the cost of one more name to keep straight in `.env`. + +A container port is exactly that kind of value: fixed for the life of a +generated project the moment its adapter is chosen, never touched again, +and a variable would only add a place for the default to drift from the +one true value. The cost is real and stated on the line that sets it: an +engineer running the image by hand gets 8080, not the port their +framework's own docs name. + +**FrankenPHP, alpine, pinned by digest, serves the Laravel images.** +`php-fpm` speaks FastCGI; this stack has no reverse proxy in front of it, so +nothing served HTTP at all. FrankenPHP is the only option that serves PHP +*and* the Vite assets in `public/build` *and* keeps the stack at one +container — a nginx sidecar would need a genuinely new kind of service in +`lib/service.sh`, one `app` both depends on and shares a volume with, which +this decision declines to build for a problem FrankenPHP already solves in +two Dockerfiles. It is not a guess: FrankenPHP has its own section in +`laravel.com/docs/13.x/deployment`, has been part of the PHP Foundation +since May 2025, runs Laravel Cloud, and Shopware has run it in production +for over a year. Alpine is not optional here — +`services/mongodb/drivers/laravel.sh` emits `apk add`, and the mongodb +driver would silently break on a base image that has no `apk`. + +**Configuration stays environment-only, composed from one password.** +`compose.yaml`'s `app` service gains an `environment:` block, written by +the selected service's driver (`service_driver_compose_env`, beside the +existing `service_driver_dockerfile`), because the shape is per adapter +family — Prisma wants one `DATABASE_URL`; Laravel wants `DB_CONNECTION` +plus a DSN and falls back to sqlite without it — and a compose fragment +belongs to a service, not a family. Each value defaults from the same +`DB_*`/`APP_KEY` variables `.env` already carries +(`DATABASE_URL: ${DATABASE_URL:-postgresql://...${DB_PASSWORD}@database:...}`), +and an operator's own `DATABASE_URL` in `.env` wins without the default ever +being evaluated — measured against a real `docker compose config`. The +password exists in exactly one place. The rejected alternative — a driver +writing a literal `DATABASE_URL=…` into `example.env` — puts the same +password in two places and asks `install.sh`'s independently-randomised +passwords to stay equal by coincidence: the `changeme`-versus-`app` defect +the 2026-09-05 run measured, reintroduced in a new costume. + +**Each adapter declares a liveness path and, where it can hold a database, +a readiness path** (`ADAPTER_LIVENESS_PATH` / `ADAPTER_READINESS_PATH` in +`adapter.env`), read by both the `HEALTHCHECK` and `scripts/deploy-check.sh` +so the two can never disagree about the route the way `nestjs`'s +`HEALTHCHECK` and its generator once did. Readiness runs one query and +reports it — `select 1` for a SQL connection, a ping command for mongodb — +returning 200 when it succeeds and 503 when it does not. **State plainly +what that proves and what it does not**: a request that reaches the +readiness route and gets a 200 has proven the listener, the environment +contract, the compose network and the credentials — four of the five links +in the chain a deploy needs. It has not proven the schema. The probe +succeeds against an empty database exactly as readily as a migrated one, +and for `nestjs` with `mongodb`, `db push` records no migration state at +all for anything to read back. The schema is proven separately: the gate +and `install.sh` both require the `migrate` service to exit 0, and a +database-bearing adapter with no `migrate` service in `compose.yaml` is +treated as a failure, not a shape with nothing to migrate. `nextjs` ships +no readiness route at all — the `web` role takes no database driver, and a +route that returns 200 without querying anything is the same +check-that-cannot-fail this record spends the next section naming. + +**The Nest runtime image carries the Prisma CLI, at a size cost, so it can +migrate itself.** `services/shared/nest.sh` installs `prisma` as a regular +dependency, not a dev dependency, specifically so `pnpm prune --prod` +leaves its binary and query engines in the runtime image the released +stack actually ships. The alternative — a second image, or a compose +service that mounts source, built only to run a migration once per deploy — +introduces a build artifact the release does not otherwise publish, for a +command that already has a home: `compose.yaml`'s `migrate` service, run +under a `migrate` compose profile so it never starts with the stack, using +the same image and environment `app` does. `install.sh` runs it once, +visibly, after the stack is up — ADR-0014 seam 5's constraint (no +entrypoint runs a migration on every start) still holds; a human running +one command on the target host is the one deploy mechanism that ADR-0014 +says exists today. + +## What running one revealed + +Two things emerged only from doing this work, not from planning it, and are +worth keeping for the next person who touches this seam. + +**The same defect shape, five times.** Each was a check that could not +fail: `lib/lint.sh`'s readiness-path enforcement had no test that failed +when the enforcement itself was deleted; the Nest post-generate wiring step +exited 0 whether or not its anchor `sed` actually matched, silently able to +ship an unregistered health route on a future generator reformat; nothing +asserted the Nest migrate command's `$$`-escaping, so tidying it to a +single `$` would have broken every Nest migration with no test to catch it; +`lib/lint.sh` still only checks that an `ADAPTER_*_PATH` line exists, not +that it names a real route; and the sharpest instance sat inside the gate +built to prevent exactly this class of defect — the deploy gate's migration +assertion was originally gated on a condition read from the same artifact +under test, so deleting the `migrate` service turned a required check into +a printed skip, and the run went green with an unmigrated schema. **Two of +the five were found only by starting a container — something nothing in +this repository had ever done before this gate.** FrankenPHP's `CMD` +silently dropped the base image's default arguments, leaving nothing +listening on 8080 while every static assertion (`EXPOSE 8080`, `HEALTHCHECK` +present) still passed; only building the image and starting it showed the +port was dead. And `nextjs`'s bundled server binding, below. + +**`nextjs` bound to the wrong address, for two independent reasons.** Its +standalone `server.js` binds to `process.env.HOSTNAME || '0.0.0.0'`, and +Docker sets `HOSTNAME` to the container's own id for every container — so +without `ENV HOSTNAME="0.0.0.0"`, the server listened on an address its own +`HEALTHCHECK` could never dial. Fixing that exposed a second, independent +cause behind the same symptom: `0.0.0.0` is an IPv4-only bind, but this +image's resolver hands `wget` the IPv6 `::1` first for `localhost`, and +busybox `wget` does not fall back to the IPv4 result — so the `HEALTHCHECK` +still failed, for a different reason, after the first fix landed. Neither +was visible from Dockerfile text; both only showed up once something +actually ran the image. + +## Consequences + +- Every adapter's runtime image serves HTTP on 8080 with a `HEALTHCHECK` + that probes the path the adapter itself declares, and `tests/compose.bats` + asserts all three (`EXPOSE 8080`, `HEALTHCHECK` present, the probed path + matches `adapter.env`) statically, cheaply, on every change — while + knowing those static assertions cannot catch the two defects above; only + the deploy gate, which starts a container, can. +- A generated project's `compose.yaml`, `example.env` and (for adapters that + need one) `mise.toml` migrate task all changed to carry the environment + contract and the migration path this record describes; + `docs/tour/07-containers.md` and + `docs/runbook/first-project-walkthrough.md` were amended to stop + describing the stack that predated this work. +- ADR-0014's seam 1 (one image to run) and seam 4 (no healthcheck is + possible for a FastCGI service) are both superseded in part; ADR-0014 + itself records where. +- ADR-0003's boundary — an adapter overlays configuration, never writes + application code — narrows to admit exactly the health routes this + record's gate requires, and nothing else; ADR-0003 records the exception. +- **One image per project still stands.** A `web`+`api` project deploys + only the role that won `set_image_context` (the last one on the command + line); the gate tests that image, and the application beside it is + generated and checked, never deployed. Unchanged by this record, and + worth restating because gate 2 looks like it covers a project when it + covers one image. +- **A readiness route is application code a client may delete.** Nothing + detects that later. The gate tests generated projects, not a client's + repository six months on. + +## Alternatives considered + +- **A nginx sidecar in front of php-fpm**, keeping FastCGI. Rejected: needs + a new service kind in `lib/service.sh` — one `app` both depends on and + shares a volume with — plus a shipped `nginx.conf` and moving the port + publish off `app`. FrankenPHP changes two Dockerfiles and nothing else. +- **Octane**, running FrankenPHP in worker mode. Rejected for now: worker + mode makes client request-handling code stateful by default and hands the + client Laravel's own memory-leak-management burden, for a throughput + problem no client has yet reported. Moving to it later is a one-line + `ENTRYPOINT` change on the same pinned base image. +- **A per-adapter `APP_CONTAINER_PORT` variable**, matching how `IMAGE_TAG` + varies. Rejected for the reason quoted above from ADR-0014: the value + does not vary the way a tag does, so a variable only adds a place for the + default to drift. +- **A driver writing a literal `DATABASE_URL` into `example.env`.** + Rejected: puts the same password in two places instead of one, which is + the exact defect class this record's environment contract exists to + close. +- **A schema marker read back per provider**, so readiness could prove + migration too. Rejected: mongodb's `db push` records no migration state + for anything to read, so a marker would need inventing per provider for a + property the gate already proves a cheaper way — requiring `migrate` to + exit 0. +- **A second image, or a compose service mounting source, to run Nest's + migration.** Rejected: introduces a build artifact the release does not + otherwise publish, for a command run once per deploy; carrying the + Prisma CLI's size cost in the one image already published is smaller. diff --git a/docs/runbook/first-project-walkthrough.md b/docs/runbook/first-project-walkthrough.md index 5231564..33a77c4 100644 --- a/docs/runbook/first-project-walkthrough.md +++ b/docs/runbook/first-project-walkthrough.md @@ -256,13 +256,59 @@ gh release list Expect: `v0.2.0`, and the image tagged `0.2.0`, `0.2`, `latest`, `sha-…`. -## 10. Run what was built +## 10. Run it ```sh -docker pull ghcr.io/ttncode/demo-app:0.2.0 +git checkout main +git pull +./install.sh +``` + +Expect: it fails immediately, printing `could not download the release +assets`. `install.sh`'s `RepoUrl` still names the CHANGEME/CHANGEME +placeholder GitHub org and repo — `scaffold new` could not have filled that +in: no repository existed yet to read a path from. This is the first point +the one-time edit ADR-0014 describes can happen, now that step 9's release +has given both `RepoUrl` and `compose.yaml`'s image line something real to +name. + +```sh +sed -i "s#github.com/CHANGEME/CHANGEME#github.com/ttncode/demo-app#" install.sh +sed -i "s#ghcr.io/CHANGEME/CHANGEME#ghcr.io/ttncode/demo-app#" compose.yaml +git checkout -b fix/point-at-published-image +git commit -am "fix: point install.sh and compose.yaml at the published image" +git push -u origin fix/point-at-published-image +gh pr create --fill +gh pr checks --watch +gh pr merge --squash --delete-branch +gh release list ``` -Expect: pulls, if the package is public or you are logged in to ghcr. +Expect: a `fix:` commit moves the patch version, so this cuts `v0.2.1` — +the release `install.sh` downloads from once it names the right repository. + +```sh +./install.sh +curl -fsS http://localhost:8080/api/health/live +``` + +Expect: `install.sh` downloads `compose.yaml` and `example.env` from +`v0.2.1`, generates passwords, starts the stack, runs the migration task, +and prints `the application is running on http://localhost:8080`. The curl +returns `200`. + +There is no readiness path to curl for this project: `--web nextjs` is the +role that won the image (the last one on the command line, back in step 5), +and `nextjs` ships no readiness route — the `web` role takes no database +driver, so there is nothing for one to query. A project whose deployed +image is `laravel-api` or `nestjs` additionally has +`curl -fsS http://localhost:8080/health/ready` return `200`. See ADR-0021 +for both routes, and for why a project that requests more than one role +still deploys only one image. + +```sh +docker compose -f app/compose.yaml down -v +``` ## 11. Add a second application to the existing project diff --git a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md index b7c9ffa..73b02d8 100644 --- a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md +++ b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md @@ -75,7 +75,13 @@ and each currently failing: the port compose publishes. - its **readiness path** returns 200 — a request reaches the database through the application, proving the listener, the environment contract, - the compose network, the credentials and the schema in one call. + the compose network and the credentials in one call. **Correction** + (Task 6/8): it does not also prove the schema — the probe is `select 1` + (or a ping), which returns 200 against an empty database exactly as + readily as a migrated one, and for `nestjs` with `mongodb`, `db push` + records no migration state for anything to read back. The schema is + proven separately: gate 2 also requires the `migrate` service (section + 8) to exit 0. Both paths are declared by the adapter (section 7), because they differ per framework and because two of them are wrong today. @@ -232,10 +238,14 @@ adapters therefore need a real liveness path, not just a corrected probe. GET /health/ready -> 200 when the query succeeds, 503 when it does not ``` -This is the only thing that proves the whole chain — listener, environment, -compose network, credentials, schema — in a single call. A liveness path -alone cannot: Laravel's `/up` never touches the database, so a project with a -wrong `DATABASE_URL` passes it. That is precisely the check-that-cannot-fail +This is the only thing that proves the listener, the environment contract, +the compose network and the credentials in a single call — not the schema +(**correction**, Task 6/8: the probe succeeds against an empty database as +readily as a migrated one, and mongodb's `db push` leaves nothing to read +back regardless; section 8's `migrate` service, required to exit 0, is what +proves that instead). A liveness path alone cannot prove even that much: +Laravel's `/up` never touches the database, so a project with a wrong +`DATABASE_URL` passes it. That is precisely the check-that-cannot-fail ADR-0014 seam 4 warned about, and shipping one as the only gate would repeat the mistake this design exists to correct. diff --git a/docs/tour/07-containers.md b/docs/tour/07-containers.md index acde00a..24f4a4e 100644 --- a/docs/tour/07-containers.md +++ b/docs/tour/07-containers.md @@ -24,16 +24,22 @@ and `tests/service.bats` fails a fragment that pins its own. ## Read this - `adapters/laravel-api/Dockerfile` — vendor stage (`composer install`) - separate from the runtime stage, and the comment explaining why it ships - with no `HEALTHCHECK` at all: it used to run `php -r 'exit(0);'`, which - only proved the PHP binary starts, not that php-fpm is serving requests, - and never failed a review or CI because it could not fail *at all*. It - was removed outright rather than kept — a check that can never fail is - worse than no check: an orchestrator with none at least knows it doesn't - know a container's state; one with an always-green check believes it - does, and routes real traffic to a dead container on that false - confidence. -- `adapters/nestjs/Dockerfile` — a real HTTP `HEALTHCHECK`, for contrast. + separate from the FrankenPHP runtime stage, and the comment explaining + why FrankenPHP replaced php-fpm: php-fpm speaks FastCGI, this stack has + no reverse proxy in front of it, and the check that used to ship here — + `php -r 'exit(0);'` — only proved the PHP binary starts, never failed a + review or CI because it could not fail *at all*, and was removed outright + rather than kept. That argument still holds: a check that can never fail + is worse than no check — an orchestrator with none at least knows it + doesn't know a container's state; one with an always-green check believes + it does, and routes real traffic to a dead container on that false + confidence. What changed is the premise underneath it, not the argument + (ADR-0014, ADR-0021): FrankenPHP serves real HTTP, so + `HEALTHCHECK … CMD wget -qO- http://localhost:8080/up` is a check that can + actually fail. +- `adapters/nestjs/Dockerfile` — the same shape, for contrast: its + `HEALTHCHECK` probes `/health/live`, the route + `adapters/nestjs/src/health/health.controller.ts` ships. - `lib/service.sh`'s `assemble_compose` — the merge described above, and `service_compose_key` for why a fragment must publish under `database` or `cache`, not its own service name: `depends_on` names the key, not @@ -54,20 +60,28 @@ and `tests/service.bats` fails a fragment that pins its own. ## Delete test -Delete the `HEALTHCHECK` line from `adapters/nestjs/Dockerfile` (or -`nextjs`'s) and nothing here notices: no test in `tests/` asserts a -Dockerfile has one, `mise run checklist` stays green, and -`docker compose up` in `common/compose.yaml` doesn't gate on it either — -only a selected service's own health is wired to `app`'s `depends_on`, and -only for that service. The consequence only shows up once a client's own -deploy target (one of -ADR-0014's seams) actually polls container health before routing traffic: -a slow-starting container gets real requests before it's ready, and -nothing in this repository would have pointed at a missing -`HEALTHCHECK` as the reason. If you're adding one to a new adapter, -delete-test it the other direction first: stop the process the check is -supposed to detect, and confirm the check actually goes unhealthy — the -laravel lesson above is what happens when nobody does. +Delete the `HEALTHCHECK` line from `adapters/nestjs/Dockerfile` (or any +other adapter's) and something notices now: `tests/compose.bats` asserts +every adapter Dockerfile has `EXPOSE 8080`, a `HEALTHCHECK`, and that the +`HEALTHCHECK` probes the exact path the adapter's own `adapter.env` +declares. Point it at a path nothing serves instead of deleting it, and the +same assertion still catches it — that is the `nestjs` defect ADR-0021 +records: it probed `/health`, which no adapter has ever served, for as long +as this repository existed, and nothing here noticed until this test was +written to compare the two values. + +What that test still cannot catch: whether the process behind the probe +ever answers for real. It reads Dockerfile text; it never builds an image +or starts a container. ADR-0021 records two defects invisible to every +static check in this repository, found only once something actually ran +the image — FrankenPHP's `CMD` silently dropping the base image's default +arguments (nothing listened on 8080 while `EXPOSE`/`HEALTHCHECK` both still +read correctly), and `nextjs` binding to an address its own `HEALTHCHECK` +could never dial. Only `scripts/deploy-check.sh`, the deploy gate, starts a +container, which is what closes that gap. If you're adding a `HEALTHCHECK` +to a new adapter, delete-test it the way that gate does: stop the process +the check is supposed to detect, and confirm the check actually goes +unhealthy — the laravel lesson above is what happens when nobody does. ## Try it From 2f4e443ec439c06e4b1a2d1bcef645515811c323 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:19:44 +0700 Subject: [PATCH 17/30] fix: reuse the nest readiness probe's PrismaClient instead of leaking one MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ready() constructed a new PrismaClient on every request and never disconnected it. Readiness is polled by every orchestrator, often every few seconds; measured against a real Postgres container, 35 polls of the old code left 35 idle connections that never closed, while the fixed code holds a single connection regardless of poll count. HealthController is a Nest singleton by default, so the client is now a field on the controller, lazily created once and reused — the same process lifetime a NestJS Prisma integration normally gives a connect-once service, without adding a full provider/module for one probe route. --- .../nestjs/src/health/health.controller.ts | 21 ++++++++++++++----- services/shared/nest.sh | 21 ++++++++++++++----- 2 files changed, 32 insertions(+), 10 deletions(-) diff --git a/adapters/nestjs/src/health/health.controller.ts b/adapters/nestjs/src/health/health.controller.ts index 102623e..f9cfe89 100644 --- a/adapters/nestjs/src/health/health.controller.ts +++ b/adapters/nestjs/src/health/health.controller.ts @@ -2,16 +2,27 @@ import { Controller, Get, HttpException, HttpStatus } from '@nestjs/common'; @Controller('health') export class HealthController { + // The client field and the probe below are both written by the selected + // service's driver: prisma has no provider-agnostic read, so a SQL + // provider gets $queryRawUnsafe and mongodb gets $runCommandRaw. A project + // generated with --db none leaves this anchor as a comment and the probe + // falls through to its throw, reporting 503 honestly. + // @DB_CLIENT@ + @Get('live') live(): { status: string } { return { status: 'ok' }; } - // The probe is written by the selected service's driver: prisma has no - // provider-agnostic read, so a SQL provider gets $queryRawUnsafe and - // mongodb gets $runCommandRaw. A project generated with --db none keeps - // the anchor's fallback and reports 503, because there is nothing here - // that could honestly report ready. + // Readiness is polled by every orchestrator, often every few seconds, so + // the probe below reuses the client field declared above instead of + // constructing a new PrismaClient per request: HealthController is a + // Nest singleton (the default provider/controller scope), so one instance + // — and one connection pool — lives for the process, the same lifetime a + // NestJS Prisma integration normally gives it via a connect-once service. + // A fresh client per call would need its own $disconnect() to avoid + // leaking a connection per poll, but tearing a real pool down and back up + // every few seconds is the wasteful version of the same fix. @Get('ready') async ready(): Promise<{ status: string }> { try { diff --git a/services/shared/nest.sh b/services/shared/nest.sh index 63fca74..57f4fc9 100644 --- a/services/shared/nest.sh +++ b/services/shared/nest.sh @@ -90,25 +90,36 @@ EOF # the generated class lacks fails tsc's "sufficient overlap" check on the # cast — caught by generating this project with a real database and # running its check task, not by lint alone. - local method preamble probe + # The client is a field on HealthController, not a local inside ready(): + # a controller is a Nest singleton by default, so one field lives for the + # whole process and every poll of /health/ready after the first reuses it. + # Constructing a PrismaClient per request and never closing it leaks one + # real database connection per poll — measured exhausting Postgres's + # max_connections well inside an hour at a 10s probe interval. + local method field preamble probe case "$PRISMA_PROVIDER" in mongodb) method='$runCommandRaw(command: object): Promise' - probe='await client.$runCommandRaw({ ping: 1 });' + probe='await this.dbClient.$runCommandRaw({ ping: 1 });' ;; *) method='$queryRawUnsafe(query: string): Promise' - probe="await client.\$queryRawUnsafe('SELECT 1');" + probe="await this.dbClient.\$queryRawUnsafe('SELECT 1');" ;; esac - preamble="const { PrismaClient } = (await import('@prisma/client')) as {\n PrismaClient: new () => { ${method} };\n };\n const client = new PrismaClient();" + field="private dbClient?: { ${method} };" + sed -i.bak "s|// @DB_CLIENT@|${field}|" \ + src/health/health.controller.ts || return 1 + + preamble="if (!this.dbClient) {\n const { PrismaClient } = (await import('@prisma/client')) as {\n PrismaClient: new () => { ${method} };\n };\n this.dbClient = new PrismaClient();\n }" sed -i.bak "s|// @DB_PROBE@|${preamble}\n ${probe}|" \ src/health/health.controller.ts || return 1 sed -i.bak "s|throw new Error('no database is configured for this project');|return { status: 'ok' };|" \ src/health/health.controller.ts || return 1 rm -f src/health/health.controller.ts.bak - grep -q "PrismaClient" src/health/health.controller.ts \ + grep -q "dbClient" src/health/health.controller.ts \ + && grep -q "PrismaClient" src/health/health.controller.ts \ && grep -q "return { status: 'ok' };" src/health/health.controller.ts \ || die "could not splice the database probe into src/health/health.controller.ts — has the anchor moved?" From c358af660715a29b9a9f15e9f5114de86f972054 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:36:45 +0700 Subject: [PATCH 18/30] fix: require adapter health paths to hold a route, not just exist MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit lint_adapters only greped that ADAPTER_LIVENESS_PATH= and ADAPTER_READINESS_PATH= appeared, so an empty value passed. That is wider than it looks: tests/compose.bats extracts the path through a pipeline where cut exits 0 regardless, so an empty value collapses its HEALTHCHECK assertion into matching any localhost probe on 8080 — the exact Dockerfile-probing-nothing defect the assertion exists to catch. Both variables now have to hold a value starting with "/" when declared. Added tests/fixtures/lint/empty-readiness-path and a test proving the new check fails on it. --- lib/lint.sh | 22 +++++++++++++++- tests/contract.bats | 11 ++++++++ .../empty-readiness-path/sample/.env.example | 1 + .../empty-readiness-path/sample/Dockerfile | 1 + .../empty-readiness-path/sample/adapter.env | 6 +++++ .../empty-readiness-path/sample/mise.toml | 26 +++++++++++++++++++ 6 files changed, 66 insertions(+), 1 deletion(-) create mode 100644 tests/fixtures/lint/empty-readiness-path/sample/.env.example create mode 100644 tests/fixtures/lint/empty-readiness-path/sample/Dockerfile create mode 100644 tests/fixtures/lint/empty-readiness-path/sample/adapter.env create mode 100644 tests/fixtures/lint/empty-readiness-path/sample/mise.toml diff --git a/lib/lint.sh b/lib/lint.sh index b792238..c86c3ff 100644 --- a/lib/lint.sh +++ b/lib/lint.sh @@ -4,7 +4,7 @@ # prints one line per problem and returns 1 when any adapter is incomplete. lint_adapters() { local dir="$1" - local adapter name file task task_body flag var role status=0 + local adapter name file task task_body flag var role value status=0 for adapter in "$dir"/*/; do [ -d "$adapter" ] || continue @@ -37,6 +37,26 @@ lint_adapters() { } ;; esac + + # A path variable that merely exists is not a route: an empty value + # satisfies every check above, and downstream that same empty value + # collapses tests/compose.bats' HEALTHCHECK assertion and the deploy + # gate's readiness curl into matching any localhost probe on 8080 — + # exactly the Dockerfile-probing-nothing defect these exist to stop. + # Only checked when the variable is declared at all: an undeclared + # ADAPTER_READINESS_PATH on a non-driven role is handled above, not + # here. + for var in ADAPTER_LIVENESS_PATH ADAPTER_READINESS_PATH; do + grep -Eq "^${var}=" "${adapter}adapter.env" || continue + value="$(sed -n "s/^${var}=\"\(.*\)\"\$/\1/p" "${adapter}adapter.env")" + case "$value" in + /*) ;; + *) + printf '%s: adapter.env sets %s to "%s", not a path starting with /\n' "$name" "$var" "$value" + status=1 + ;; + esac + done fi [ -f "${adapter}mise.toml" ] || continue diff --git a/tests/contract.bats b/tests/contract.bats index f1d9c6a..694d188 100644 --- a/tests/contract.bats +++ b/tests/contract.bats @@ -155,6 +155,17 @@ setup() { [ -z "$output" ] } +@test "lint_adapters rejects a readiness path declared but left empty" { + # Both checks above only grep that the line is present, not that it holds + # a route: an empty value passed both, and downstream that same empty + # value collapses tests/compose.bats' HEALTHCHECK assertion and the deploy + # gate's readiness curl into matching any localhost probe on 8080 — the + # exact defect those checks exist to stop. + run lint_adapters "${SCAFFOLD_ROOT}/tests/fixtures/lint/empty-readiness-path" + [ "$status" -eq 1 ] + [[ "$output" == *'sample: adapter.env sets ADAPTER_READINESS_PATH to "", not a path starting with /'* ]] +} + @test "scaffold lint covers the services that ship" { run scaffold lint assert_ok diff --git a/tests/fixtures/lint/empty-readiness-path/sample/.env.example b/tests/fixtures/lint/empty-readiness-path/sample/.env.example new file mode 100644 index 0000000..ae255d7 --- /dev/null +++ b/tests/fixtures/lint/empty-readiness-path/sample/.env.example @@ -0,0 +1 @@ +APP_ENV=local diff --git a/tests/fixtures/lint/empty-readiness-path/sample/Dockerfile b/tests/fixtures/lint/empty-readiness-path/sample/Dockerfile new file mode 100644 index 0000000..c35f1b5 --- /dev/null +++ b/tests/fixtures/lint/empty-readiness-path/sample/Dockerfile @@ -0,0 +1 @@ +FROM scratch diff --git a/tests/fixtures/lint/empty-readiness-path/sample/adapter.env b/tests/fixtures/lint/empty-readiness-path/sample/adapter.env new file mode 100644 index 0000000..fea4eb2 --- /dev/null +++ b/tests/fixtures/lint/empty-readiness-path/sample/adapter.env @@ -0,0 +1,6 @@ +ADAPTER_NAME="sample" +ADAPTER_ROLE="api" +ADAPTER_FAMILY="laravel" +ADAPTER_GENERATOR='mkdir -p "$APP_DIR"' +ADAPTER_LIVENESS_PATH="/health/live" +ADAPTER_READINESS_PATH="" diff --git a/tests/fixtures/lint/empty-readiness-path/sample/mise.toml b/tests/fixtures/lint/empty-readiness-path/sample/mise.toml new file mode 100644 index 0000000..8ec8128 --- /dev/null +++ b/tests/fixtures/lint/empty-readiness-path/sample/mise.toml @@ -0,0 +1,26 @@ +[tasks.install] +run = "true" + +[tasks.format] +run = "true" + +[tasks."format-fix"] +run = "true" + +[tasks.lint] +run = "true" + +[tasks.check] +run = "true" + +[tasks.test] +run = "true" + +[tasks.build] +run = "true" + +[tasks.ci-unit] +run = "true" + +[tasks.checklist] +run = "true" From 0a4976d670c1bb5394ece58d674a9ce3c8bfab6d Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:37:00 +0700 Subject: [PATCH 19/30] test: prove lint_services' driver-function loop can fail The missing-driver fixture proved a family with no driver file at all fails lint. Nothing proved the mirror case: a driver file that exists, sources cleanly, but omits one of REQUIRED_DRIVER_FUNCTIONS. Deleting the whole `for fn` loop left the suite passing, the same "gate that cannot fail" shape already named for the readiness-path check one loop above. Added tests/fixtures/lint-services/missing-driver-function (a complete laravel driver alongside a nest driver missing service_driver_compose_migrate) and a test asserting the specific "does not define" message. Verified by deleting the loop, confirming only this new test fails, then restoring it. --- tests/contract.bats | 13 +++++++++++++ .../sample/compose.dev.fragment.yaml | 2 ++ .../sample/compose.fragment.yaml | 2 ++ .../sample/compose.prod.fragment.yaml | 2 ++ .../sample/compose.test.fragment.yaml | 2 ++ .../sample/drivers/laravel.sh | 5 +++++ .../missing-driver-function/sample/drivers/nest.sh | 7 +++++++ .../missing-driver-function/sample/env.fragment | 1 + .../missing-driver-function/sample/service.env | 3 +++ 9 files changed, 37 insertions(+) create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/compose.dev.fragment.yaml create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/compose.fragment.yaml create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/compose.prod.fragment.yaml create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/compose.test.fragment.yaml create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/drivers/laravel.sh create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/drivers/nest.sh create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/env.fragment create mode 100644 tests/fixtures/lint-services/missing-driver-function/sample/service.env diff --git a/tests/contract.bats b/tests/contract.bats index 694d188..b878ff3 100644 --- a/tests/contract.bats +++ b/tests/contract.bats @@ -100,6 +100,19 @@ setup() { [[ "$output" == *"sample: no driver for laravel"* ]] } +@test "lint_services reports a driver that does not define a required function" { + # The missing-driver fixture above proves a family with no driver file + # fails; nothing proved the mirror case — a driver file that exists and + # sources cleanly but omits one of REQUIRED_DRIVER_FUNCTIONS. Deleting the + # whole `for fn` loop in lint_services left this suite green, which is the + # same "gate that cannot fail" shape task 1's own ruling already named. + run lint_services \ + "${SCAFFOLD_ROOT}/tests/fixtures/lint-services/missing-driver-function" \ + "${SCAFFOLD_ROOT}/adapters" + [ "$status" -eq 1 ] + [[ "$output" == *"sample: nest driver does not define service_driver_compose_migrate"* ]] +} + @test "lint_services reports a missing required file" { run lint_services \ "${SCAFFOLD_ROOT}/tests/fixtures/lint-services/missing-file" \ diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/compose.dev.fragment.yaml b/tests/fixtures/lint-services/missing-driver-function/sample/compose.dev.fragment.yaml new file mode 100644 index 0000000..034d8e3 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/compose.dev.fragment.yaml @@ -0,0 +1,2 @@ +services: + database: {} diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/compose.fragment.yaml b/tests/fixtures/lint-services/missing-driver-function/sample/compose.fragment.yaml new file mode 100644 index 0000000..034d8e3 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/compose.fragment.yaml @@ -0,0 +1,2 @@ +services: + database: {} diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/compose.prod.fragment.yaml b/tests/fixtures/lint-services/missing-driver-function/sample/compose.prod.fragment.yaml new file mode 100644 index 0000000..034d8e3 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/compose.prod.fragment.yaml @@ -0,0 +1,2 @@ +services: + database: {} diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/compose.test.fragment.yaml b/tests/fixtures/lint-services/missing-driver-function/sample/compose.test.fragment.yaml new file mode 100644 index 0000000..034d8e3 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/compose.test.fragment.yaml @@ -0,0 +1,2 @@ +services: + database: {} diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/drivers/laravel.sh b/tests/fixtures/lint-services/missing-driver-function/sample/drivers/laravel.sh new file mode 100644 index 0000000..7a65504 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/drivers/laravel.sh @@ -0,0 +1,5 @@ +# shellcheck shell=bash +service_driver_apply() { :; } +service_driver_dockerfile() { :; } +service_driver_compose_env() { :; } +service_driver_compose_migrate() { :; } diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/drivers/nest.sh b/tests/fixtures/lint-services/missing-driver-function/sample/drivers/nest.sh new file mode 100644 index 0000000..e2cb2e6 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/drivers/nest.sh @@ -0,0 +1,7 @@ +# shellcheck shell=bash +# Deliberately incomplete: this driver exists and sources cleanly, but omits +# service_driver_compose_migrate — the fixture for the "a driver exists and +# does not define a required function" gap in lint_services' function loop. +service_driver_apply() { :; } +service_driver_dockerfile() { :; } +service_driver_compose_env() { :; } diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/env.fragment b/tests/fixtures/lint-services/missing-driver-function/sample/env.fragment new file mode 100644 index 0000000..75cc45b --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/env.fragment @@ -0,0 +1 @@ +DB_PASSWORD=changeme diff --git a/tests/fixtures/lint-services/missing-driver-function/sample/service.env b/tests/fixtures/lint-services/missing-driver-function/sample/service.env new file mode 100644 index 0000000..5e7d6d4 --- /dev/null +++ b/tests/fixtures/lint-services/missing-driver-function/sample/service.env @@ -0,0 +1,3 @@ +SERVICE_NAME="sample" +SERVICE_KIND="database" +SERVICE_IMAGE="docker.io/library/busybox:1@sha256:0000000000000000000000000000000000000000000000000000000000000000" From 3b48bd3b22bf8b8b4115be007c61cf6c8f390bfb Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:44:01 +0700 Subject: [PATCH 20/30] fix: make the deploy gate generate real passwords, not changeme MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit deploy-check.sh copied example.env to .env and never called generate_service_passwords, so the gate ran with DB_PASSWORD=changeme and APP_KEY=changeme literal — the password loop, the APP_KEY branch, and run_migrations could all break with every check in this branch staying green. The gate now sources common/install.sh and calls generate_service_passwords on the copied .env. Verified directly (a changeme value becomes a real random one) and end to end (nestjs+postgres and laravel-api+mysql both still pass the gate with generated credentials). ADR-0021 claimed common/install.sh "is the same sequence handed to a client", which was false — the gate builds and starts the stack directly and never downloads a release or runs install.sh end to end. Corrected to say what the gate actually shares with install.sh (generate_service_passwords) and what it does not. --- docs/decisions/0021-the-released-stack-must-run.md | 10 ++++++++-- scripts/deploy-check.sh | 9 +++++++++ 2 files changed, 17 insertions(+), 2 deletions(-) diff --git a/docs/decisions/0021-the-released-stack-must-run.md b/docs/decisions/0021-the-released-stack-must-run.md index 2c8f589..f774281 100644 --- a/docs/decisions/0021-the-released-stack-must-run.md +++ b/docs/decisions/0021-the-released-stack-must-run.md @@ -28,8 +28,14 @@ the existing `smoke` lane:** released stack's own start-up sequence against it, then: the **liveness path** returns 200 (the container serves HTTP on the port compose publishes), and the **readiness path** returns 200 where the adapter - declares one. `scripts/deploy-check.sh` is gate 2's implementation; - `common/install.sh` is the same sequence handed to a client. + declares one. `scripts/deploy-check.sh` is gate 2's implementation. It + builds the image locally and starts the stack directly rather than + running `common/install.sh` end to end — it does not download a + release, call `create_directory`, or check `check_image_configured` — + but it does call `install.sh`'s own `generate_service_passwords` on the + copied `.env`, so the password loop and the `APP_KEY` branch run under + the same substitution a client's install would perform, not against + every credential left at `changeme`. **The container port is fixed at 8080, not a variable.** Every adapter serves HTTP on container port 8080; `common/compose.yaml` publishes diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh index 9227527..4d76177 100755 --- a/scripts/deploy-check.sh +++ b/scripts/deploy-check.sh @@ -121,7 +121,16 @@ assert_image_is_built_tag app yq -e '.services.migrate' "${PROJECT_DIR}/compose.yaml" >/dev/null 2>&1 \ && assert_image_is_built_tag migrate +# common/install.sh's own generate_service_passwords, not a second copy of +# the substitution: a gate that leaves every password at the literal +# "changeme" runs a sequence no real deploy ever runs, and proves nothing +# about the password loop, the APP_KEY branch, or anything downstream that +# depends on either. cp "${PROJECT_DIR}/example.env" "${PROJECT_DIR}/.env" +# shellcheck source=/dev/null # path is this toolbox's own common/install.sh +source "${ROOT}/common/install.sh" +generate_service_passwords "${PROJECT_DIR}/.env" \ + || die "could not generate service passwords for ${ADAPTER}" cd "$PROJECT_DIR" From bc559251a907daadefc4ae34a3b026fa6ac57ee8 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:48:31 +0700 Subject: [PATCH 21/30] fix: skip the readiness curl for a driven adapter generated with --db none A driven adapter (api/app) declares a readiness path regardless of --db, so `./scripts/deploy-check.sh laravel-api --db none` curled /health/ready expecting 200 against a route that has no driver spliced in and correctly returns 503. Confirmed both ways: the gate now passes laravel-api --db none, and reverting the guard reproduces the exact "returned 503, not 200" failure. The spec claimed --db none ships no readiness route at all; it does, and it answers honestly. Corrected to describe what actually happens and why the gate now skips the check for that combination instead of curling a route designed to fail. --- .../2026-09-06-deployable-stack-design.md | 21 +++++++++++-------- scripts/deploy-check.sh | 8 +++++++ 2 files changed, 20 insertions(+), 9 deletions(-) diff --git a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md index 73b02d8..b1f856d 100644 --- a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md +++ b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md @@ -256,15 +256,18 @@ boundary moves from "no application code" to "no application code except a readiness route the deploy gate requires", which is narrow, stated, and testable. -A project generated with `--db none` ships no readiness route, and neither -does `nextjs` in any shape: there is nothing for either to query, and a route -that returns 200 without doing anything is the same worthless check in a -different place. The gate curls readiness only when the image it built serves -one — which is decided by the role that won `set_image_context`, not by -whether the project has a database. A `--api laravel-api --web nextjs --db -mysql` project builds the **nextjs** image, so its gate is liveness only, and -the Laravel app beside it is generated and checked but never deployed. Section -13 says why that is a limit worth naming. +`nextjs` ships no readiness route in any shape — there is nothing for it to +query, and a route that returns 200 without doing anything is the same +worthless check in a different place. A driven adapter (api/app) ships one +unconditionally, including with `--db none`: the probe's anchor is never +spliced, so its `throw` survives and the route honestly reports 503. The gate +curls readiness only when the adapter declares a path and the project was not +generated with `--db none` — a driven adapter with no database would +otherwise fail the gate against a route correctly reporting itself unready. A +`--api laravel-api --web nextjs --db mysql` project builds the **nextjs** +image, so its gate is liveness only, and the Laravel app beside it is +generated and checked but never deployed. Section 13 says why that is a limit +worth naming. ## 8. Migrations at install time diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh index 4d76177..d9dc15a 100755 --- a/scripts/deploy-check.sh +++ b/scripts/deploy-check.sh @@ -59,6 +59,12 @@ else READINESS_PATH="" fi +# A driven adapter still declares a readiness path with --db none: the route +# ships unconditionally and correctly reports 503 (nothing to connect to), +# but a gate that curls it expecting 200 would fail a combination the spec +# says is fine. Skipped the same way a non-driven role's absent path is. +[ "$DB_SERVICE" = none ] && READINESS_PATH="" + TMP_DIR="$(mktemp -d)" PROJECT_DIR="${TMP_DIR}/demo" IMAGE_TAG="deploy-check/${ADAPTER}:local" @@ -198,6 +204,8 @@ check_path liveness "$LIVENESS_PATH" if [ -n "$READINESS_PATH" ]; then check_path readiness "$READINESS_PATH" +elif [ "$DB_SERVICE" = none ]; then + log "--db none — skipping readiness check" else log "${ADAPTER} declares no readiness path — skipping readiness check" fi From 0718e51b3252e0e2a6d568c79941bcba79531ccf Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 00:57:59 +0700 Subject: [PATCH 22/30] fix: assert absence of a literal password, not presence of an interpolation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit tests/service.bats' password check had two holes: `[ -z "$block" ] && continue` skipped any driver emitting nothing, and the assertion checked for the presence of `${DB_PASSWORD}`/`${REDIS_PASSWORD}` rather than the absence of a literal, so a driver emitting both would have passed. The mysql and postgres laravel drivers carried a `DB_PASSWORD: ${DB_PASSWORD}` line whose own comment said it existed so this test could see it — the password already reaches the container through compose's env_file. Removed both lines; the corrected test still passes on the drivers' real merits, and a real laravel-api deploy against both mysql and postgres still authenticates and runs its migrations with a generated (non-changeme) password, so nothing depended on the restated line. --- services/mysql/drivers/laravel.sh | 11 ++++------ services/postgres/drivers/laravel.sh | 11 ++++------ tests/service.bats | 33 ++++++++++++++++++++++++---- 3 files changed, 37 insertions(+), 18 deletions(-) diff --git a/services/mysql/drivers/laravel.sh b/services/mysql/drivers/laravel.sh index df617ad..5dd3964 100644 --- a/services/mysql/drivers/laravel.sh +++ b/services/mysql/drivers/laravel.sh @@ -6,13 +6,10 @@ LARAVEL_PORT="3306" # image already carries. LARAVEL_PACKAGE="" LARAVEL_SETUP="RUN docker-php-ext-install pdo_mysql" -# DB_PASSWORD already reaches the container via compose.yaml's env_file (it -# is in the project's example.env, assembled from this service's own -# env.fragment) — restated here, interpolated rather than baked, so the -# service_driver_compose_env test that guards against a literal password can -# see it the same way the URL-based drivers show theirs. +# DB_PASSWORD is not restated here: it already reaches the container via +# compose.yaml's env_file (it is in the project's example.env, assembled +# from this service's own env.fragment). LARAVEL_COMPOSE_ENV="DB_HOST: database -DB_PORT: 3306 -DB_PASSWORD: \${DB_PASSWORD}" +DB_PORT: 3306" # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/laravel.sh" diff --git a/services/postgres/drivers/laravel.sh b/services/postgres/drivers/laravel.sh index 4955e96..0cc1af1 100644 --- a/services/postgres/drivers/laravel.sh +++ b/services/postgres/drivers/laravel.sh @@ -5,13 +5,10 @@ LARAVEL_PORT="5432" LARAVEL_PACKAGE="" LARAVEL_SETUP="RUN apk add --no-cache postgresql-dev \\ && docker-php-ext-install pdo_pgsql" -# DB_PASSWORD already reaches the container via compose.yaml's env_file (it -# is in the project's example.env, assembled from this service's own -# env.fragment) — restated here, interpolated rather than baked, so the -# service_driver_compose_env test that guards against a literal password can -# see it the same way the URL-based drivers show theirs. +# DB_PASSWORD is not restated here: it already reaches the container via +# compose.yaml's env_file (it is in the project's example.env, assembled +# from this service's own env.fragment). LARAVEL_COMPOSE_ENV="DB_HOST: database -DB_PORT: 5432 -DB_PASSWORD: \${DB_PASSWORD}" +DB_PORT: 5432" # shellcheck source=/dev/null . "${SCAFFOLD_ROOT}/services/shared/laravel.sh" diff --git a/tests/service.bats b/tests/service.bats index d0ff008..bfbe735 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -587,18 +587,43 @@ EOF || { echo "allowBuilds (line ${allow_builds}) must come before the first pnpm add (line ${first_add})"; false; } } -@test "a driver's compose environment interpolates rather than embedding a password" { +@test "a driver's compose environment never bakes a literal password" { # The password must exist in exactly one place — .env — so compose composes # the URL at `up` time. A literal baked here is the changeme-versus-app # mismatch that made the dev stack unable to authenticate. + # + # No skip for an empty block: an empty block has no literal password by + # construction, so the checks below already cover it without a special + # case — the previous `[ -z "$block" ] && continue` let a driver that + # regressed to emitting nothing pass unseen, for a reason unrelated to + # passwords. + # + # Asserts the absence of a literal, not the presence of an interpolation: + # a block could carry `${DB_PASSWORD}` somewhere else and a hardcoded + # value where the credential actually goes, and the old presence-only + # check could not tell the two apart. local bad="" for driver in "${SCAFFOLD_ROOT}"/services/*/drivers/*.sh; do block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" SERVICE_DIR="$(dirname "$(dirname "$driver")")" . "$driver"; service_driver_compose_env )" - [ -z "$block" ] && continue - grep -q '\${DB_PASSWORD' <<<"$block" || grep -q '\${REDIS_PASSWORD' <<<"$block" \ - || bad="${bad}${driver}"$'\n' + + # A top-level *_PASSWORD key whose value is not exactly an interpolation. + while IFS= read -r line; do + case "$line" in + *_PASSWORD:\ \$\{*_PASSWORD\}) ;; + *) bad="${bad}${driver} (${line})"$'\n' ;; + esac + done < <(grep -E '^[A-Za-z_]*_PASSWORD:' <<<"$block") + + # A DSN's user:password@ slot whose password is not exactly an + # interpolation — the same shape, embedded in a URL instead of a key. + while IFS= read -r segment; do + case "$segment" in + :\$\{*_PASSWORD\}@) ;; + *) bad="${bad}${driver} (${segment})"$'\n' ;; + esac + done < <(grep -oE ':[^:@]*@' <<<"$block") done [ -z "$bad" ] || { echo "embeds a literal password:"; echo "$bad"; false; } } From 766c111866528e343e48f67b4e538e7a588fb7e0 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 01:11:57 +0700 Subject: [PATCH 23/30] fix: four small corrections from the final review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - tests/compose.bats:177 — anchor the HEALTHCHECK grep to the start of the line, so it cannot match the word inside an unrelated comment. - docs/superpowers/specs/2026-09-06-deployable-stack-design.md:225 — nextjs's liveness path table entry still said `/`; it has been `/api/health/live` since Task 1. - adapters/nextjs/adapter.env:15 — "created by the next task" was plan-speak that shipped into a generated project's own file. Rewritten to name the route's actual path. - common/compose.yaml:17 — "replace it once, by hand" is wrong for a file install.sh re-downloads and overwrites on every run. Points at the runbook's actual flow (fix the repository's compose.yaml and cut a release) instead. --- adapters/nextjs/adapter.env | 6 +++--- common/compose.yaml | 8 +++++--- .../specs/2026-09-06-deployable-stack-design.md | 2 +- tests/compose.bats | 2 +- 4 files changed, 10 insertions(+), 8 deletions(-) diff --git a/adapters/nextjs/adapter.env b/adapters/nextjs/adapter.env index 4bdc652..a9fbf7c 100644 --- a/adapters/nextjs/adapter.env +++ b/adapters/nextjs/adapter.env @@ -11,7 +11,7 @@ ADAPTER_GENERATOR='pnpm create next-app@16 "$APP_DIR" --ts --app --eslint --tail # formatted — the contract's format and test tasks need both. ADAPTER_POST_GENERATE='pnpm add -D prettier vitest && pnpm exec prettier --write .' # A route handler answers without rendering the home page, so the probe does -# not depend on whatever the client later puts on `/`. The route itself is -# created by the next task; declaring the final value here avoids this task -# shipping a comment that the next task makes false. +# not depend on whatever the client later puts on `/`. The route lives at +# src/app/api/health/live/route.ts (copied in by apply_adapter's directory +# loop, not this adapter's flat file list). ADAPTER_LIVENESS_PATH="/api/health/live" diff --git a/common/compose.yaml b/common/compose.yaml index 1b3b472..6ad6652 100644 --- a/common/compose.yaml +++ b/common/compose.yaml @@ -14,9 +14,11 @@ services: app: # ghcr.io/CHANGEME/CHANGEME is a placeholder: scaffold generates this # project before it has a github repository, so it cannot know its own - # registry path. replace it once, by hand, after the repository exists - # and its first image has been published (see the scaffold toolbox's - # ADR-0014, not shipped here). + # registry path. install.sh re-downloads this file from the latest + # release on every run, so a hand-edit to a deployed copy is undone the + # next time it runs — fix it in this repository's own compose.yaml and + # cut a release instead (see the scaffold toolbox's + # docs/runbook/first-project-walkthrough.md step 10, not shipped here). image: ghcr.io/CHANGEME/CHANGEME:${IMAGE_TAG:-latest} env_file: # required: false so this validates before a .env exists; install.sh diff --git a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md index b1f856d..e8d5c9c 100644 --- a/docs/superpowers/specs/2026-09-06-deployable-stack-design.md +++ b/docs/superpowers/specs/2026-09-06-deployable-stack-design.md @@ -222,7 +222,7 @@ deploy gate curls both. | `laravel-api` | `/up` (shipped by Laravel since 11.x) | `/health/ready` | | `laravel-inertia` | `/up` | `/health/ready` | | `nestjs` | `/health/live` | `/health/ready` | -| `nextjs` | `/` (the generated home page) | none — the `web` role takes no database driver | +| `nextjs` | `/api/health/live` | none — the `web` role takes no database driver | **Two of the four are wrong today, in the same way.** `nestjs`'s Dockerfile probes `http://localhost:3001/health`, and no adapter ships a `/health` route diff --git a/tests/compose.bats b/tests/compose.bats index 35f0367..e1ecaee 100644 --- a/tests/compose.bats +++ b/tests/compose.bats @@ -174,7 +174,7 @@ INNER_EOF path="$(grep '^ADAPTER_LIVENESS_PATH=' "${dir}adapter.env" | cut -d'"' -f2)" for file in "${dir}"Dockerfile "${dir}"Dockerfile.workspace; do [ -f "$file" ] || continue - grep -q "HEALTHCHECK" "$file" \ + grep -q '^HEALTHCHECK' "$file" \ || { wrong="${wrong}${file}: no HEALTHCHECK"$'\n'; continue; } # localhost or 127.0.0.1: nextjs's HEALTHCHECK dials 127.0.0.1 because # this image's resolver hands "localhost" the IPv6 ::1 first and the From 27d1df5d50439aa6446fd42958a9e536da769919 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 09:32:30 +0700 Subject: [PATCH 24/30] fix: correct two false claims about what is enforced MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit tests/service.bats claimed the removed empty-block skip had closed a hole — it had not: an empty block matches neither grep in the check, so the skip's removal changed no behaviour. docs/decisions/0021 named three things the deploy gate does not share with common/install.sh's own logic, and left out the one that matters most: the gate does not call install.sh's run_migrations at all, it carries a second implementation that has already drifted from it (install.sh falls back to grepping for a database service, the gate falls back to ROLE/DB_SERVICE). --- docs/decisions/0021-the-released-stack-must-run.md | 9 ++++++++- tests/service.bats | 8 +++----- 2 files changed, 11 insertions(+), 6 deletions(-) diff --git a/docs/decisions/0021-the-released-stack-must-run.md b/docs/decisions/0021-the-released-stack-must-run.md index f774281..3fdfaa7 100644 --- a/docs/decisions/0021-the-released-stack-must-run.md +++ b/docs/decisions/0021-the-released-stack-must-run.md @@ -35,7 +35,14 @@ the existing `smoke` lane:** but it does call `install.sh`'s own `generate_service_passwords` on the copied `.env`, so the password loop and the `APP_KEY` branch run under the same substitution a client's install would perform, not against - every credential left at `changeme`. + every credential left at `changeme`. It also does not call + `install.sh`'s `run_migrations`: the gate carries its own second + implementation (`scripts/deploy-check.sh`'s own migrate block), and the + two have already drifted — `install.sh` falls back to checking for a + `database` service when no `migrate` service is found, the gate falls + back to the adapter's own `ROLE`/`DB_SERVICE`. Unifying them is deferred + until a real-project run has exercised `install.sh` against a published + release. **The container port is fixed at 8080, not a variable.** Every adapter serves HTTP on container port 8080; `common/compose.yaml` publishes diff --git a/tests/service.bats b/tests/service.bats index bfbe735..a52bd2d 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -592,11 +592,9 @@ EOF # the URL at `up` time. A literal baked here is the changeme-versus-app # mismatch that made the dev stack unable to authenticate. # - # No skip for an empty block: an empty block has no literal password by - # construction, so the checks below already cover it without a special - # case — the previous `[ -z "$block" ] && continue` let a driver that - # regressed to emitting nothing pass unseen, for a reason unrelated to - # passwords. + # No skip for an empty block: an empty block matches neither grep below, + # so this loop already treats "emits nothing" as "nothing to flag" without + # a special case for it. # # Asserts the absence of a literal, not the presence of an interpolation: # a block could carry `${DB_PASSWORD}` somewhere else and a hardcoded From 0c800128df0a76b0e93f22f23026e31c74bc7275 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 09:33:06 +0700 Subject: [PATCH 25/30] test: prove the literal-password check can fail MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both while loops in the password check exist only for the eight real shipping drivers, and every one of those already writes an interpolation. Deleting the loops leaves the suite green — the same "gate that cannot fail" shape task 0a4976d fixed for the driver-function loop. Factors the check into _password_literal_report so both the real-driver sweep and a new fixture test call the identical logic, then adds a fixture driver that bakes in a literal password and a test asserting the check reports it. Mutation-proved: deleting both loops' bodies left the real-driver test green (an empty report still satisfies `[ -z "$bad" ]`) but failed the new fixture test; restoring the loops made it pass again. --- .../sample/drivers/laravel.sh | 8 +++ tests/service.bats | 64 +++++++++++++------ 2 files changed, 52 insertions(+), 20 deletions(-) create mode 100644 tests/fixtures/lint-services/literal-password/sample/drivers/laravel.sh diff --git a/tests/fixtures/lint-services/literal-password/sample/drivers/laravel.sh b/tests/fixtures/lint-services/literal-password/sample/drivers/laravel.sh new file mode 100644 index 0000000..abb29c5 --- /dev/null +++ b/tests/fixtures/lint-services/literal-password/sample/drivers/laravel.sh @@ -0,0 +1,8 @@ +# shellcheck shell=bash +# Exists only to prove tests/service.bats' literal-password check can fail: +# a check exercised solely by drivers already written to pass it is a check +# that cannot fail, which is the same gap this fixture's sibling +# missing-driver-function closes for the driver-function loop. +service_driver_compose_env() { + printf 'DB_PASSWORD: hunter2\n' +} diff --git a/tests/service.bats b/tests/service.bats index a52bd2d..418ce2a 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -587,6 +587,35 @@ EOF || { echo "allowBuilds (line ${allow_builds}) must come before the first pnpm add (line ${first_add})"; false; } } +# Shared by the two tests below, so a regression in the check itself fails +# both: a copy of this logic kept only in the fixture test could no-op right +# alongside a broken check while still reporting green on its own. +_password_literal_report() { + local driver="$1" block bad="" + block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" + SERVICE_DIR="$(dirname "$(dirname "$driver")")" + . "$driver"; service_driver_compose_env )" + + # A top-level *_PASSWORD key whose value is not exactly an interpolation. + while IFS= read -r line; do + case "$line" in + *_PASSWORD:\ \$\{*_PASSWORD\}) ;; + *) bad="${bad}${driver} (${line})"$'\n' ;; + esac + done < <(grep -E '^[A-Za-z_]*_PASSWORD:' <<<"$block") + + # A DSN's user:password@ slot whose password is not exactly an + # interpolation — the same shape, embedded in a URL instead of a key. + while IFS= read -r segment; do + case "$segment" in + :\$\{*_PASSWORD\}@) ;; + *) bad="${bad}${driver} (${segment})"$'\n' ;; + esac + done < <(grep -oE ':[^:@]*@' <<<"$block") + + printf '%s' "$bad" +} + @test "a driver's compose environment never bakes a literal password" { # The password must exist in exactly one place — .env — so compose composes # the URL at `up` time. A literal baked here is the changeme-versus-app @@ -602,30 +631,25 @@ EOF # check could not tell the two apart. local bad="" for driver in "${SCAFFOLD_ROOT}"/services/*/drivers/*.sh; do - block="$( . "${SCAFFOLD_ROOT}/lib/service.sh" - SERVICE_DIR="$(dirname "$(dirname "$driver")")" - . "$driver"; service_driver_compose_env )" - - # A top-level *_PASSWORD key whose value is not exactly an interpolation. - while IFS= read -r line; do - case "$line" in - *_PASSWORD:\ \$\{*_PASSWORD\}) ;; - *) bad="${bad}${driver} (${line})"$'\n' ;; - esac - done < <(grep -E '^[A-Za-z_]*_PASSWORD:' <<<"$block") - - # A DSN's user:password@ slot whose password is not exactly an - # interpolation — the same shape, embedded in a URL instead of a key. - while IFS= read -r segment; do - case "$segment" in - :\$\{*_PASSWORD\}@) ;; - *) bad="${bad}${driver} (${segment})"$'\n' ;; - esac - done < <(grep -oE ':[^:@]*@' <<<"$block") + bad="${bad}$(_password_literal_report "$driver")" done [ -z "$bad" ] || { echo "embeds a literal password:"; echo "$bad"; false; } } +@test "the literal-password check reports a driver that bakes one in" { + # The test above only proves the check accepts what ships today — deleting + # its loops leaves that test green too, which is the same "gate that + # cannot fail" shape the missing-driver-function fixture exists to rule + # out for the driver-function loop. This drives the identical check + # against a fixture driver that hardcodes a password, so a regression to + # "matches nothing" fails here even while every real driver still passes. + local driver="${SCAFFOLD_ROOT}/tests/fixtures/lint-services/literal-password/sample/drivers/laravel.sh" + local bad + bad="$(_password_literal_report "$driver")" + [[ "$bad" == *"DB_PASSWORD: hunter2"* ]] \ + || { echo "expected a literal password to be reported, got:"; echo "$bad"; false; } +} + @test "apply_service_compose_env merges into the app service" { local project="${BATS_TEST_TMPDIR}/p" mkdir -p "$project" From 7631df718fb1155fd5ad07d6444fd274d1ee6c50 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 09:33:45 +0700 Subject: [PATCH 26/30] fix: anchor the password-key grep past leading whitespace ^[A-Za-z_]*_PASSWORD: required the key at column zero, so an indented key ( DB_PASSWORD: hunter2) never reached the case check that would have flagged its literal value. The case pattern itself already tolerated a prefix via its leading *; only the grep that selects candidate lines needed the same tolerance. Left PGPASSWORD/DB_PASS/APP_KEY and friends alone, in a comment: a name blacklist is never complete, and the drivers are the only writers of this block and already go through review. --- tests/service.bats | 12 ++++++++++-- 1 file changed, 10 insertions(+), 2 deletions(-) diff --git a/tests/service.bats b/tests/service.bats index 418ce2a..473a702 100644 --- a/tests/service.bats +++ b/tests/service.bats @@ -596,13 +596,21 @@ _password_literal_report() { SERVICE_DIR="$(dirname "$(dirname "$driver")")" . "$driver"; service_driver_compose_env )" - # A top-level *_PASSWORD key whose value is not exactly an interpolation. + # A *_PASSWORD key whose value is not exactly an interpolation. Anchored + # with optional leading whitespace, not a bare ^, so an indented key still + # gets selected for the case check below. + # + # Only *_PASSWORD, deliberately: PGPASSWORD, DB_PASS, and APP_KEY's own + # base64 secret would also slip past this, but a name blacklist is never + # complete, and the services/*/drivers/*.sh files are the only writers of + # this block and already go through review — widen the blacklist here and + # the next unlisted name just becomes the new hole. while IFS= read -r line; do case "$line" in *_PASSWORD:\ \$\{*_PASSWORD\}) ;; *) bad="${bad}${driver} (${line})"$'\n' ;; esac - done < <(grep -E '^[A-Za-z_]*_PASSWORD:' <<<"$block") + done < <(grep -E '^[[:space:]]*[A-Za-z_]*_PASSWORD:' <<<"$block") # A DSN's user:password@ slot whose password is not exactly an # interpolation — the same shape, embedded in a URL instead of a key. From 344b29b47243b130096dfc8090f26e2f5298ede0 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 09:34:38 +0700 Subject: [PATCH 27/30] fix: drop the redundant REDIS_PASSWORD line from redis's laravel driver MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Measured: the block passes the literal-password check with the line and without it — env_file already delivers REDIS_PASSWORD to the container from the same .env compose interpolates from, so restating it here was dead weight. mysql and postgres's laravel drivers dropped the identical DB_PASSWORD line already; redis was the last holdout. Done after the literal-password fixture (tests/fixtures/lint-services/ literal-password) landed: this line was the only real driver exercising the key-form half of the check, so removing it first would have left that half checking nothing live. --- services/redis/drivers/laravel.sh | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/services/redis/drivers/laravel.sh b/services/redis/drivers/laravel.sh index fbf9b65..814804c 100644 --- a/services/redis/drivers/laravel.sh +++ b/services/redis/drivers/laravel.sh @@ -30,12 +30,11 @@ service_driver_dockerfile() { # REDIS_PASSWORD already reaches the container via compose.yaml's env_file # (it is in the project's example.env, assembled from this service's own -# env.fragment) — restated here, interpolated rather than baked, so the -# service_driver_compose_env test that guards against a literal password can -# see it the same way the URL-based drivers show theirs. +# env.fragment) — restating it here would only be redundant, the same +# reasoning that already dropped DB_PASSWORD from the mysql and postgres +# laravel drivers. service_driver_compose_env() { printf 'REDIS_HOST: cache\n' - printf 'REDIS_PASSWORD: ${REDIS_PASSWORD}\n' } # a cache has no schema to migrate — printing nothing keeps the migrate From 25f9f432b6923a78469e86b5639a2043cefeb41d Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 13:53:11 +0700 Subject: [PATCH 28/30] fix: satisfy Laravel Pint on the health route fully_qualified_strict_types rejected the shipped \RuntimeException and \Throwable now that pint runs in CI. Fixing the source template alone was not enough: the mysql/postgres and mongodb drivers splice a fully qualified DB:: probe into the same file, which pint also rejects, and which then needs its own use import and a blank line before the following return (blank_line_before_statement). Verified against laravel-api's own generated project with pint's writing mode, then against tests/new-laravel-api.bats and tests/new-laravel-inertia.bats. --- adapters/laravel-api/routes/health.php | 4 ++-- adapters/laravel-inertia/routes/health.php | 4 ++-- services/mongodb/drivers/laravel.sh | 16 ++++++++++++++-- services/shared/laravel.sh | 16 ++++++++++++++-- 4 files changed, 32 insertions(+), 8 deletions(-) diff --git a/adapters/laravel-api/routes/health.php b/adapters/laravel-api/routes/health.php index c607230..a113445 100644 --- a/adapters/laravel-api/routes/health.php +++ b/adapters/laravel-api/routes/health.php @@ -14,8 +14,8 @@ Route::get('/health/ready', function () { try { // @DB_PROBE@ - throw new \RuntimeException('no database is configured for this project'); - } catch (\Throwable $e) { + throw new RuntimeException('no database is configured for this project'); + } catch (Throwable $e) { return response()->json(['status' => 'unavailable', 'reason' => $e->getMessage()], 503); } }); diff --git a/adapters/laravel-inertia/routes/health.php b/adapters/laravel-inertia/routes/health.php index c607230..a113445 100644 --- a/adapters/laravel-inertia/routes/health.php +++ b/adapters/laravel-inertia/routes/health.php @@ -14,8 +14,8 @@ Route::get('/health/ready', function () { try { // @DB_PROBE@ - throw new \RuntimeException('no database is configured for this project'); - } catch (\Throwable $e) { + throw new RuntimeException('no database is configured for this project'); + } catch (Throwable $e) { return response()->json(['status' => 'unavailable', 'reason' => $e->getMessage()], 503); } }); diff --git a/services/mongodb/drivers/laravel.sh b/services/mongodb/drivers/laravel.sh index 8580ed3..37faf8c 100644 --- a/services/mongodb/drivers/laravel.sh +++ b/services/mongodb/drivers/laravel.sh @@ -47,9 +47,21 @@ service_driver_apply() { # always returns first. The throw is replaced in place instead, so a # --db none project keeps it — unreachable in no project this driver ever # touches. - sed -i.bak 's|// @DB_PROBE@|\\Illuminate\\Support\\Facades\\DB::connection(\x27mongodb\x27)->getMongoDB()->command([\x27ping\x27 => 1]);|' \ + # + # The probe is spliced in as a short class name with its own `use` added + # here, not the FQCN a --db none project ships: pint's + # fully_qualified_strict_types rejects an inline FQCN once the file already + # has imports, and a --db none project never runs this substitution (or + # carries an import it would leave unused). + sed -i.bak 's|use Illuminate\\Support\\Facades\\Route;|use Illuminate\\Support\\Facades\\DB;\nuse Illuminate\\Support\\Facades\\Route;|' \ + routes/health.php || return 1 + sed -i.bak 's|// @DB_PROBE@|DB::connection(\x27mongodb\x27)->getMongoDB()->command([\x27ping\x27 => 1]);|' \ routes/health.php || return 1 - sed -i.bak "s|throw new \\\\RuntimeException('no database is configured for this project');|return response()->json(['status' => 'ok']);|" \ + # Matched with its leading indentation so the replacement's `\n` opens a + # bare blank line rather than one trailing the throw statement's own + # indentation — pint's blank_line_before_statement wants a blank line + # between the probe call above and this return. + sed -i.bak "s| throw new RuntimeException('no database is configured for this project');|\\n return response()->json(['status' => 'ok']);|" \ routes/health.php || return 1 rm -f routes/health.php.bak diff --git a/services/shared/laravel.sh b/services/shared/laravel.sh index ab2c452..2ffaf99 100644 --- a/services/shared/laravel.sh +++ b/services/shared/laravel.sh @@ -50,9 +50,21 @@ service_driver_apply() { # always returns first. The throw is replaced in place instead, so a # --db none project keeps it — unreachable in no project this driver ever # touches. - sed -i.bak 's|// @DB_PROBE@|\\Illuminate\\Support\\Facades\\DB::connection()->select(\x27select 1\x27);|' \ + # + # The probe is spliced in as a short class name with its own `use` added + # here, not the FQCN a --db none project ships: pint's + # fully_qualified_strict_types rejects an inline FQCN once the file already + # has imports, and a --db none project never runs this substitution (or + # carries an import it would leave unused). + sed -i.bak 's|use Illuminate\\Support\\Facades\\Route;|use Illuminate\\Support\\Facades\\DB;\nuse Illuminate\\Support\\Facades\\Route;|' \ + routes/health.php || return 1 + sed -i.bak 's|// @DB_PROBE@|DB::connection()->select(\x27select 1\x27);|' \ routes/health.php || return 1 - sed -i.bak "s|throw new \\\\RuntimeException('no database is configured for this project');|return response()->json(['status' => 'ok']);|" \ + # Matched with its leading indentation so the replacement's `\n` opens a + # bare blank line rather than one trailing the throw statement's own + # indentation — pint's blank_line_before_statement wants a blank line + # between the probe call above and this return. + sed -i.bak "s| throw new RuntimeException('no database is configured for this project');|\\n return response()->json(['status' => 'ok']);|" \ routes/health.php || return 1 rm -f routes/health.php.bak From a54ab32c6c2bd7672d219c78243d5827a1f391d5 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 14:01:09 +0700 Subject: [PATCH 29/30] fix: give deploy-check.sh its own git identity MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit scaffold new commits the project it generates, so it needs a git identity. tests/helpers/setup.bash sets one up for bats, but this script runs outside bats and never got it — every deploy/deploy-tier-b job failed in seconds with "git has no user.name". Reproduced with GIT_CONFIG_GLOBAL unset and HOME pointed at an empty sandbox (no fallback identity anywhere), matching the runner; fixed the same way setup.bash does, and only when a global isn't already configured, so a developer with a real identity keeps theirs. --- scripts/deploy-check.sh | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh index d9dc15a..77ebb0b 100755 --- a/scripts/deploy-check.sh +++ b/scripts/deploy-check.sh @@ -69,6 +69,18 @@ TMP_DIR="$(mktemp -d)" PROJECT_DIR="${TMP_DIR}/demo" IMAGE_TAG="deploy-check/${ADAPTER}:local" +# A runner has no git identity either, and `scaffold new` commits what it +# creates — tests/helpers/setup.bash gives bats the same thing, but this +# script runs outside bats and never picked it up. Owned by this run rather +# than written into a real global config; skipped when one is already set, +# so a developer with a real identity keeps theirs. +if [ -z "${GIT_CONFIG_GLOBAL:-}" ]; then + GIT_CONFIG_GLOBAL="${TMP_DIR}/gitconfig" + export GIT_CONFIG_GLOBAL + git config --global user.name "deploy-check" + git config --global user.email "deploy-check@scaffold.invalid" +fi + # Every generated project's compose.yaml is `name: app` (common/compose.yaml) # — without this, a local run reconciles against, and `down -v`s, any real # "app" project already running on this machine, database volumes included. From 27bb6880513520f5038244e5ca8446528e108ff8 Mon Sep 17 00:00:00 2001 From: iam-truongtrungnghia Date: Mon, 7 Sep 2026 14:20:42 +0700 Subject: [PATCH 30/30] fix: give deploy-check.sh a GitHub owner and its own mise trust store MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The git-identity fix revealed the next thing tests/helpers/setup.bash supplies that this script, running scaffold new outside bats, never picked up: a GitHub account for resolve_github_owner's guard, which a runner has no gh login or git config github.user to satisfy — every deploy/deploy-tier-b job died in seconds on "no GitHub account to substitute for 'you/'". Fixed the same way as the git identity, plus the trust store setup.bash also isolates for the same reason: mise trust records every config it trusts in real state keyed by path, and a repeated local run of this script would otherwise grow the developer's machine the same way bats' suites had, past 7600 stale entries. Reproduced by scrubbing only GH_CONFIG_DIR and SCAFFOLD_GITHUB_OWNER (not HOME, which moves mise's own data dir and breaks yq resolution instead of reproducing the bug) and confirmed the same error the CI log shows; fixed, confirmed nextjs passes clean, and confirmed nestjs --db postgres still passes. --- scripts/deploy-check.sh | 29 ++++++++++++++++++++++++----- 1 file changed, 24 insertions(+), 5 deletions(-) diff --git a/scripts/deploy-check.sh b/scripts/deploy-check.sh index 77ebb0b..4ffdbcc 100755 --- a/scripts/deploy-check.sh +++ b/scripts/deploy-check.sh @@ -69,11 +69,15 @@ TMP_DIR="$(mktemp -d)" PROJECT_DIR="${TMP_DIR}/demo" IMAGE_TAG="deploy-check/${ADAPTER}:local" -# A runner has no git identity either, and `scaffold new` commits what it -# creates — tests/helpers/setup.bash gives bats the same thing, but this -# script runs outside bats and never picked it up. Owned by this run rather -# than written into a real global config; skipped when one is already set, -# so a developer with a real identity keeps theirs. +# `scaffold new` needs an identity, an account, and a trust store that a +# runner has none of on its own — tests/helpers/setup.bash hands bats all +# three for exactly this reason, but this script runs outside bats and +# never picked any of them up. Each is owned by this run rather than +# written into real state, and skipped when the caller already supplied +# one, so a developer with a real identity, account, or trust store keeps +# theirs. + +# scaffold new commits what it creates, and git refuses without an identity. if [ -z "${GIT_CONFIG_GLOBAL:-}" ]; then GIT_CONFIG_GLOBAL="${TMP_DIR}/gitconfig" export GIT_CONFIG_GLOBAL @@ -81,6 +85,21 @@ if [ -z "${GIT_CONFIG_GLOBAL:-}" ]; then git config --global user.email "deploy-check@scaffold.invalid" fi +# resolve_github_owner (lib/project.sh) substitutes this for the generated +# workflows' placeholder `you/` account, and falls back to `gh auth login` +# or git's github.user before giving up — a runner has none of the three. +export SCAFFOLD_GITHUB_OWNER="${SCAFFOLD_GITHUB_OWNER:-deploy-check}" + +# mise records every config it trusts (`mise trust`, below) under its state +# directory keyed by path; a throwaway project dir trusted here has no +# reason to outlive this run, and a repeated local run would otherwise grow +# the developer's real store the way tests/helpers/setup.bash found bats +# had — past 7600 stale entries. +if [ -z "${MISE_STATE_DIR:-}" ]; then + MISE_STATE_DIR="${TMP_DIR}/mise-state" + export MISE_STATE_DIR +fi + # Every generated project's compose.yaml is `name: app` (common/compose.yaml) # — without this, a local run reconciles against, and `down -v`s, any real # "app" project already running on this machine, database volumes included.