feat: make a generated project's released stack actually run - #7
Merged
Merged
Conversation
An acceptance run started a generated project's released stack for the first time. It does not work, and never has, in any shape: no application is told how to reach its database, container port 8080 matches no adapter in this repository, the Laravel images end at php-fpm with no web server anywhere in the stack, and nestjs has been probing a /health route nothing serves. None of it is a regression. ADR-0014 seam 2 already required that every value a container needs arrives through env_file/environment; nothing was ever made to produce the values. It survived because no check could fail — tests/compose.bats runs `docker compose config --quiet`, which validates YAML, not reachability, and no test starts a container. The design turns "green immediately, deployable immediately" into two mechanically checkable gates, and settles the three decisions that follow: every adapter serves on container port 8080; the Laravel images move to FrankenPHP alpine pinned by digest; and the app's connection variables are composed in compose.yaml from the service variables already in .env, by the driver, which is the only thing that knows both the service and the adapter family. It also moves ADR-0003's boundary. An adapter has never written application code; it now writes one readiness route, because a liveness path that does not touch the database is the check-that-cannot-fail ADR-0014 seam 4 warned about, and it is what let all of this ship. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Ten tasks, each ending in something independently testable, and each container check using `docker compose up -d` rather than `docker run`, which this environment denies. Writing it surfaced two things the spec had wrong, both fixed in the spec as well rather than papered over in the plan: The published nestjs image cannot migrate itself. The driver installs the CLI with `pnpm add -D prisma@6` and the Dockerfile runs `pnpm prune --prod` before the runtime stage copies node_modules, so the image carries @prisma/client and no prisma binary. The driver now installs it as a regular dependency, which is what Prisma's own guidance assumes when the migration runs from the image, and the image-size cost is recorded rather than discovered later. And install.sh cannot run a mise task, because no image carries mise. The driver writes a profiled `migrate` compose service instead, sharing the app's image and environment, so it never starts with the stack — ADR-0014 seam 5 forbids migrating from an entrypoint, not from the operator's one command. Also removes a dangling ADR-0021 reference from both documents. The decision record is Task 8's output, and tests/documentation.bats was right to fail on a spec citing a file that does not exist yet. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
ADAPTER_LIVENESS_PATH (every adapter) and ADAPTER_READINESS_PATH (adapters whose role is in DRIVEN_ROLES) are now declared once per adapter and read by lib/lint.sh, so the Dockerfile and the CI gate stop each keeping their own idea of the route — the defect that let nestjs probe a /health nothing served. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
The case branch in lint_adapters that requires ADAPTER_READINESS_PATH only for DRIVEN_ROLES had no fixture exercising it: every existing lint fixture is role api, and the tests that did touch readiness paths checked the real adapter.env files, not the enforcement code, so deleting the branch would have gone unnoticed. Adds a role-api fixture missing the var (must be reported) and a role-web fixture missing it (must stay silent), each with its own test. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
nestjs and nextjs both listened on their generator's default port while common/compose.yaml has always published 8080 — nestjs additionally probed a /health route no adapter ever shipped, so its image reported unhealthy from first boot. Both adapters now listen on 8080, ship a real liveness route, and their HEALTHCHECKs probe it there. apply_adapter's docker/-only special case is now a loop over every directory an adapter ships, since nestjs's health routes and nextjs's health route both need the same directory-delivery mechanism the Laravel pair already used for docker/opcache.ini. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
sed -i "1i ..." always succeeds regardless of match, and the following substitution for imports: [] silently no-ops if the nest generator ever reformats app.module.ts — leaving HealthModule unregistered with no build or lint failure, only a 404 on /health/live discovered by the HEALTHCHECK this task just pointed at that route. Grep for both halves of the wiring after the sed calls and fail loudly, naming the file and the likely cause, so a generator reformat cannot silently re-arm the unhealthy-image defect this task exists to remove. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Replace the php-fpm runtime stage in both laravel adapters with FrankenPHP (alpine, pinned by digest), so the images actually speak HTTP instead of FastCGI with nothing in front of them. Adds SERVER_NAME, the XDG_* pair Caddy needs to start as www-data, ownership fixes for storage/ and bootstrap/cache, and a real HEALTHCHECK against /up. The base image's default CMD carries --config/--adapter flags that point frankenphp at its Caddyfile; a bare `frankenphp run` drops them and starts only the admin API, so CMD restates them explicitly. laravel-inertia also moves COPY . . before the assets copy so a developer's local public/build can no longer shadow the one the assets stage built, and .dockerignore excludes public/build from the build context for the same reason. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
compose.yaml only ever received DB_DATABASE/DB_USERNAME/DB_PASSWORD, and neither Prisma nor Laravel reads those directly, so a released stack could not reach its database. apply_service_compose_env merges a driver's service_driver_compose_env block into compose.yaml's app service with `yq -P` so the merge keeps block style and the file's comments survive. Every driver is required to define the hook (REQUIRED_DRIVER_FUNCTIONS, checked by lint_services), so a driver shipping only two of the three fails at generation or at lint instead of silently producing an app with no DATABASE_URL. No driver implements the hook yet — that is task 5. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
lint_services sourced each driver with no SERVICE_DIR set, unlike every
other call site that sources a driver — a Task-5 driver reading it at
sourcing time would die under the inherited `set -u` and get misreported as
"does not define" the function it never got a chance to declare. Set it the
way load_service does, and capture the subshell's stderr so a sourcing
failure reports its own text instead of being folded into the
missing-function message.
Dropped compose-env's `-P`: measured, not assumed. The fragment it merges is
always this function's own block-style output, never the flow style
merge_lefthook_fragment's `-P` guards against, so there is nothing here for
it to protect. Merging without it left compose.yaml's untouched lines
byte-identical; with it, yq rewrote nodes this merge never touched — an
unquoted `${APP_PORT:-8080}:8080` and a healthcheck array reshaped from flow
to block. Both happened to pass prettier --check, but a merge with no
business editing those lines should not be reshaping them anyway.
Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
All eight drivers now implement service_driver_compose_env: nest gets DATABASE_URL/REDIS_URL built from the same DB_*/REDIS_* variables the database/cache containers read, operator-overridable via .env; laravel gets DB_CONNECTION plus whatever host/port/DSN that family needs, since config/database.php defaults to sqlite and its absence was a silent wrong answer, not a failure. Every interpolation reads the password rather than embedding it, so it exists in exactly one place. Both frameworks also get a real readiness probe, spliced into the anchor each ships pre-generation. Two substitutions, not one, in both: the anchor becomes the probe, and the shipped fallback (`throw`) is replaced in place rather than left as dead code below an added return — nest's controller would otherwise fail eslint's no-unreachable under --max-warnings 0. Nest's probe is cast to an explicit method signature instead of a bare dynamic import, because before `prisma generate` runs (lint runs first) the generated client module doesn't exist, and an untyped access to it is `any` under the no-unsafe-* rules; only the one method the provider actually calls is declared, since a real PrismaClient's mongodb build has no $queryRawUnsafe and its SQL builds have no $runCommandRaw. Laravel needed a route to splice into. Both laravel adapters now ship routes/health.php, wired into bootstrap/app.php's withRouting(then: ...) — laravel only auto-loads routes/web.php and routes/console.php, so a file dropped there with nothing pointing at it 404s forever. APP_KEY, per-family rather than per-service, is appended to the project's example.env by every laravel driver, since no service env.fragment can carry it and laravel will not boot without one. Verified live: generating a laravel+mongodb project and hitting /health/ready surfaced a real bug independent of this task — `pecl install mongodb` resolves latest (2.x) while mongodb/laravel-mongodb is built against the 1.x extension's BSON model classes, so any code touching MongoDB\Client fatals before reaching this project's own try/catch. Pinned to 1.21.0, matching the platform.ext-mongodb declaration already beside it. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
install.sh now generates a laravel-valid APP_KEY (base64: + 32 bytes, handled in generate_service_passwords' own loop, with the sed delimiter changed to | since a base64 value can contain /) and runs a profiled migrate compose service after the stack is up, before the success message. Every service driver now implements service_driver_compose_migrate (added to REQUIRED_DRIVER_FUNCTIONS), and a new apply_service_compose_service merges a driver's complete services: fragment — apply_service_compose_env cannot do this itself, since it hardcodes services.app.environment. The migrate service shares the app's image (read back off compose.yaml) and its environment. install.sh's guard for whether a migrate service exists uses `docker compose --profile migrate config --services`, not the profile-less form, which never lists a profiled service at all. The nest driver now installs prisma as a regular dependency, not -D, since the Dockerfile's `pnpm prune --prod` would otherwise strip the CLI the migrate service needs out of the published image. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
The compose migrate command written by service_driver_compose_migrate used `pnpm exec`, but the runtime image (adapters/nestjs/Dockerfile[.workspace]) never installs pnpm — only the build stage runs corepack enable, and the runtime stage copies node_modules/dist alone. Measured against a real built image: `which pnpm` exits 1. prisma's own bin does survive `pnpm prune --prod` (it is a regular dependency for exactly this), but at a location that depends on which Dockerfile shape wins, a decision made after this driver runs: apps/<app>/node_modules/.bin for the typescript-workspace shape, node_modules/.bin at the container root for the standalone one. The new command tries both, and cd's into whichever matched before running it, since prisma resolves its schema relative to its own cwd and that schema lives nested under the app directory in the workspace shape. That surfaced a second defect once the binary was found: `schema.prisma` was never copied into the runtime image at all, in either Dockerfile shape, since it is not an artifact `nest build` produces. Both Dockerfiles now copy prisma/ into the runtime stage via a bracket glob (`pris[m]a`), which copies nothing rather than failing the build on a --db none project that has no prisma/ directory to copy. Also fixed: `$` in the compose-yaml command needs to be `$$` to survive compose's own variable interpolation before the container ever sees it — a single `$` here silently resolved to an unset variable and blanked the whole loop, caught by inspecting `docker compose config`'s resolved output before trusting it. Verified on a real nestjs+postgres stack: readiness 200 before and after migrating (the SQL probe is connectivity-only, matching Task 5's own documented finding — left alone per this round's ruling), and `docker compose --profile migrate run --rm migrate` exits 0 and reaches the real database. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Everything before this validated compose.yaml as YAML; nothing ever started a container, which is how a stack that could never serve a request shipped for weeks. scripts/deploy-check.sh generates an adapter, builds its released image, brings the compose stack up, waits for the app container to report healthy, asserts the migrate profile service exits 0 (the readiness probe alone proves connectivity, not schema), then asserts both the liveness and declared readiness path return 200 — tearing the stack down on every one of those paths via a trap, not a trailing cleanup line. Both paths are read out of the adapter's own adapter.env, never hardcoded, and readiness is skipped only where the adapter declares none, with that skip stated in the output. adapters.yml gains deploy/deploy-tier-b jobs mirroring smoke/smoke-tier-b's discover-driven matrices, running this script per tier-a and tier-b adapter. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Task 7's deploy gate found this on its first real run: Docker sets HOSTNAME to the container's own id for every container, and standalone server.js binds to `process.env.HOSTNAME || '0.0.0.0'` — so without an explicit override, the app listened on that id-derived address, reachable through compose's published port (routed to the container's real interface either way) but not from the HEALTHCHECK, which dials localhost from inside the same container and got "connection refused" forever. ENV HOSTNAME="0.0.0.0" fixes the bind, but exposed a second, previously masked mismatch: that bind is IPv4-only, while this image's resolver hands wget the ::1 (IPv6) address for "localhost" first, and busybox wget does not fall back to the IPv4 result. Measured with `docker run` at each step. Pointing the HEALTHCHECK at 127.0.0.1 instead closes that gap. PORT was already correct — `ss -tlnp` showed the server bound on :8080 throughout; the defect was only ever the bind address, not the port. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Review round 2. Critical: the migrate check gated on a condition derived from the artifact under test — a renamed service, a broken profile, or a driver returning no command all read as "no database", the readiness probe (select 1) then returns 200 against unapplied schema, and the run exits 0. ROLE and DB_SERVICE are already known, so absence is now a failure whenever a database was actually requested. common/install.sh's run_migrations had the identical hole (a database service with no migrate beside it now fails loudly instead of returning 0), since it ships to every generated project. Also, in order of how they'd bite: - COMPOSE_PROJECT_NAME is set per adapter — every generated stack is `name: app` (common/compose.yaml), so a local run used to reconcile against, and `down -v` any real "app" project already on the machine. - ADAPTER_LIVENESS_PATH and ADAPTER_READINESS_PATH are asserted non-empty: lib/lint.sh only checks the line is present, and an empty liveness path silently probes "/", which nextjs happens to answer 200 for reasons unrelated to its real liveness. - compose.yaml's image is asserted equal to the tag just built, not just "no CHANGEME left" — a real registry reference in common/compose.yaml would match neither check and the stack would come up on a pulled image. - notify-on-schedule-failure now watches deploy and deploy-tier-b too. - APP_PORT's read no longer dies silently under pipefail when a .env lacks it; the health poll uses `docker inspect` instead of `docker compose ps --format json`, whose shape is compose-version-dependent; the cleanup trap now covers INT/TERM, not just EXIT. tests/compose.bats' liveness-probe test is updated to accept 127.0.0.1 alongside localhost, matching fix round 1's nextjs HEALTHCHECK change. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Task 8 of the deployable-stack plan: ADR-0014 seams 1 and 4, ADR-0003's no-application-code boundary, the tour's healthcheck example, the runbook's pull-only step 10, and the design spec's overclaimed readiness proof were all falsified by the work in commits c61600f..0b59f6a. Records the corrected boundary in a new ADR-0021, marks the superseded parts of ADR-0014 and ADR-0003 rather than editing them silently, and fixes the tour/runbook/spec to describe the stack that actually runs today. Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
… one ready() constructed a new PrismaClient on every request and never disconnected it. Readiness is polled by every orchestrator, often every few seconds; measured against a real Postgres container, 35 polls of the old code left 35 idle connections that never closed, while the fixed code holds a single connection regardless of poll count. HealthController is a Nest singleton by default, so the client is now a field on the controller, lazily created once and reused — the same process lifetime a NestJS Prisma integration normally gives a connect-once service, without adding a full provider/module for one probe route.
lint_adapters only greped that ADAPTER_LIVENESS_PATH= and ADAPTER_READINESS_PATH= appeared, so an empty value passed. That is wider than it looks: tests/compose.bats extracts the path through a pipeline where cut exits 0 regardless, so an empty value collapses its HEALTHCHECK assertion into matching any localhost probe on 8080 — the exact Dockerfile-probing-nothing defect the assertion exists to catch. Both variables now have to hold a value starting with "/" when declared. Added tests/fixtures/lint/empty-readiness-path and a test proving the new check fails on it.
The missing-driver fixture proved a family with no driver file at all fails lint. Nothing proved the mirror case: a driver file that exists, sources cleanly, but omits one of REQUIRED_DRIVER_FUNCTIONS. Deleting the whole `for fn` loop left the suite passing, the same "gate that cannot fail" shape already named for the readiness-path check one loop above. Added tests/fixtures/lint-services/missing-driver-function (a complete laravel driver alongside a nest driver missing service_driver_compose_migrate) and a test asserting the specific "does not define" message. Verified by deleting the loop, confirming only this new test fails, then restoring it.
deploy-check.sh copied example.env to .env and never called generate_service_passwords, so the gate ran with DB_PASSWORD=changeme and APP_KEY=changeme literal — the password loop, the APP_KEY branch, and run_migrations could all break with every check in this branch staying green. The gate now sources common/install.sh and calls generate_service_passwords on the copied .env. Verified directly (a changeme value becomes a real random one) and end to end (nestjs+postgres and laravel-api+mysql both still pass the gate with generated credentials). ADR-0021 claimed common/install.sh "is the same sequence handed to a client", which was false — the gate builds and starts the stack directly and never downloads a release or runs install.sh end to end. Corrected to say what the gate actually shares with install.sh (generate_service_passwords) and what it does not.
… none A driven adapter (api/app) declares a readiness path regardless of --db, so `./scripts/deploy-check.sh laravel-api --db none` curled /health/ready expecting 200 against a route that has no driver spliced in and correctly returns 503. Confirmed both ways: the gate now passes laravel-api --db none, and reverting the guard reproduces the exact "returned 503, not 200" failure. The spec claimed --db none ships no readiness route at all; it does, and it answers honestly. Corrected to describe what actually happens and why the gate now skips the check for that combination instead of curling a route designed to fail.
…lation
tests/service.bats' password check had two holes: `[ -z "$block" ] &&
continue` skipped any driver emitting nothing, and the assertion checked
for the presence of `${DB_PASSWORD}`/`${REDIS_PASSWORD}` rather than the
absence of a literal, so a driver emitting both would have passed.
The mysql and postgres laravel drivers carried a `DB_PASSWORD:
${DB_PASSWORD}` line whose own comment said it existed so this test could
see it — the password already reaches the container through compose's
env_file. Removed both lines; the corrected test still passes on the
drivers' real merits, and a real laravel-api deploy against both mysql
and postgres still authenticates and runs its migrations with a
generated (non-changeme) password, so nothing depended on the restated
line.
- tests/compose.bats:177 — anchor the HEALTHCHECK grep to the start of the line, so it cannot match the word inside an unrelated comment. - docs/superpowers/specs/2026-09-06-deployable-stack-design.md:225 — nextjs's liveness path table entry still said `/`; it has been `/api/health/live` since Task 1. - adapters/nextjs/adapter.env:15 — "created by the next task" was plan-speak that shipped into a generated project's own file. Rewritten to name the route's actual path. - common/compose.yaml:17 — "replace it once, by hand" is wrong for a file install.sh re-downloads and overwrites on every run. Points at the runbook's actual flow (fix the repository's compose.yaml and cut a release) instead.
tests/service.bats claimed the removed empty-block skip had closed a hole — it had not: an empty block matches neither grep in the check, so the skip's removal changed no behaviour. docs/decisions/0021 named three things the deploy gate does not share with common/install.sh's own logic, and left out the one that matters most: the gate does not call install.sh's run_migrations at all, it carries a second implementation that has already drifted from it (install.sh falls back to grepping for a database service, the gate falls back to ROLE/DB_SERVICE).
Both while loops in the password check exist only for the eight real shipping drivers, and every one of those already writes an interpolation. Deleting the loops leaves the suite green — the same "gate that cannot fail" shape task 0a4976d fixed for the driver-function loop. Factors the check into _password_literal_report so both the real-driver sweep and a new fixture test call the identical logic, then adds a fixture driver that bakes in a literal password and a test asserting the check reports it. Mutation-proved: deleting both loops' bodies left the real-driver test green (an empty report still satisfies `[ -z "$bad" ]`) but failed the new fixture test; restoring the loops made it pass again.
^[A-Za-z_]*_PASSWORD: required the key at column zero, so an indented key ( DB_PASSWORD: hunter2) never reached the case check that would have flagged its literal value. The case pattern itself already tolerated a prefix via its leading *; only the grep that selects candidate lines needed the same tolerance. Left PGPASSWORD/DB_PASS/APP_KEY and friends alone, in a comment: a name blacklist is never complete, and the drivers are the only writers of this block and already go through review.
Measured: the block passes the literal-password check with the line and without it — env_file already delivers REDIS_PASSWORD to the container from the same .env compose interpolates from, so restating it here was dead weight. mysql and postgres's laravel drivers dropped the identical DB_PASSWORD line already; redis was the last holdout. Done after the literal-password fixture (tests/fixtures/lint-services/ literal-password) landed: this line was the only real driver exercising the key-form half of the check, so removing it first would have left that half checking nothing live.
fully_qualified_strict_types rejected the shipped \RuntimeException and \Throwable now that pint runs in CI. Fixing the source template alone was not enough: the mysql/postgres and mongodb drivers splice a fully qualified DB:: probe into the same file, which pint also rejects, and which then needs its own use import and a blank line before the following return (blank_line_before_statement). Verified against laravel-api's own generated project with pint's writing mode, then against tests/new-laravel-api.bats and tests/new-laravel-inertia.bats.
scaffold new commits the project it generates, so it needs a git identity. tests/helpers/setup.bash sets one up for bats, but this script runs outside bats and never got it — every deploy/deploy-tier-b job failed in seconds with "git has no user.name". Reproduced with GIT_CONFIG_GLOBAL unset and HOME pointed at an empty sandbox (no fallback identity anywhere), matching the runner; fixed the same way setup.bash does, and only when a global isn't already configured, so a developer with a real identity keeps theirs.
The git-identity fix revealed the next thing tests/helpers/setup.bash supplies that this script, running scaffold new outside bats, never picked up: a GitHub account for resolve_github_owner's guard, which a runner has no gh login or git config github.user to satisfy — every deploy/deploy-tier-b job died in seconds on "no GitHub account to substitute for 'you/'". Fixed the same way as the git identity, plus the trust store setup.bash also isolates for the same reason: mise trust records every config it trusts in real state keyed by path, and a repeated local run of this script would otherwise grow the developer's machine the same way bats' suites had, past 7600 stale entries. Reproduced by scrubbing only GH_CONFIG_DIR and SCAFFOLD_GITHUB_OWNER (not HOME, which moves mise's own data dir and breaks yq resolution instead of reproducing the bug) and confirmed the same error the CI log shows; fixed, confirmed nextjs passes clean, and confirmed nestjs --db postgres still passes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
A generated project's released stack — the
compose.yaml,example.envandinstall.shattached to every release — could never serve a request. Measured on 2026-09-05 against real published images:Laravel did not even report an error:
config/database.phpisenv('DB_CONNECTION', 'sqlite'), so with that variable absent it read theDB_DATABASE=appit did receive as a sqlite filename and never contacted the service. Three more, each independently fatal:EXPOSEsaid9000/3001/3000while compose published8080:8080and nothing rewrote it, so container port 8080 matched no adapter in this repository; both Laravel images ended atCMD ["php-fpm"]with no web server anywhere in the stack; andstorage/was root-owned while the server ran aswww-data.None of it was a regression. ADR-0014's seam 2 has required since 2026-08-27 that "every value a container needs arrives through
env_file/environmentat run time". Nothing was ever built to produce the values. It survived because no check could fail —tests/compose.batsrandocker compose config --quiet, which validates YAML, not reachability, and no test started a container.The branch closes that:
IMAGE_REPOSITORY: it "would only add a place for the default to silently drift from the real value".laravel.com/docs/13.x/deployment, part of the PHP Foundation since May 2025, and what Laravel Cloud runs. Alpine is not cosmetic: the mongodb driver emitsapk add.uptime.HEALTHCHECKand the gate read one value.nestjshad been probing a/healthno adapter shipped, so everynestjsimage reported unhealthy from first boot.install.sh— never from an entrypoint (ADR-0014 seam 5). The Nest image now carries the Prisma CLI so it can migrate itself.scripts/deploy-check.shplusdeploy/deploy-tier-b) that generates, builds, starts the stack, migrates, and requires both paths to return 200. Tier-gated per ADR-0012.Found by this pull request's own CI
The first real run of the
deploygate — the reason for opening this rather than merging locally — turned up two failures no local check could have caught, and one thing I got wrong in this description.deployjobs failed in 6-12 seconds:git has no user.name to commit the new project with.scaffold newcommits what it generates. Thesmokejobs survive because they run through bats, andtests/helpers/setup.bashgives them aGIT_CONFIG_GLOBAL;deploy-check.shcalls the script directly and gets nothing. Invisible locally, because a developer machine has an identity.smoke (laravel-api)andsmoke-tier-bfailed on Pint:routes/health.phpviolatesfully_qualified_strict_typesandblank_line_before_statement. Both adapters ship a byte-identical copy, so both fail the same way.Fixes are in flight on this branch.
How it was verified
mise run lintclean.mise run test-runner199/0 at344b29b.That number is narrower than it sounds, and I overstated it when I opened this pull request.
test-runnerdoes not runtests/new-*.bats— the task's own description says so, and those per-adapter suites are run only by CI'ssmokejobs. A greentest-runnertherefore says nothing about any adapter's ownci-unit, which is exactly how the Pint failure above reached CI instead of my laptop.scripts/deploy-check.shgreen against real containers forlaravel-api(default mysql),nestjs --db postgres,nextjs, andlaravel-inertia— the tier-B adapter whose PHP+Node assets stage had never been started.Green-immediately proven from a
git clone, not a working tree: all four config roots'ci-unitpass in a clone. A working tree keeps artifacts that make checks pass for the wrong reason — that is howlaravel-inertiashipped for weeks unable to pass its own CI.Three defects were found the only way they could be, by starting a container: FrankenPHP's
CMDdropping the base image's args so nothing listened;nextjsbinding toprocess.env.HOSTNAME, which Docker sets to the container id; and the Nest runtime image having neitherprismanorschema.prisma.The connection leak in the readiness route was measured, not argued: unfixed, 35 polls left 35 idle Postgres connections (1→36); fixed, it holds at 1.
Every new check was mutation-proved — deleted, watched fail, restored.
Known and deliberately not fixed
install.sh'srun_migrations; it carries a second implementation, and the two have drifted in their fallback condition. Deferred on purpose: a real-project run is the first time that code will ever execute, and choosing which copy is right deserves evidence rather than reasoning. ADR-0021 states the exclusion.docs/decisions/0021-the-released-stack-must-run.mdrecords the rest.Checklist
mise run lintpassesmise run test-runnerpasses — 199/0 at344b29b; see the note above on what it does not coverhttps://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv