Skip to content

feat: make a generated project's released stack actually run - #7

Merged
ttncode merged 30 commits into
mainfrom
feat/deployable-stack
Sep 7, 2026
Merged

ttncode merged 30 commits into
mainfrom
feat/deployable-stack

Conversation

@ttncode

@ttncode ttncode commented Sep 7, 2026 •

Copy link
Copy Markdown
Owner

What this changes

A generated project's released stack — the compose.yaml, example.env and install.sh attached to every release — could never serve a request. Measured on 2026-09-05 against real published images:

$ docker compose exec app node -e "... new PrismaClient(); p.$connect() ..."
error: Environment variable not found: DATABASE_URL.

Laravel did not even report an error: config/database.php is env('DB_CONNECTION', 'sqlite'), so with that variable absent it read the DB_DATABASE=app it did receive as a sqlite filename and never contacted the service. Three more, each independently fatal: EXPOSE said 9000/3001/3000 while compose published 8080:8080 and nothing rewrote it, so container port 8080 matched no adapter in this repository; both Laravel images ended at CMD ["php-fpm"] with no web server anywhere in the stack; and storage/ was root-owned while the server ran as www-data.

None of it was a regression. ADR-0014's seam 2 has required since 2026-08-27 that "every value a container needs arrives through env_file/environment at run time". Nothing was ever built to produce the values. It survived because no check could fail — tests/compose.bats ran docker compose config --quiet, which validates YAML, not reachability, and no test started a container.

The branch closes that:

  • A container port contract. Every adapter serves on 8080. The alternative — a variable written at generation time — is rejected for the reason this project already recorded when it rejected a parameterised IMAGE_REPOSITORY: it "would only add a place for the default to silently drift from the real value".
  • An HTTP listener in the Laravel images. FrankenPHP, alpine, pinned by digest. Documented on laravel.com/docs/13.x/deployment, part of the PHP Foundation since May 2025, and what Laravel Cloud runs. Alpine is not cosmetic: the mongodb driver emits apk add.
  • An environment contract produced by the driver, the only thing that knows both the service and the adapter family. The password exists in exactly one place; compose interpolates it at up time.
  • Liveness and readiness paths declared by the adapter, so the HEALTHCHECK and the gate read one value. nestjs had been probing a /health no adapter shipped, so every nestjs image reported unhealthy from first boot.
  • Migrations from a profiled compose service, run by install.sh — never from an entrypoint (ADR-0014 seam 5). The Nest image now carries the Prisma CLI so it can migrate itself.
  • A deploy gate (scripts/deploy-check.sh plus deploy/deploy-tier-b) that generates, builds, starts the stack, migrates, and requires both paths to return 200. Tier-gated per ADR-0012.

Found by this pull request's own CI

The first real run of the deploy gate — the reason for opening this rather than merging locally — turned up two failures no local check could have caught, and one thing I got wrong in this description.

  1. All four deploy jobs failed in 6-12 seconds: git has no user.name to commit the new project with. scaffold new commits what it generates. The smoke jobs survive because they run through bats, and tests/helpers/setup.bash gives them a GIT_CONFIG_GLOBAL; deploy-check.sh calls the script directly and gets nothing. Invisible locally, because a developer machine has an identity.
  2. smoke (laravel-api) and smoke-tier-b failed on Pint: routes/health.php violates fully_qualified_strict_types and blank_line_before_statement. Both adapters ship a byte-identical copy, so both fail the same way.
  3. My verification claim was wider than my evidence — see the note below.

Fixes are in flight on this branch.

How it was verified

  • mise run lint clean. mise run test-runner 199/0 at 344b29b.

    That number is narrower than it sounds, and I overstated it when I opened this pull request. test-runner does not run tests/new-*.bats — the task's own description says so, and those per-adapter suites are run only by CI's smoke jobs. A green test-runner therefore says nothing about any adapter's own ci-unit, which is exactly how the Pint failure above reached CI instead of my laptop.

  • scripts/deploy-check.sh green against real containers for laravel-api (default mysql), nestjs --db postgres, nextjs, and laravel-inertia — the tier-B adapter whose PHP+Node assets stage had never been started.

  • Green-immediately proven from a git clone, not a working tree: all four config roots' ci-unit pass in a clone. A working tree keeps artifacts that make checks pass for the wrong reason — that is how laravel-inertia shipped for weeks unable to pass its own CI.

  • Three defects were found the only way they could be, by starting a container: FrankenPHP's CMD dropping the base image's args so nothing listened; nextjs binding to process.env.HOSTNAME, which Docker sets to the container id; and the Nest runtime image having neither prisma nor schema.prisma.

  • The connection leak in the readiness route was measured, not argued: unfixed, 35 polls left 35 idle Postgres connections (1→36); fixed, it holds at 1.

  • Every new check was mutation-proved — deleted, watched fail, restored.

Known and deliberately not fixed

  • The gate does not run install.sh's run_migrations; it carries a second implementation, and the two have drifted in their fallback condition. Deferred on purpose: a real-project run is the first time that code will ever execute, and choosing which copy is right deserves evidence rather than reasoning. ADR-0021 states the exclusion.
  • No CI lane starts a mongodb container, and the gate only exercises each adapter's default database.
  • docs/decisions/0021-the-released-stack-must-run.md records the rest.

Checklist

  • mise run lint passes
  • mise run test-runner passes — 199/0 at 344b29b; see the note above on what it does not cover
  • New behaviour has a test that fails without the change
  • Docs that describe changed behaviour were updated in the same commit
  • No unrelated changes

https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv

An acceptance run started a generated project's released stack for the first
time. It does not work, and never has, in any shape: no application is told
how to reach its database, container port 8080 matches no adapter in this
repository, the Laravel images end at php-fpm with no web server anywhere in
the stack, and nestjs has been probing a /health route nothing serves.

None of it is a regression. ADR-0014 seam 2 already required that every value
a container needs arrives through env_file/environment; nothing was ever made
to produce the values. It survived because no check could fail —
tests/compose.bats runs `docker compose config --quiet`, which validates YAML,
not reachability, and no test starts a container.

The design turns "green immediately, deployable immediately" into two
mechanically checkable gates, and settles the three decisions that follow:
every adapter serves on container port 8080; the Laravel images move to
FrankenPHP alpine pinned by digest; and the app's connection variables are
composed in compose.yaml from the service variables already in .env, by the
driver, which is the only thing that knows both the service and the adapter
family.

It also moves ADR-0003's boundary. An adapter has never written application
code; it now writes one readiness route, because a liveness path that does not
touch the database is the check-that-cannot-fail ADR-0014 seam 4 warned about,
and it is what let all of this ship.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Ten tasks, each ending in something independently testable, and each
container check using `docker compose up -d` rather than `docker run`, which
this environment denies.

Writing it surfaced two things the spec had wrong, both fixed in the spec as
well rather than papered over in the plan:

The published nestjs image cannot migrate itself. The driver installs the CLI
with `pnpm add -D prisma@6` and the Dockerfile runs `pnpm prune --prod` before
the runtime stage copies node_modules, so the image carries @prisma/client and
no prisma binary. The driver now installs it as a regular dependency, which is
what Prisma's own guidance assumes when the migration runs from the image, and
the image-size cost is recorded rather than discovered later.

And install.sh cannot run a mise task, because no image carries mise. The
driver writes a profiled `migrate` compose service instead, sharing the app's
image and environment, so it never starts with the stack — ADR-0014 seam 5
forbids migrating from an entrypoint, not from the operator's one command.

Also removes a dangling ADR-0021 reference from both documents. The decision
record is Task 8's output, and tests/documentation.bats was right to fail on a
spec citing a file that does not exist yet.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
ADAPTER_LIVENESS_PATH (every adapter) and ADAPTER_READINESS_PATH (adapters
whose role is in DRIVEN_ROLES) are now declared once per adapter and read by
lib/lint.sh, so the Dockerfile and the CI gate stop each keeping their own
idea of the route — the defect that let nestjs probe a /health nothing
served.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
The case branch in lint_adapters that requires ADAPTER_READINESS_PATH only
for DRIVEN_ROLES had no fixture exercising it: every existing lint fixture
is role api, and the tests that did touch readiness paths checked the real
adapter.env files, not the enforcement code, so deleting the branch would
have gone unnoticed. Adds a role-api fixture missing the var (must be
reported) and a role-web fixture missing it (must stay silent), each with
its own test.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
nestjs and nextjs both listened on their generator's default port while
common/compose.yaml has always published 8080 — nestjs additionally probed a
/health route no adapter ever shipped, so its image reported unhealthy from
first boot. Both adapters now listen on 8080, ship a real liveness route, and
their HEALTHCHECKs probe it there. apply_adapter's docker/-only special case
is now a loop over every directory an adapter ships, since nestjs's health
routes and nextjs's health route both need the same directory-delivery
mechanism the Laravel pair already used for docker/opcache.ini.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
sed -i "1i ..." always succeeds regardless of match, and the following
substitution for imports: [] silently no-ops if the nest generator ever
reformats app.module.ts — leaving HealthModule unregistered with no build
or lint failure, only a 404 on /health/live discovered by the HEALTHCHECK
this task just pointed at that route. Grep for both halves of the wiring
after the sed calls and fail loudly, naming the file and the likely cause,
so a generator reformat cannot silently re-arm the unhealthy-image defect
this task exists to remove.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Replace the php-fpm runtime stage in both laravel adapters with FrankenPHP
(alpine, pinned by digest), so the images actually speak HTTP instead of
FastCGI with nothing in front of them. Adds SERVER_NAME, the XDG_* pair
Caddy needs to start as www-data, ownership fixes for storage/ and
bootstrap/cache, and a real HEALTHCHECK against /up.

The base image's default CMD carries --config/--adapter flags that point
frankenphp at its Caddyfile; a bare `frankenphp run` drops them and starts
only the admin API, so CMD restates them explicitly.

laravel-inertia also moves COPY . . before the assets copy so a developer's
local public/build can no longer shadow the one the assets stage built, and
.dockerignore excludes public/build from the build context for the same
reason.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
compose.yaml only ever received DB_DATABASE/DB_USERNAME/DB_PASSWORD, and
neither Prisma nor Laravel reads those directly, so a released stack could
not reach its database. apply_service_compose_env merges a driver's
service_driver_compose_env block into compose.yaml's app service with
`yq -P` so the merge keeps block style and the file's comments survive.

Every driver is required to define the hook (REQUIRED_DRIVER_FUNCTIONS,
checked by lint_services), so a driver shipping only two of the three fails
at generation or at lint instead of silently producing an app with no
DATABASE_URL. No driver implements the hook yet — that is task 5.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
lint_services sourced each driver with no SERVICE_DIR set, unlike every
other call site that sources a driver — a Task-5 driver reading it at
sourcing time would die under the inherited `set -u` and get misreported as
"does not define" the function it never got a chance to declare. Set it the
way load_service does, and capture the subshell's stderr so a sourcing
failure reports its own text instead of being folded into the
missing-function message.

Dropped compose-env's `-P`: measured, not assumed. The fragment it merges is
always this function's own block-style output, never the flow style
merge_lefthook_fragment's `-P` guards against, so there is nothing here for
it to protect. Merging without it left compose.yaml's untouched lines
byte-identical; with it, yq rewrote nodes this merge never touched — an
unquoted `${APP_PORT:-8080}:8080` and a healthcheck array reshaped from flow
to block. Both happened to pass prettier --check, but a merge with no
business editing those lines should not be reshaping them anyway.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
All eight drivers now implement service_driver_compose_env: nest gets
DATABASE_URL/REDIS_URL built from the same DB_*/REDIS_* variables the
database/cache containers read, operator-overridable via .env; laravel gets
DB_CONNECTION plus whatever host/port/DSN that family needs, since
config/database.php defaults to sqlite and its absence was a silent wrong
answer, not a failure. Every interpolation reads the password rather than
embedding it, so it exists in exactly one place.

Both frameworks also get a real readiness probe, spliced into the anchor
each ships pre-generation. Two substitutions, not one, in both: the anchor
becomes the probe, and the shipped fallback (`throw`) is replaced in place
rather than left as dead code below an added return — nest's controller
would otherwise fail eslint's no-unreachable under --max-warnings 0. Nest's
probe is cast to an explicit method signature instead of a bare dynamic
import, because before `prisma generate` runs (lint runs first) the
generated client module doesn't exist, and an untyped access to it is `any`
under the no-unsafe-* rules; only the one method the provider actually calls
is declared, since a real PrismaClient's mongodb build has no
$queryRawUnsafe and its SQL builds have no $runCommandRaw.

Laravel needed a route to splice into. Both laravel adapters now ship
routes/health.php, wired into bootstrap/app.php's withRouting(then: ...) —
laravel only auto-loads routes/web.php and routes/console.php, so a file
dropped there with nothing pointing at it 404s forever. APP_KEY, per-family
rather than per-service, is appended to the project's example.env by every
laravel driver, since no service env.fragment can carry it and laravel will
not boot without one.

Verified live: generating a laravel+mongodb project and hitting
/health/ready surfaced a real bug independent of this task — `pecl install
mongodb` resolves latest (2.x) while mongodb/laravel-mongodb is built
against the 1.x extension's BSON model classes, so any code touching
MongoDB\Client fatals before reaching this project's own try/catch. Pinned
to 1.21.0, matching the platform.ext-mongodb declaration already beside it.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
install.sh now generates a laravel-valid APP_KEY (base64: + 32 bytes,
handled in generate_service_passwords' own loop, with the sed delimiter
changed to | since a base64 value can contain /) and runs a profiled
migrate compose service after the stack is up, before the success message.

Every service driver now implements service_driver_compose_migrate
(added to REQUIRED_DRIVER_FUNCTIONS), and a new apply_service_compose_service
merges a driver's complete services: fragment — apply_service_compose_env
cannot do this itself, since it hardcodes services.app.environment. The
migrate service shares the app's image (read back off compose.yaml) and its
environment. install.sh's guard for whether a migrate service exists uses
`docker compose --profile migrate config --services`, not the profile-less
form, which never lists a profiled service at all.

The nest driver now installs prisma as a regular dependency, not -D, since
the Dockerfile's `pnpm prune --prod` would otherwise strip the CLI the
migrate service needs out of the published image.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
The compose migrate command written by service_driver_compose_migrate used
`pnpm exec`, but the runtime image (adapters/nestjs/Dockerfile[.workspace])
never installs pnpm — only the build stage runs corepack enable, and the
runtime stage copies node_modules/dist alone. Measured against a real built
image: `which pnpm` exits 1.

prisma's own bin does survive `pnpm prune --prod` (it is a regular
dependency for exactly this), but at a location that depends on which
Dockerfile shape wins, a decision made after this driver runs:
apps/<app>/node_modules/.bin for the typescript-workspace shape,
node_modules/.bin at the container root for the standalone one. The new
command tries both, and cd's into whichever matched before running it, since
prisma resolves its schema relative to its own cwd and that schema lives
nested under the app directory in the workspace shape.

That surfaced a second defect once the binary was found: `schema.prisma`
was never copied into the runtime image at all, in either Dockerfile shape,
since it is not an artifact `nest build` produces. Both Dockerfiles now
copy prisma/ into the runtime stage via a bracket glob (`pris[m]a`), which
copies nothing rather than failing the build on a --db none project that
has no prisma/ directory to copy.

Also fixed: `$` in the compose-yaml command needs to be `$$` to survive
compose's own variable interpolation before the container ever sees it —
a single `$` here silently resolved to an unset variable and blanked the
whole loop, caught by inspecting `docker compose config`'s resolved output
before trusting it.

Verified on a real nestjs+postgres stack: readiness 200 before and after
migrating (the SQL probe is connectivity-only, matching Task 5's own
documented finding — left alone per this round's ruling), and
`docker compose --profile migrate run --rm migrate` exits 0 and reaches the
real database.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Everything before this validated compose.yaml as YAML; nothing ever started
a container, which is how a stack that could never serve a request shipped
for weeks. scripts/deploy-check.sh generates an adapter, builds its released
image, brings the compose stack up, waits for the app container to report
healthy, asserts the migrate profile service exits 0 (the readiness probe
alone proves connectivity, not schema), then asserts both the liveness and
declared readiness path return 200 — tearing the stack down on every one of
those paths via a trap, not a trailing cleanup line. Both paths are read out
of the adapter's own adapter.env, never hardcoded, and readiness is skipped
only where the adapter declares none, with that skip stated in the output.

adapters.yml gains deploy/deploy-tier-b jobs mirroring smoke/smoke-tier-b's
discover-driven matrices, running this script per tier-a and tier-b adapter.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Task 7's deploy gate found this on its first real run: Docker sets HOSTNAME
to the container's own id for every container, and standalone server.js
binds to `process.env.HOSTNAME || '0.0.0.0'` — so without an explicit
override, the app listened on that id-derived address, reachable through
compose's published port (routed to the container's real interface either
way) but not from the HEALTHCHECK, which dials localhost from inside the
same container and got "connection refused" forever.

ENV HOSTNAME="0.0.0.0" fixes the bind, but exposed a second, previously
masked mismatch: that bind is IPv4-only, while this image's resolver hands
wget the ::1 (IPv6) address for "localhost" first, and busybox wget does not
fall back to the IPv4 result. Measured with `docker run` at each step.
Pointing the HEALTHCHECK at 127.0.0.1 instead closes that gap.

PORT was already correct — `ss -tlnp` showed the server bound on :8080
throughout; the defect was only ever the bind address, not the port.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Review round 2. Critical: the migrate check gated on a condition derived
from the artifact under test — a renamed service, a broken profile, or a
driver returning no command all read as "no database", the readiness probe
(select 1) then returns 200 against unapplied schema, and the run exits 0.
ROLE and DB_SERVICE are already known, so absence is now a failure whenever
a database was actually requested. common/install.sh's run_migrations had
the identical hole (a database service with no migrate beside it now fails
loudly instead of returning 0), since it ships to every generated project.

Also, in order of how they'd bite:
- COMPOSE_PROJECT_NAME is set per adapter — every generated stack is
  `name: app` (common/compose.yaml), so a local run used to reconcile
  against, and `down -v` any real "app" project already on the machine.
- ADAPTER_LIVENESS_PATH and ADAPTER_READINESS_PATH are asserted non-empty:
  lib/lint.sh only checks the line is present, and an empty liveness path
  silently probes "/", which nextjs happens to answer 200 for reasons
  unrelated to its real liveness.
- compose.yaml's image is asserted equal to the tag just built, not just
  "no CHANGEME left" — a real registry reference in common/compose.yaml
  would match neither check and the stack would come up on a pulled image.
- notify-on-schedule-failure now watches deploy and deploy-tier-b too.
- APP_PORT's read no longer dies silently under pipefail when a .env lacks
  it; the health poll uses `docker inspect` instead of `docker compose ps
  --format json`, whose shape is compose-version-dependent; the cleanup
  trap now covers INT/TERM, not just EXIT.

tests/compose.bats' liveness-probe test is updated to accept 127.0.0.1
alongside localhost, matching fix round 1's nextjs HEALTHCHECK change.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
Task 8 of the deployable-stack plan: ADR-0014 seams 1 and 4, ADR-0003's
no-application-code boundary, the tour's healthcheck example, the
runbook's pull-only step 10, and the design spec's overclaimed readiness
proof were all falsified by the work in commits c61600f..0b59f6a. Records
the corrected boundary in a new ADR-0021, marks the superseded parts of
ADR-0014 and ADR-0003 rather than editing them silently, and fixes the
tour/runbook/spec to describe the stack that actually runs today.

Claude-Session: https://claude.ai/code/session_01J4HB8qJjdZwaAv42k6HMpv
… one

ready() constructed a new PrismaClient on every request and never
disconnected it. Readiness is polled by every orchestrator, often every
few seconds; measured against a real Postgres container, 35 polls of the
old code left 35 idle connections that never closed, while the fixed code
holds a single connection regardless of poll count.

HealthController is a Nest singleton by default, so the client is now a
field on the controller, lazily created once and reused — the same
process lifetime a NestJS Prisma integration normally gives a
connect-once service, without adding a full provider/module for one
probe route.
lint_adapters only greped that ADAPTER_LIVENESS_PATH= and
ADAPTER_READINESS_PATH= appeared, so an empty value passed. That is wider
than it looks: tests/compose.bats extracts the path through a pipeline
where cut exits 0 regardless, so an empty value collapses its HEALTHCHECK
assertion into matching any localhost probe on 8080 — the exact
Dockerfile-probing-nothing defect the assertion exists to catch.

Both variables now have to hold a value starting with "/" when declared.
Added tests/fixtures/lint/empty-readiness-path and a test proving the new
check fails on it.
The missing-driver fixture proved a family with no driver file at all
fails lint. Nothing proved the mirror case: a driver file that exists,
sources cleanly, but omits one of REQUIRED_DRIVER_FUNCTIONS. Deleting the
whole `for fn` loop left the suite passing, the same "gate that cannot
fail" shape already named for the readiness-path check one loop above.

Added tests/fixtures/lint-services/missing-driver-function (a complete
laravel driver alongside a nest driver missing
service_driver_compose_migrate) and a test asserting the specific
"does not define" message. Verified by deleting the loop, confirming
only this new test fails, then restoring it.
deploy-check.sh copied example.env to .env and never called
generate_service_passwords, so the gate ran with DB_PASSWORD=changeme and
APP_KEY=changeme literal — the password loop, the APP_KEY branch, and
run_migrations could all break with every check in this branch staying
green.

The gate now sources common/install.sh and calls
generate_service_passwords on the copied .env. Verified directly (a
changeme value becomes a real random one) and end to end (nestjs+postgres
and laravel-api+mysql both still pass the gate with generated
credentials).

ADR-0021 claimed common/install.sh "is the same sequence handed to a
client", which was false — the gate builds and starts the stack directly
and never downloads a release or runs install.sh end to end. Corrected to
say what the gate actually shares with install.sh (generate_service_passwords)
and what it does not.
… none

A driven adapter (api/app) declares a readiness path regardless of --db,
so `./scripts/deploy-check.sh laravel-api --db none` curled
/health/ready expecting 200 against a route that has no driver spliced
in and correctly returns 503. Confirmed both ways: the gate now passes
laravel-api --db none, and reverting the guard reproduces the exact
"returned 503, not 200" failure.

The spec claimed --db none ships no readiness route at all; it does, and
it answers honestly. Corrected to describe what actually happens and why
the gate now skips the check for that combination instead of curling a
route designed to fail.
…lation

tests/service.bats' password check had two holes: `[ -z "$block" ] &&
continue` skipped any driver emitting nothing, and the assertion checked
for the presence of `${DB_PASSWORD}`/`${REDIS_PASSWORD}` rather than the
absence of a literal, so a driver emitting both would have passed.

The mysql and postgres laravel drivers carried a `DB_PASSWORD:
${DB_PASSWORD}` line whose own comment said it existed so this test could
see it — the password already reaches the container through compose's
env_file. Removed both lines; the corrected test still passes on the
drivers' real merits, and a real laravel-api deploy against both mysql
and postgres still authenticates and runs its migrations with a
generated (non-changeme) password, so nothing depended on the restated
line.
- tests/compose.bats:177 — anchor the HEALTHCHECK grep to the start of
  the line, so it cannot match the word inside an unrelated comment.
- docs/superpowers/specs/2026-09-06-deployable-stack-design.md:225 —
  nextjs's liveness path table entry still said `/`; it has been
  `/api/health/live` since Task 1.
- adapters/nextjs/adapter.env:15 — "created by the next task" was
  plan-speak that shipped into a generated project's own file. Rewritten
  to name the route's actual path.
- common/compose.yaml:17 — "replace it once, by hand" is wrong for a
  file install.sh re-downloads and overwrites on every run. Points at
  the runbook's actual flow (fix the repository's compose.yaml and cut a
  release) instead.
tests/service.bats claimed the removed empty-block skip had closed a
hole — it had not: an empty block matches neither grep in the check, so
the skip's removal changed no behaviour.

docs/decisions/0021 named three things the deploy gate does not share
with common/install.sh's own logic, and left out the one that matters
most: the gate does not call install.sh's run_migrations at all, it
carries a second implementation that has already drifted from it
(install.sh falls back to grepping for a database service, the gate
falls back to ROLE/DB_SERVICE).
Both while loops in the password check exist only for the eight real
shipping drivers, and every one of those already writes an interpolation.
Deleting the loops leaves the suite green — the same "gate that cannot
fail" shape task 0a4976d fixed for the driver-function loop.

Factors the check into _password_literal_report so both the real-driver
sweep and a new fixture test call the identical logic, then adds a
fixture driver that bakes in a literal password and a test asserting the
check reports it.

Mutation-proved: deleting both loops' bodies left the real-driver test
green (an empty report still satisfies `[ -z "$bad" ]`) but failed the
new fixture test; restoring the loops made it pass again.
^[A-Za-z_]*_PASSWORD: required the key at column zero, so an indented
key (  DB_PASSWORD: hunter2) never reached the case check that would
have flagged its literal value. The case pattern itself already
tolerated a prefix via its leading *; only the grep that selects
candidate lines needed the same tolerance.

Left PGPASSWORD/DB_PASS/APP_KEY and friends alone, in a comment: a name
blacklist is never complete, and the drivers are the only writers of
this block and already go through review.
Measured: the block passes the literal-password check with the line and
without it — env_file already delivers REDIS_PASSWORD to the container
from the same .env compose interpolates from, so restating it here was
dead weight. mysql and postgres's laravel drivers dropped the identical
DB_PASSWORD line already; redis was the last holdout.

Done after the literal-password fixture (tests/fixtures/lint-services/
literal-password) landed: this line was the only real driver exercising
the key-form half of the check, so removing it first would have left
that half checking nothing live.
fully_qualified_strict_types rejected the shipped \RuntimeException and
\Throwable now that pint runs in CI. Fixing the source template alone
was not enough: the mysql/postgres and mongodb drivers splice a fully
qualified DB:: probe into the same file, which pint also rejects, and
which then needs its own use import and a blank line before the
following return (blank_line_before_statement). Verified against
laravel-api's own generated project with pint's writing mode, then
against tests/new-laravel-api.bats and tests/new-laravel-inertia.bats.
scaffold new commits the project it generates, so it needs a git
identity. tests/helpers/setup.bash sets one up for bats, but this
script runs outside bats and never got it — every deploy/deploy-tier-b
job failed in seconds with "git has no user.name". Reproduced with
GIT_CONFIG_GLOBAL unset and HOME pointed at an empty sandbox (no
fallback identity anywhere), matching the runner; fixed the same way
setup.bash does, and only when a global isn't already configured, so a
developer with a real identity keeps theirs.
The git-identity fix revealed the next thing tests/helpers/setup.bash
supplies that this script, running scaffold new outside bats, never
picked up: a GitHub account for resolve_github_owner's guard, which a
runner has no gh login or git config github.user to satisfy — every
deploy/deploy-tier-b job died in seconds on "no GitHub account to
substitute for 'you/'". Fixed the same way as the git identity, plus
the trust store setup.bash also isolates for the same reason: mise
trust records every config it trusts in real state keyed by path, and
a repeated local run of this script would otherwise grow the
developer's machine the same way bats' suites had, past 7600 stale
entries.

Reproduced by scrubbing only GH_CONFIG_DIR and SCAFFOLD_GITHUB_OWNER
(not HOME, which moves mise's own data dir and breaks yq resolution
instead of reproducing the bug) and confirmed the same error the CI
log shows; fixed, confirmed nextjs passes clean, and confirmed
nestjs --db postgres still passes.
@ttncode
ttncode merged commit 17bcb36 into main Sep 7, 2026
18 checks passed
@ttncode
ttncode deleted the feat/deployable-stack branch September 7, 2026 08:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants