Skip to content

048: the spec interview now keeps a live draft and a plain-words check before anything is written - #712

Merged
soydachi merged 56 commits into
mainfrom
ship/048-handshake-intake
Aug 29, 2026
Merged

048: the spec interview now keeps a live draft and a plain-words check before anything is written#712
soydachi merged 56 commits into
mainfrom
ship/048-handshake-intake

Conversation

@soydachi

Copy link
Copy Markdown
Member

What changed, in plain words: before the framework writes a specification, it now interviews you the way a careful person would — it asks one question at a time, keeps a running draft on disk so a crashed session can pick up where it stopped instead of starting over, and before writing anything down it repeats your idea back in two plain sentences and waits for you to say "yes, that's it". If nobody is at the keyboard (an unattended run), it no longer hangs waiting for that yes; it writes down that the check went unconfirmed and carries on, so the record always says what was and wasn't human-verified. Nothing new was added to the command list: this all lives inside the spec-writing step that already exists.

Spec: specs/048-handshake-intake-mechanisms/spec.md (approved at digests by docs/adr/0036 and docs/adr/0037; the critics' rounds are folded in its ## Grill and ## Council sections).

Look at first: .agents/skills/ai-spec/references/intake.md — the new interview rules — and the rewritten step 0 in .agents/skills/ai-spec/SKILL.md, whose exact bytes are pinned twice in tests/test_contracts.py (the verbatim step text and the whole-file digest).

What I am not confident about: the two stored pins make step 0 cheap to change in this repository and expensive to change wrongly — the digest test is the only early warning; and the read-back confirmation is still enforced by prose, not by code (spec 048's risks name what mechanizing it must not repeat: 037's validator was built and deleted as an orphan). The /ai-verify and /ai-review verdicts and the security lane's named residuals (tick-digest binding, plan-id collision, history-scan gap on old tags) are recorded in the spec's sections, not hidden.

Gate: just check exit 0; ai-eng audit verify WARN (22 accounted test-buffer links); just security exit 0 at pinned versions; all four plan tasks ticked by sealed commands over the 0037 digest.

…auditors, the orphan-delete pattern as the central finding

The claim verified against the tree, not the ticks: twelve modules from specs
029-037 were written, checked, ticked [x], never wired, and deleted by 044 the
same day 042 built the register whose plan forbade exactly that deletion
(plan.md:38 quoted). 035 never existed; 023-T6/7 and the B4 batched gate are
promised prose; ~30 research P0s have neither adoption nor recorded refusal.
045/046 survive independent refutation intact.
…ule — spec 044 cut twelve modules and left six sentences pointing at the corpses

trim, skillify, versions.verify_against_installed and loopgate are gone from
src/, but ai-note, ai-review and ai-security corpora still told the agent to
call them, ai-challenge and ai-council still named loopgate as the
orchestrator's instrument, and test_skill_bounds pinned the word loopgate so
the dead reference could not be repaired without going red. The quoted
situations are untouched — skill_eval counts only those, so the routing
baseline stays 428 with delta zero. The cap now names the rule itself: two
digest-equal green runs.
…ork file — and the test stops accepting the Spanish it translated

The reference carried the four basics in Spanish while every skill, spec and
corpus speaks English, and test_reference_names_the_rule accepted either
('keyboard' or 'teclado'), which is how the drift hid. The reference is
rewritten in English, the disjunction loses its second arm, and the seven
targeted tests stay green.
…ad, five dead with a named reopening condition

The owner asked whether the deleted modules were necessary at all. Per module,
in a note (append-only, with still_true_when so the verdict rots visibly if a
restoration happens): the rule survives as prose where the module was
scaffolding (trim, versions, skillify, constellation, decision_fw, evidencing,
decision_boundary), and where the instrument is genuinely wanted (loopgate,
lane_merge, verify_cold+answer_key, intake+bail_out) it returns only together
with its consumer in the same commit — the orphan shape is what this refuses.
…il contract.py carries the number

The owner set 180 minutes for a giant block (normal ones far under) against
the measured 19 h 42 min of the only cycle ever timed. The row is INCOMPLETE
and says which half: the ceiling exists as a decision and a postmortem, not as
a constant a recap can fail on. Its evidence command greps contract.py for
CYCLE_WALL_BUDGET_MINUTES, so the row closes only when spec 047 lands the
number beside the other named budgets. COMMITMENTS moves 26 to 27 in the same
commit — the count and the row are one change.
…ers as an outcome

current_facts carried two claims the tree had already contradicted — four
inherited test_madr failures (ADR 0025 closed them; 37 passed today) and the
22-link audit red (audit verify exits 0). The owner's ceiling becomes an
intended outcome: a governed cycle fits 180 minutes for a giant block and
less for normal ones, held open by PO-27 until contract.py carries the
number. The page is re-forged over the same bytes.
… derived report

The postmortem's five priced cuts, turned into eight atomic tasks: three named
budgets in contract.py (PO-27's evidence command answers at task 1), a red
fixture first, a stdlib vitals reader over the event stream the hooks already
stamp, the report verb, the five critics carrying their box and a TIMEBOXED
exit, the batched check-all with the verified 1.58 prefix, the goal's
anti-stall rule, and a close that lets the clock disqualify and never
approve. Draft until the owner signs the digests.
… three real defects and the spec answers all three

Grill caught the context naming `stamp` for wall arithmetic when `stamp` is the
chain-seal hash and `ts` is the time (task 3 rewritten around `ts`, the
attribution rule named), caught the 120-calls figure cited to a measurement the
postmortem does not carry (now stated as derived: 409/5, rounded), and caught a
stale line count. The council's CLAIMED-NOT-REAL sweep over the eight tasks was
refuted by the cross-read — a draft with every box `[ ]` is not a lie — and its
one surviving gap candidate refuted itself. The council reader passes: RAN
council=64/88 with 047's counts stated and reproduced.
…xecutes must exit zero, and a red fixture cannot satisfy a tick by being red

The plan's own mechanics caught what the critics did not: --tick refuses a
non-zero check, so the red-fixture task needed a check that proves the shape
(collect-only, imports inside test bodies) while the run itself stays red.
The house precedent is 046 task 3, whose check names the three refusals
passing — red-then-green ticks on the green half.
…, execution directed

The owner approved in conversation on 2026-08-29 at spec 1ec7e171… and plan
e4252e37… (canonical, tick column masked). The record names the nine defects
the critics and the tick mechanics found before signing — stamp-is-not-time,
the invented 120-call measurement, the host-transcript buckets, the red
fixture that could never seal — and what the approval knowingly costs:
derived boxes, gaps-not-phases for silent forks, and check-all degrading to
check if just ever drops the prefix.
CYCLE_WALL_BUDGET_MINUTES 180 (the owner's decision, PO-27's row carries it),
CRITIC_TIMEBOX_MINUTES 40 and CRITIC_CALLS_MAX 120 (derived, and the comment
says so — 409 measured calls over five critics, rounded down). The only place
the numbers live; the clock disqualifies and never approves.
…failing assertions, imports inside test bodies

--collect-only exits zero (six tests gathered) while the run is red naming
every missing half: the five critics carry no box and vitals does not exist.
The plan's check is the collection precisely because --tick seals a zero
exit: a red fixture cannot satisfy a tick by being red, so the check proves
the shape and the done-when says which half stays red on purpose.
…o the earlier event's cls

stdlib-only, reads the text of the event record and a session id; verdict
PASS inside CYCLE_WALL_BUDGET_MINUTES, INCOMPLETE [OVER_BUDGET] naming the
largest bucket, INCOMPLETE [NO_DATA] when the session has no readable pair —
never an empty green. The fixture's fifth assertion was itself wrong (40x5
does not fit 180: the five critics never share one box — grill and council
are the spec phase, the post-build trio pairs); the comment and the assert
now carry the per-phase derivation. The approved plan's prose keeps the
older sentence — a correction to approved bytes moves its digest and needs
re-approval (ADR 0032's precedent), which is a decision for the owner, not
a quiet edit mid-execution.
…verdict in the result, TIMEBOXED at the bound

The same section in each of the five: the numbers contract.py owns, the I/O
contract (return the verdict, write no files), and the bounded exit — a
partial verdict is a verdict, silence is not. The 045 postmortem's worst
number was Security burning 102 minutes and dying with no verdict delivered;
this is the sentence that makes that shape impossible to leave silent. The
quoted corpus situations are untouched: skilleval stays 428 delta 0, the fog
ratchet passes.
…o field, and --tick refused every task that named two

The envelope requires file/check/rollback/done when per task (TASK_FIELDS);
tasks 4 and 7 named two files under a plural key the parser does not know,
so the verb that executes an approved plan could not tick its own tasks.
The house format is singular **file** with the list inside — 046 and 045
carry it 23 times. The correction moves the canonical plan digest; ADR 0034
re-approves the corrected bytes, the 0032 precedent.
…gest, superseding 0033

The 0032 shape: the field-name correction moves the canonical plan digest, so
the record names the bytes --tick verifies against instead of the ones it
refused. The drift reader learns 0033's moved plan row with the reason beside
it, the way it learned 0021, 0023 and 0031 before.
…ts on the arithmetic

The subcommand, the will banner naming the wall reading (D-046-05: a surface
the banner omits is a false statement about a run), and the hermetic test
that dispatches through report.main over a tmp root. The first check the
plan named was a shell substitution --tick cannot run without a shell, over
a file CI does not have — the plan's third defect, corrected beside the
code it corrects; the re-approval record follows. vitals types its returns
dict[str, Any] so no type: ignore is needed anywhere (rule 3).
…hanical defect, superseding 0034

Task 4's check was a shell substitution --tick cannot run without a shell,
over a gitignored file CI has none of; the correction is the hermetic pytest
command shipped with the verb. The drift reader learns 0034's moved row the
way it learned 0033's, with the reason beside each.
…ed list printed at the end

B4 of the 045 postmortem: fail-fast made k independent defects cost k full
passes of thirteen minutes, measured three whole gates between 03:59 and
04:06Z. Each step keeps just's verified 1.58 '-' prefix (continue after a
failed step) and names itself into .ai/check-all-red.txt — just runs every
line in its own shell, so a shell variable would die with the line that set
it and the prefix alone swallows the exit code the recipe exists to return.
The contract test pins the step list to check's own dependencies, so the two
cannot drift. check itself is unchanged: fail-fast stays the keyboard shape.
…te per red

ai-goal gains the unattended-cycle rule: no step waits for input, a refusal
is repaired or the turn closes with one BLOCKED line naming the unblock, a
fork past its box returns TIMEBOXED, and the handover prints report vitals
beside audit verify — the clock disqualifies and never approves. ai-build's
step 5 names the batch rule the 045 postmortem priced at three full gates in
eight minutes: one red joins the others, check-all runs once at close. The
fixture pins both skills to their real commands; skilleval stays 428.
…s to decide

The clock disqualifies and never approves is the sentence the entry closes
on; the two plan re-signatures (0034, 0035) are named because an executor
who checks the digests will find three, and the reason belongs beside the
number. The intent page is re-forged over the shipped budgets.
…the reading the pins ask for

GOVERNING_SKILL_TEXT re-pinned for ai-build at 9d83df38… — step 5 now names
just check-all at block close, step 7 still records the stop before stopping,
and the comment beside the digest says that reading, which is what the pin
asks for when it goes red. The report verb's declared surface grows from
seven to eight with vitals, and the docstring says why it belongs in the
sends-nothing class.
…ical claim is corrected, the draft names its pattern and the unattended note its landing
… repoint their authority, rewrite the last shared idiom, drop the dormant map pairs
@github-actions

Copy link
Copy Markdown

Recap: Handshake intake mechanisms

  • range ed3c836472b35426cdc89ca7109781c5e3463b4e..ddceda2c9a5f17f50943807efce788c7abe366e146 files changed
  • rendered by ai-eng report recap from the diff and nothing else
  • GitHub Pages is off for this repository, so the page is a run artifact: open the recap run

@github-actions

Copy link
Copy Markdown

Recap: Handshake intake mechanisms

  • range ed3c836472b35426cdc89ca7109781c5e3463b4e..a5a473a315ac32bcfb1da03fb98307dc4914255e47 files changed
  • rendered by ai-eng report recap from the diff and nothing else
  • GitHub Pages is off for this repository, so the page is a run artifact: open the recap run

@sonarqubecloud

Copy link
Copy Markdown

@soydachi
soydachi merged commit c0088fd into main Aug 29, 2026
57 of 69 checks passed
@soydachi
soydachi deleted the ship/048-handshake-intake branch August 29, 2026 16:43
@github-project-automation github-project-automation Bot moved this from Backlog to Done in ai-engineering Aug 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant