048: the spec interview now keeps a live draft and a plain-words check before anything is written - #712
Merged
Merged
Conversation
…auditors, the orphan-delete pattern as the central finding The claim verified against the tree, not the ticks: twelve modules from specs 029-037 were written, checked, ticked [x], never wired, and deleted by 044 the same day 042 built the register whose plan forbade exactly that deletion (plan.md:38 quoted). 035 never existed; 023-T6/7 and the B4 batched gate are promised prose; ~30 research P0s have neither adoption nor recorded refusal. 045/046 survive independent refutation intact.
…ule — spec 044 cut twelve modules and left six sentences pointing at the corpses trim, skillify, versions.verify_against_installed and loopgate are gone from src/, but ai-note, ai-review and ai-security corpora still told the agent to call them, ai-challenge and ai-council still named loopgate as the orchestrator's instrument, and test_skill_bounds pinned the word loopgate so the dead reference could not be repaired without going red. The quoted situations are untouched — skill_eval counts only those, so the routing baseline stays 428 with delta zero. The cap now names the rule itself: two digest-equal green runs.
…ork file — and the test stops accepting the Spanish it translated
The reference carried the four basics in Spanish while every skill, spec and
corpus speaks English, and test_reference_names_the_rule accepted either
('keyboard' or 'teclado'), which is how the drift hid. The reference is
rewritten in English, the disjunction loses its second arm, and the seven
targeted tests stay green.
…ad, five dead with a named reopening condition The owner asked whether the deleted modules were necessary at all. Per module, in a note (append-only, with still_true_when so the verdict rots visibly if a restoration happens): the rule survives as prose where the module was scaffolding (trim, versions, skillify, constellation, decision_fw, evidencing, decision_boundary), and where the instrument is genuinely wanted (loopgate, lane_merge, verify_cold+answer_key, intake+bail_out) it returns only together with its consumer in the same commit — the orphan shape is what this refuses.
…il contract.py carries the number The owner set 180 minutes for a giant block (normal ones far under) against the measured 19 h 42 min of the only cycle ever timed. The row is INCOMPLETE and says which half: the ceiling exists as a decision and a postmortem, not as a constant a recap can fail on. Its evidence command greps contract.py for CYCLE_WALL_BUDGET_MINUTES, so the row closes only when spec 047 lands the number beside the other named budgets. COMMITMENTS moves 26 to 27 in the same commit — the count and the row are one change.
…ers as an outcome current_facts carried two claims the tree had already contradicted — four inherited test_madr failures (ADR 0025 closed them; 37 passed today) and the 22-link audit red (audit verify exits 0). The owner's ceiling becomes an intended outcome: a governed cycle fits 180 minutes for a giant block and less for normal ones, held open by PO-27 until contract.py carries the number. The page is re-forged over the same bytes.
… derived report The postmortem's five priced cuts, turned into eight atomic tasks: three named budgets in contract.py (PO-27's evidence command answers at task 1), a red fixture first, a stdlib vitals reader over the event stream the hooks already stamp, the report verb, the five critics carrying their box and a TIMEBOXED exit, the batched check-all with the verified 1.58 prefix, the goal's anti-stall rule, and a close that lets the clock disqualify and never approve. Draft until the owner signs the digests.
… three real defects and the spec answers all three Grill caught the context naming `stamp` for wall arithmetic when `stamp` is the chain-seal hash and `ts` is the time (task 3 rewritten around `ts`, the attribution rule named), caught the 120-calls figure cited to a measurement the postmortem does not carry (now stated as derived: 409/5, rounded), and caught a stale line count. The council's CLAIMED-NOT-REAL sweep over the eight tasks was refuted by the cross-read — a draft with every box `[ ]` is not a lie — and its one surviving gap candidate refuted itself. The council reader passes: RAN council=64/88 with 047's counts stated and reproduced.
…xecutes must exit zero, and a red fixture cannot satisfy a tick by being red The plan's own mechanics caught what the critics did not: --tick refuses a non-zero check, so the red-fixture task needed a check that proves the shape (collect-only, imports inside test bodies) while the run itself stays red. The house precedent is 046 task 3, whose check names the three refusals passing — red-then-green ticks on the green half.
…, execution directed The owner approved in conversation on 2026-08-29 at spec 1ec7e171… and plan e4252e37… (canonical, tick column masked). The record names the nine defects the critics and the tick mechanics found before signing — stamp-is-not-time, the invented 120-call measurement, the host-transcript buckets, the red fixture that could never seal — and what the approval knowingly costs: derived boxes, gaps-not-phases for silent forks, and check-all degrading to check if just ever drops the prefix.
CYCLE_WALL_BUDGET_MINUTES 180 (the owner's decision, PO-27's row carries it), CRITIC_TIMEBOX_MINUTES 40 and CRITIC_CALLS_MAX 120 (derived, and the comment says so — 409 measured calls over five critics, rounded down). The only place the numbers live; the clock disqualifies and never approves.
…sts and sealed it
…failing assertions, imports inside test bodies --collect-only exits zero (six tests gathered) while the run is red naming every missing half: the five critics carry no box and vitals does not exist. The plan's check is the collection precisely because --tick seals a zero exit: a red fixture cannot satisfy a tick by being red, so the check proves the shape and the done-when says which half stays red on purpose.
…o the earlier event's cls stdlib-only, reads the text of the event record and a session id; verdict PASS inside CYCLE_WALL_BUDGET_MINUTES, INCOMPLETE [OVER_BUDGET] naming the largest bucket, INCOMPLETE [NO_DATA] when the session has no readable pair — never an empty green. The fixture's fifth assertion was itself wrong (40x5 does not fit 180: the five critics never share one box — grill and council are the spec phase, the post-build trio pairs); the comment and the assert now carry the per-phase derivation. The approved plan's prose keeps the older sentence — a correction to approved bytes moves its digest and needs re-approval (ADR 0032's precedent), which is a decision for the owner, not a quiet edit mid-execution.
…verdict in the result, TIMEBOXED at the bound The same section in each of the five: the numbers contract.py owns, the I/O contract (return the verdict, write no files), and the bounded exit — a partial verdict is a verdict, silence is not. The 045 postmortem's worst number was Security burning 102 minutes and dying with no verdict delivered; this is the sentence that makes that shape impossible to leave silent. The quoted corpus situations are untouched: skilleval stays 428 delta 0, the fog ratchet passes.
…o field, and --tick refused every task that named two The envelope requires file/check/rollback/done when per task (TASK_FIELDS); tasks 4 and 7 named two files under a plural key the parser does not know, so the verb that executes an approved plan could not tick its own tasks. The house format is singular **file** with the list inside — 046 and 045 carry it 23 times. The correction moves the canonical plan digest; ADR 0034 re-approves the corrected bytes, the 0032 precedent.
…gest, superseding 0033 The 0032 shape: the field-name correction moves the canonical plan digest, so the record names the bytes --tick verifies against instead of the ones it refused. The drift reader learns 0033's moved plan row with the reason beside it, the way it learned 0021, 0023 and 0031 before.
…ts on the arithmetic The subcommand, the will banner naming the wall reading (D-046-05: a surface the banner omits is a false statement about a run), and the hermetic test that dispatches through report.main over a tmp root. The first check the plan named was a shell substitution --tick cannot run without a shell, over a file CI does not have — the plan's third defect, corrected beside the code it corrects; the re-approval record follows. vitals types its returns dict[str, Any] so no type: ignore is needed anywhere (rule 3).
…hanical defect, superseding 0034 Task 4's check was a shell substitution --tick cannot run without a shell, over a gitignored file CI has none of; the correction is the hermetic pytest command shipped with the verb. The drift reader learns 0034's moved row the way it learned 0033's, with the reason beside each.
…ed list printed at the end B4 of the 045 postmortem: fail-fast made k independent defects cost k full passes of thirteen minutes, measured three whole gates between 03:59 and 04:06Z. Each step keeps just's verified 1.58 '-' prefix (continue after a failed step) and names itself into .ai/check-all-red.txt — just runs every line in its own shell, so a shell variable would die with the line that set it and the prefix alone swallows the exit code the recipe exists to return. The contract test pins the step list to check's own dependencies, so the two cannot drift. check itself is unchanged: fail-fast stays the keyboard shape.
…te per red ai-goal gains the unattended-cycle rule: no step waits for input, a refusal is repaired or the turn closes with one BLOCKED line naming the unblock, a fork past its box returns TIMEBOXED, and the handover prints report vitals beside audit verify — the clock disqualifies and never approves. ai-build's step 5 names the batch rule the 045 postmortem priced at three full gates in eight minutes: one red joins the others, check-all runs once at close. The fixture pins both skills to their real commands; skilleval stays 428.
…s to decide The clock disqualifies and never approves is the sentence the entry closes on; the two plan re-signatures (0034, 0035) are named because an executor who checks the digests will find three, and the reason belongs beside the number. The intent page is re-forged over the shipped budgets.
…dence command answers
…the reading the pins ask for GOVERNING_SKILL_TEXT re-pinned for ai-build at 9d83df38… — step 5 now names just check-all at block close, step 7 still records the stop before stopping, and the comment beside the digest says that reading, which is what the pin asks for when it goes red. The report verb's declared surface grows from seven to eight with vitals, and the docstring says why it belongs in the sends-nothing class.
… build will create
…, execution directed
…nd references/intake.md
…rity_role, register the moved plan row
…ical claim is corrected, the draft names its pattern and the unattended note its landing
…run plus the per-surface feature gaps
… repoint their authority, rewrite the last shared idiom, drop the dormant map pairs
Recap: Handshake intake mechanisms
|
Recap: Handshake intake mechanisms
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



What changed, in plain words: before the framework writes a specification, it now interviews you the way a careful person would — it asks one question at a time, keeps a running draft on disk so a crashed session can pick up where it stopped instead of starting over, and before writing anything down it repeats your idea back in two plain sentences and waits for you to say "yes, that's it". If nobody is at the keyboard (an unattended run), it no longer hangs waiting for that yes; it writes down that the check went unconfirmed and carries on, so the record always says what was and wasn't human-verified. Nothing new was added to the command list: this all lives inside the spec-writing step that already exists.
Spec:
specs/048-handshake-intake-mechanisms/spec.md(approved at digests bydocs/adr/0036anddocs/adr/0037; the critics' rounds are folded in its## Grilland## Councilsections).Look at first:
.agents/skills/ai-spec/references/intake.md— the new interview rules — and the rewritten step 0 in.agents/skills/ai-spec/SKILL.md, whose exact bytes are pinned twice intests/test_contracts.py(the verbatim step text and the whole-file digest).What I am not confident about: the two stored pins make step 0 cheap to change in this repository and expensive to change wrongly — the digest test is the only early warning; and the read-back confirmation is still enforced by prose, not by code (spec 048's risks name what mechanizing it must not repeat: 037's validator was built and deleted as an orphan). The
/ai-verifyand/ai-reviewverdicts and the security lane's named residuals (tick-digest binding, plan-id collision, history-scan gap on old tags) are recorded in the spec's sections, not hidden.Gate:
just checkexit 0;ai-eng audit verifyWARN (22 accounted test-buffer links);just securityexit 0 at pinned versions; all four plan tasks ticked by sealed commands over the 0037 digest.