Skip to content

feat: a tauceti-coverage:v1 header with the report's verdict on each layer - #14

Merged
kim-em merged 7 commits into
TauCetiProject:mainfrom
roed-math:coverage-header
Sep 25, 2026
Merged

kim-em merged 7 commits into
TauCetiProject:mainfrom
roed-math:coverage-header

Conversation

@roed-math

@roed-math roed-math commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Adds a second machine header to STATUS.md, tauceti-coverage:v1: the report's verdict on each layer of the roadmap (done / partial / untouched / unassessed, with one line on what remains), in a form a script can read. The prose already says these things; the header exists so forty roadmaps' worth of them can be put on one page without a person re-reading every report. The consumer is the Progress page on the TauCeti site (TauCetiProject/TauCeti#7502, merged), which today reads the same verdicts from a hand-transcribed file and reads them from this header when a report carries one.

What changes

The model supplies untrusted prose and assessments; code decides identity, window, validation, serialisation and publication.

  • plan records the roadmap's layers in the plan: the README's Layer / Lane / Part / Stage headings (or L0A-style labels), each with its id, title and line, plus readme_sha, a SHA-256 of the README's text. For an umbrella area it also lists sub_roadmaps, each with its own layers and README hash (below). A README with no such headings gives no layers. The extraction rule is the consumer's.
  • The writing prompt asks the model to end the status prose with a fenced ```coverage block holding a JSON array, one {"id", "state", "remaining"?} object per listed id. For an umbrella it is instead an object keyed by roadmap id, each value such an array. The state is one of four words; the note is optional, one line, at most 200 characters.
  • apply removes the block from the prose, decodes it, hands the plan-owned envelope (roadmap, commit, README hash) to files.require_coverage, then checks the one thing the gate cannot: that every listed layer appears exactly once and nothing else, restoring README order. A block that is present but not JSON (including JSON null), or does not fit, is refused on the worker. So is a missing block when the plan lists layers: a new report retires the site's hand transcription of the old one, so a headerless report would turn assessed layers into unassessed ones, and the model always has unassessed for a layer it cannot judge.
  • The format and the gate (files.py) treat the header as part of the canonical prefix: the line after the status header and nowhere else, the same roadmap and commit, a closed schema (bounded layer count, restricted ids, notes without angle brackets so nothing can close the HTML comment). require_coverage is the one schema definition and runs at both trust boundaries. A status file without the header is unchanged and still passes the gate.

The gate proves the header's shape, never its truth, as for the prose. It cannot check readme_sha (it never checks out the roadmap repository); the consumer does, and refuses a header whose hash does not match the README it reads, showing those layers as unassessed rather than applying an old verdict to new requirements. Layer ids alone are not a specification identity: requirements change under unchanged headings.

Umbrella areas

RepresentationTheory is one labelled area with one report, so that report is the only account of its twelve sub-roadmaps (105 layers). A sub-roadmap is a directory directly below the area with a README.md and a Suggested.lean, which is the consumer's rule. plan lists each one under sub_roadmaps, with its Area/Child id, README path, layers and its own readme_sha. The model reads those READMEs and answers with a block keyed by roadmap id. apply writes one tauceti-coverage:v1 line per sub-roadmap, with roadmap set to RepresentationTheory/<Child> and bound to that child's README, after the area's own line (if its README has layers) and in ascending order of name. The gate holds each to the same schema, requires a child directly below the status file's own area, caps them at 32, and treats all of them as part of the canonical prefix. The consumer change is TauCetiProject/TauCeti#8786; the consumer before it ignores these lines.

The standalone prompt status template is handed no layer list, stays prose-only, and says so.

Files

  • progress/layers.py (new): heading extraction, the block splitter, and the assembly of validated payloads, one per roadmap for an umbrella.
  • progress/files.py: the marker, its schema, the prefix, parsing and the reserved-marker scan.
  • progress/plan.py, progress/apply.py: carrying the layers and sub-roadmaps, and turning the block into the headers.
  • progress/prompts/progress.md: the block instructions. README.md: the schema, trust boundary, missing-block behaviour, README hash, scope and rollout, in one place.
  • Tests: tests/test_layers.py (new), and additions to test_files.py, test_apply_announce.py, test_window.py and test_prompts.py. Refusals are labeled tables. tests/fixtures/coverage-contract/ holds a README, a model body with its block, and the exact header line the consumer was recorded accepting; the apply tests prove the producer still emits it byte for byte, that the gate accepts the pair, and that editing the README changes the hash under stable ids. tests/fixtures/coverage-contract-umbrella/ does the same for an umbrella with two sub-roadmaps and a references/ folder that is not one. Nothing is fetched in CI. ./tests/run passes in full.

Cross-check with the consumer

Run locally, not in CI: the fixture header is accepted by TauCeti's scripts/roadmap_progress.py at b48d6c9 with the intended states and notes, and refused by it against a README with an added requirement, a README with a retitled layer under a stable id, and another library commit. The umbrella fixture's two sub-roadmap lines are accepted by the consumer with TauCeti#8786 (8809a66), and a child's line is refused once that child's README is edited. Against TauCetiRoadmap 9fad8ac, the producer and the consumer extract identical layer ids, lines and README hashes for all 50 roadmaps and all 12 sub-roadmaps. A worst-case RepresentationTheory report, with every layer carrying a 200-character note, is 34 KB of the 64 KB cap and passes the gate.

After merging

  • TauCetiRoadmap's progress-merge.yml pins this repository by SHA in two places; bump them to the merged SHA, together with the worker's PROGRESS_REF, so the generator and the gate run one version. Until then a report carrying the header would be refused by the old gate, and one without it is accepted by both.
  • No report changes until the pins move. Afterwards the gate still accepts a report without the headers, but the worker no longer writes one for a roadmap with layers.
  • Nothing is backfilled. Each roadmap, the twelve sub-roadmaps included, moves from its hand transcription in TauCeti's scripts/roadmap_coverage.json to its header at its next report, and its entry there can then be deleted. The sub-roadmaps need TauCeti#8786 merged by then.

🤖 Generated with Claude Code

roed-math and others added 3 commits September 19, 2026 03:03
…layer

A STATUS.md may now carry a second machine header beside tauceti-status:v1:
the report's state for each layer of the roadmap (done, partial, untouched,
or unassessed when the material says nothing), with one line on what remains,
in a form a script can read. The prose already says these things; the header
exists so forty roadmaps' worth of them can be put on one page.

The model still only writes prose. `plan` extracts the roadmap's layers from
its README (the Layer / Lane / Part / Stage headings, or L0A-style labels) and
records them with a SHA-256 of the README; the prompt asks for a trailing
```coverage block, one line per listed layer; `apply` removes the block from
the prose, checks it names every listed layer once and nothing else, and
writes the header; the gate treats the header as part of the canonical prefix
and validates its closed schema. A body with no block is a report with no
header, as before. The README hash is what binds an assessment to the
specification it was made against, since a layer's requirements can change
under an unchanged heading; the consumer refuses a header whose hash does not
match the README it read the layers from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e-only status prompt

Three things review asked for, none a change to the header itself.

The scope of this first version is now stated: `plan` reads the selected
labelled area's own README and nothing below it. An umbrella area whose README
is an index of sub-roadmaps (RepresentationTheory) has no layer headings, so
its reports carry no header and the children's layers are not in the plan at
all; on the Progress page they stay on hand transcriptions. The README says
so, and a planner test builds an umbrella with two layered children and
asserts the plan lists no layers. A second planner test asserts the exact
inventory and README hash for an ordinary area, not merely that a plan builds.

The standalone `prompt status` template is a second interface and was
silent about the block. It now says it is prose-only (it is handed no layer
list, so it must not invent ids) and that a report written from it carries no
coverage header; the CLI help says the same, and the prompt tests are named
by the interface they exercise, with a new one holding `status.md` to that.

tests/fixtures/coverage-contract/ is an offline producer/consumer fixture: a
README, the model's body with its block, and expected.json recording what
plan+apply emit and what TauCeti's scripts/roadmap_progress.py at 9477d5c
made of it (ids, lines, states, notes, and the refusals against an edited
README body, an edited title with stable ids, and another library commit).
test_contract.py proves the producer still emits that header byte for byte,
that the gate accepts the pair, and that a body without a block yields a
plain report. Nothing is fetched in CI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The model now writes a JSON array inside the trailing ```coverage block
instead of a line mini-language: split_block decodes it and coverage()
hands the plan-owned envelope to files.require_coverage, the gate's own
validator, then checks the plan's inventory (the one thing the gate cannot
check) and restores README order. A present block that is not JSON is an
error, never "no coverage". The wire format is unchanged.

The contract test and its fixtures shrink to a README, a model body and
the recorded header line, folded into the apply tests; the schema and
block refusals become labeled tables. The README carries the schema,
trust boundary, missing-block behaviour, README hash, scope and rollout
once, and the module prose that repeated them is cut.

Checked: tests/run green; the fixture header accepted by TauCeti's
scripts/roadmap_progress.py at b48d6c9, and refused against an edited
README and another library commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@CBirkbeck

Copy link
Copy Markdown

Review of 26edc5bb08ef8fa9f90654eae48ccd8ea629039c

The overall design looks sound, but I recommend fixing one validation bug before approval. There is also a small, nonblocking mathematical correction in the test fixture.

1. Fix before approval: JSON null is incorrectly treated as an absent coverage block

Location: progress/layers.py, in split_block(), and progress/apply.py:363–371, in render_update().

split_block() uses Python None to indicate that no coverage block was supplied. However, json.loads("null") also returns None.

Consequently, an otherwise valid status body ending with:

```coverage
null
```

is accepted even when the plan has a nonempty layers list. The block is removed, render_update() treats it as absent, and layers_mod.coverage() / files.require_coverage() are never called. The report is then accepted without a coverage header, rather than rejecting the invalid payload.

This contradicts the documented distinction between:

  • A genuinely absent block, which is intentionally allowed.
  • A present block whose payload does not satisfy the schema, which should be refused.

This is a limited validation/data-loss bug, not a way to forge a successful assessment.

Suggested fix: distinguish block absence from its decoded JSON value using a separate presence flag or a unique sentinel. Alternatively, explicitly reject decoded None inside split_block().

Please add an apply-level regression test, using a plan with layers and otherwise valid prose, checking that:

  • A block containing null raises files.FormatError.
  • A genuinely missing block still produces an accepted headerless report.
  • A valid array naming every planned layer still round-trips successfully.

2. Nonblocking: incorrect point count in the fixture

Location: tests/fixtures/coverage-contract/README.md:27–29.

The example y² = x³ + 1 over F₇ says “Six points”. It has 11 affine points, hence 12 rational points including infinity.

For x = 0, 1, 2, 3, 4, 5, 6, the numbers of possible y values are respectively:

2, 2, 2, 1, 2, 1, 1

Please correct the sentence and update the fixture’s expected README hash and canonical header in expected.json accordingly. This does not affect the coverage implementation itself.

Otherwise

The shared worker/gate schema, canonical header placement, README and library-commit binding, and complete planned-layer inventory check look sound. Headerless-report compatibility and the initial exclusion of umbrella sub-roadmaps are intentional, documented scope decisions—not blockers.

CI passed for the reviewed head. The review did not include a live worker publication or complete website build.

Once the null case is fixed and regression-tested, I found no further blocker within the reviewed scope. On deployment, retain the documented coordinated update of the two gate pins and the worker’s PROGRESS_REF, since the old gate rejects the new coverage header.

split_block returned None both for "no block" and for a block whose JSON
decoded to null, so a null block was dropped and the report published
without a coverage header instead of being refused. Reject null in
split_block, with an apply-level test that null is refused, an absent
block still gives a headerless report, and a full block still round-trips.

Also correct the fixture's point count (y^2 = x^3 + 1 over F_7 has 11
affine points, 12 projective) and re-record its README hash and header;
the consumer at the recorded revision accepts the new header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@roed-math

Copy link
Copy Markdown
Contributor Author

Addressed in 5ba2711.

1. null coverage block. split_block now refuses a block whose JSON decodes to null, so it can no longer collide with the None that means "no block". New apply-level regression test in tests/test_apply_announce.py, with a plan that lists layers and otherwise valid prose:

  • a block holding null raises files.FormatError;
  • a genuinely missing block still gives an accepted headerless report;
  • a full block naming every planned layer still round-trips.

tests/test_layers.py also gains a null case among the blocks split_block must refuse.

2. Fixture point count. tests/fixtures/coverage-contract/README.md now reads "Eleven affine points, so twelve with the point at infinity." The README changed, so readme_sha and header_line in expected.json are re-recorded. I re-ran the new header through the consumer (scripts/roadmap_progress.py at b48d6c9): states_from_marker accepts it with the same layer ids and line numbers. The wire format is unchanged, so consumer_revision stays.

./tests/run passes. The deployment note stands: bump both gate pins in progress-merge.yml and the worker's PROGRESS_REF together.

@kim-em

kim-em commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

The producer/consumer contract looks sound, and the hand-transcription gap on the Progress page is real. I have two follow-ups:

  1. Require a coverage block for a new worker-generated report when the plan lists layers. render_update() currently accepts an absent block and publishes a headerless STATUS.md. That preserves compatibility with older reports, but the newly published report changes report_sha, so the site's existing hand transcription is retired and its layers become unassessed. Since the new prompt requests a block (and unassessed is a valid state), could the worker reject an omission for a nonempty plan["layers"]? The merge gate can continue accepting older headerless files. An apply-level test for the omitted-block case would pin this distinction.

  2. Clarify the path off hand transcriptions. Existing reports will not gain headers until regenerated, and this version does not emit coverage for the umbrella area's child roadmaps. The current companion file has 12 child-roadmap entries. Could the rollout note state whether those entries will remain maintained by hand, be backfilled, or be addressed in follow-up work? This is a scope question, not a request to add child-roadmap support to this PR.

A new report changes the report hash, which retires the site's hand
transcription of the old one, so publishing it headerless would turn
assessed layers into unassessed ones. When the plan lists layers the
worker now refuses a body with no ```coverage block; the model can
always say `unassessed`. The gate still accepts a status file without
the header, as every report before the header looks. The prompt says
the block is required, and an apply-level test pins both halves.

The README gains a rollout note on leaving the hand transcriptions:
nothing is backfilled, top-level areas move over at their next report,
and the twelve RepresentationTheory children stay hand-maintained
(retiring together at each umbrella report) until a follow-up gives
each child a carrier of its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@roed-math
roed-math requested a review from a team as a code owner September 25, 2026 05:59
RepresentationTheory is one labelled area with one report, and its twelve
sub-roadmaps (105 layers) could only be assessed by hand transcription,
all of which a new umbrella report retires at once. Now:

- plan lists `sub_roadmaps`: each directory directly below the area with
  a README.md and a Suggested.lean (the consumer's rule), with layers,
  its `Area/Child` id, README path and its own readme_sha;
- the prompt points the model at those READMEs and, for an umbrella,
  asks for a JSON object keyed by roadmap id instead of an array;
- apply writes one tauceti-coverage:v1 header per sub-roadmap, bound to
  its own README, after the area's own and in ascending order of name,
  and refuses a block that leaves a sub-roadmap or a layer out;
- the gate validates each (roadmap `Area/Child` directly below the status
  file's area, at most 32) as part of the canonical prefix.

A consumer that does not know these headers ignores them. The new
coverage-contract-umbrella fixture records the header lines TauCeti's
roadmap_progress.py (with the matching change, 8809a66) was checked
accepting; against TauCetiRoadmap 9fad8ac the planner and the consumer
agree on all twelve children, and a worst-case report is 34 KB of 64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@roed-math

Copy link
Copy Markdown
Contributor Author

Addressed in b9cad58 and d5a4e09.

1. When the plan lists layers, render_update now refuses a body with no ```coverage block, and the error points the model to unassessed. The gate still accepts headerless status files. There is an apply-level test for both sides: no block with layers is refused, and a headerless file still passes validate_update. The prompt and the README's "Missing block" paragraph say the same.

2. Rather than leave the children on hand transcription, this PR now covers them. For an umbrella area the plan lists each sub-roadmap, a directory below the area with a README and a Suggested.lean (the consumer's rule), with its own layers and README hash. The model answers with a block keyed by roadmap id. apply writes one tauceti-coverage:v1 line per child, with roadmap set to RepresentationTheory/<Child> and bound to that child's README. The gate validates each as part of the canonical prefix, at most 32, each naming a child directly below the report's own area.

The consumer change is TauCetiProject/TauCeti#8786: a child reads its own marker from the inherited report. The consumer before that change ignores these lines, so the two can merge in either order. Against TauCetiRoadmap 9fad8ac, the producer and the consumer agree on all 12 children (105 layers), and a worst-case umbrella report is 34 KB of the 64 KB cap. A new fixture, coverage-contract-umbrella, records the header lines the consumer was checked accepting.

Nothing is backfilled: each roadmap switches to headers at its next report, and its transcription entry can then be deleted. The README's rollout note and the PR description are updated to match.

The prompt called the facts file ground truth and said a result not in it
did not land, while the coverage block describes the whole roadmap. Read
cautiously, that turns every layer finished in an earlier window into
`unassessed` on the Progress page, which matters far more now that an
umbrella report states all 105 RepresentationTheory layers. The facts file
is now ground truth for this window; the previous STATUS.md (and its
coverage headers) and PROGRESS.md are the evidence for earlier ones, and
their verdicts stand unless this window or a changed README gives a reason
to revise them. A prompt test pins the wording.

A `remaining` note could hold a lone surrogate spelled as a JSON escape,
which passed validation and then failed with UnicodeEncodeError when the
header was written unescaped. REMAINING_RE now excludes surrogates, so it
is a FormatError at the block like any other malformed note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kim-em
kim-em merged commit 6d26dd3 into TauCetiProject:main Sep 25, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants