Skip to content

Commit 278c94b

Browse files
songoowclaude
andcommitted
fix(semantics): cover only the vocabularies that declare a slot
B5 is optional in #4447 and gated on concrete slot identity. The scan fell back to the vocabulary id when a registry entry declared no `literal_scan.field`, which analysed 25 of 26 vocabularies against a field name nobody had claimed exists; every row downstream inherited that guess. The fallback is gone. A vocabulary with no declared slot is reported by name under `missing_slot_identity`, counted in the header, and not analysed at all -- coverage is now the 1 of 26 the registry actually declares, matching the `literal_scan_fields:1/1` the drift smoke already reports. Three further claims the evidence did not establish: * A call argument was reported `pass_through`. `Kind(value)`, `int(value)` and `sink(value)` are the same syntax, and none of them show whether the value comes out unchanged; an f-string and a further attribute are conversions the same way, and `.value` was read as an enum unwrap on an object the scan had not established was an enum member. All of these now end the climb as `unknown` with the construct named. `pass_through` means the AST shows the value itself relocated -- returned, stored, placed in a structure -- with nothing applied to it. An observed branch still wins over an unresolved sibling use, because interpreting is the top of the rank. * An owner-member comparison resolved against module-level imports only, so a parameter, a local assignment, a local import or an `except` target that had taken the owner's name over was still credited to the registered owner. Scope shadows now accumulate outwards-in, as `scan_python_production` already did, and a shadowed operand is `unknown`, not the owner it resembles. * The printed row listing dropped every computed-key and unattributable site, so the table read as a complete census of the slot's readers. Both populations are printed now, in their own blocks under the same `--top` budget, since a single location-ordered list would bury 111 classified rows under 602 unattributed ones -- the same concealment spelled differently. Measured on the merged tree: 713 rows over 485 of 1213 tracked sources -- 0 read, 61 interpret, 9 pass_through, 643 unknown, a 90.2% unknown share; of the 111 rows attributed to `effective_action`, 41 are unknown (36.9%). The guessing implementation reported 921 rows at 78.9%. Smaller and less confident is the direction the evidence supports. The cross-vocabulary parse and scope caches go with the narrowing: they existed so 26 vocabularies shared one parse, and one covered vocabulary does not need an unbounded module-level cache keyed on a text hash. Removing them costs nothing measurable (5.9-6.7s over three runs, unchanged). Still advisory: reachable only via `--report --consumer-evidence`, still refusing to run without `--report`, still not referenced by the drift smoke. Refs #4447 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
1 parent bfdbda7 commit 278c94b

4 files changed

Lines changed: 483 additions & 316 deletions

File tree

‎docs/architecture/rfcs/semantic-vocabulary-convergence-v0.md‎

Lines changed: 64 additions & 32 deletions
Original file line numberDiff line numberDiff line change
@@ -824,7 +824,7 @@ on the next full-tree scan; genuine shared-contract changes still need review.
824824
| Measurement covers both carrier shapes and filters local naming | `uv run --extra test python -m pytest tests/architecture/test_semantic_inventory.py` | pass, including the collision and module-local-convention fixtures | Rules come from this RFC, not from scanner output |
825825
| No behavior change from the two owner fixes | `uv run --extra test python -m pytest tests/test_loopx_turn_transaction.py tests/test_loop_turn_loop_controller.py tests/test_turn_loop_disposition.py tests/test_loopx_turn_managed_step.py tests/control_plane -k authority` and `uv run --extra test loopx canary premerge --from-git-diff` | pass | Environment failures already present on `main` are excluded when reproduced on a clean tree |
826826
| Docs governance accepts the RFC pair | `python3 examples/docs-governance-smoke.py` | pass | Checks mirror, links, index |
827-
| Consumer roles are reported per site, with the unknown stated (B5) | `uv run python scripts/generate_semantic_inventory.py --report --consumer-evidence` and `uv run --extra test python -m pytest tests/architecture/test_semantic_consumer_report.py` | the report prints read/interpret/pass-through/unknown counts, both unknown shares, and every unknown reason with its site count; the tests pass | Advisory only: syntactic use over a two-root scan reach, never data flow and never a gate |
827+
| Consumer roles are reported per site for vocabularies that declare a slot, with the unknown stated (B5, optional) | `uv run python scripts/generate_semantic_inventory.py --report --consumer-evidence` and `uv run --extra test python -m pytest tests/architecture/test_semantic_consumer_report.py` | the report names its coverage and the vocabularies it refused to analyse, then prints read/interpret/pass-through/unknown counts, both unknown shares, and every unknown reason with its site count; the tests pass | Advisory only, and optional in #4447: syntactic use over a two-root scan reach, never data flow and never a gate. Coverage is the 1 of 26 registered vocabularies that declares `literal_scan.field`; the other 25 are reported as `missing_slot_identity` and are not analysed |
828828
| Retirement budgets use standalone field tokens | `count_identifier_modules()` uses identifier boundaries for the six fields | `goal_boundary`: 30 Python modules under the new metric; the old substring metric was 35 | Conservative lexical measure; it removes compound-name false positives but does not prove semantic reader absence |
829829
| The module-local convention filter is a code edit | Widen `MODULE_LOCAL_CONVENTION` in `inventory.py` and scan | `*_semantic` budgets fall with no code change elsewhere | Known boundary; the regex is in code so the widening is a reviewed diff, and the unfiltered totals stay budgeted |
830830
| A registered value nobody produces fails (M0.5) | Run the production-form scan on the baseline | Fails naming `effective_action` and `skip`; passes after `skip` is removed or listed `compatibility_only` | First expected I12 failure; a compared-only value is not carried |
@@ -1124,47 +1124,78 @@ introduce a competing target state.
11241124

11251125
## Appendix A: Execution ledger (non-normative)
11261126

1127-
### 2026-09-17 — B5: consumer roles reported per site, with the unknown stated
1127+
### 2026-09-17 — B5: consumer roles for the vocabularies that declare a slot
11281128

1129-
Non-normative; advisory evidence only. No check changes its pass/fail result on
1130-
the current tree, and nothing added here gates a merge.
1129+
Non-normative; advisory evidence only, and B5 is an optional item of #4447. No
1130+
check changes its pass/fail result on the current tree, and nothing added here
1131+
gates a merge.
11311132

11321133
- `consumer_ranking` counts modules that *mention* a symbol. Section 5 already
11331134
says that number "does not classify roles or prove data flow", so B5 adds
11341135
`loopx/semantics/consumer_report.py`: a bounded AST scan that classifies each
11351136
consuming site as `read`, `interpret`, `pass_through` or `unknown` and carries
11361137
per row the location (`module::symbol` and line), the source SHA the scan ran
11371138
against, the value domain, and the limitation that applies to that row.
1139+
- **Coverage is the registry's declaration, not a guess, and it is 1 of 26.**
1140+
B5 is gated on concrete slot identity, and only `effective_action` declares
1141+
`literal_scan.field` — the same `literal_scan_fields:1/1` the drift smoke
1142+
already reports. The first implementation fell back to the vocabulary id when
1143+
no field was declared, which analysed 25 vocabularies against a field name
1144+
nobody had claimed exists and let every downstream row inherit the guess. The
1145+
fallback is removed: a vocabulary with no declared slot is reported by name
1146+
under `missing_slot_identity`, counted in the header, and not analysed at all
1147+
— not partially, and not by owner class alone.
11381148
- Anchors are identities the registry already carries, which is why B2 was the
1139-
precondition: the vocabulary's slot name (`literal_scan.field`, else the
1140-
vocabulary id) and its registered owner class, bound through the same
1141-
one-unrenamed-hop import discipline the producer scanner uses. No module
1142-
registers itself as a consumer, and the scan reach is code owned exactly as
1143-
`PRODUCER_ROOTS` is, so registry data cannot widen it.
1144-
- Measured on `440b002fb` across `loopx/control_plane` and `loopx/cli_commands`,
1145-
484 of 1207 tracked sources: **921 rows — 3 `read`, 141 `interpret`,
1146-
50 `pass_through`, 727 `unknown`, a 78.9% unknown share.** Restricted to the
1147-
319 rows that belong to a single vocabulary, the unknown share is 39.2%.
1149+
precondition: the declared slot name and the registered owner class, bound
1150+
through the same one-unrenamed-hop import discipline the producer scanner
1151+
uses, and only where no nearer binding has taken the owner's name over. The
1152+
scan reach is code owned exactly as `PRODUCER_ROOTS` is, so registry data
1153+
cannot widen it.
1154+
- **A role is claimed only from syntax that establishes it.** A call, an
1155+
f-string and a further attribute all end the climb as `unknown` with the
1156+
construct named. `Kind(value)` and `sink(value)` are the same syntax, so
1157+
neither may be read as proof that the value came out unchanged; the earlier
1158+
implementation reported both as `pass_through`, and `.value` as a
1159+
meaning-preserving enum unwrap on an object it had not established was an
1160+
enum member. `pass_through` now means the AST shows the value itself
1161+
relocated — returned, stored, placed in a structure — with nothing applied.
1162+
An observed branch still wins over an unresolved sibling use, because
1163+
interpreting is the top of the rank and nothing stronger could be hidden.
1164+
- Measured on `d8e7af141` across `loopx/control_plane` and `loopx/cli_commands`,
1165+
485 of 1213 tracked sources: **713 rows — 0 `read`, 61 `interpret`,
1166+
9 `pass_through`, 643 `unknown`, a 90.2% unknown share.** Restricted to the
1167+
111 rows that belong to `effective_action`, the unknown share is 36.9%.
1168+
Against the guessing implementation this is 921 rows down to 713 and a 78.9%
1169+
unknown share up to 90.2%: the report became smaller and less confident, which
1170+
is the direction the evidence supports.
11481171
- The unknown is most of the measurement, not a residue to tidy away. 602 sites
1149-
read a mapping under a computed key and are therefore unresolved for every
1150-
vocabulary at once; 76 modules spell a slot with no recognized anchor;
1151-
46 tracked TypeScript sources carrying a slot are not walked because there is
1152-
no TypeScript parser on this path; 3 Python sites are an unstable local or an
1153-
unclassified context. Each is a row with a location and a recorded reason,
1154-
following the pattern B2 set for unresolved producer sites.
1155-
- What the report establishes is **syntactic use, not data flow**. A row says
1156-
this location performs a recognized read of this slot and that the syntax
1157-
around the read branches on the value or forwards it. It does not prove the
1158-
value came from a registered producer, that the branch is reachable, or that
1159-
a vocabulary with no rows has no reader — the computed-key population is
1160-
precisely why that last claim cannot be made from a name-keyed scan.
1172+
read a mapping under a computed key and are therefore unresolved for the
1173+
covered vocabulary; 18 modules spell the slot with no recognized anchor;
1174+
13 sites hand the value to a callee this scan does not follow; 9 tracked
1175+
TypeScript sources carrying the slot are not walked because there is no
1176+
TypeScript parser on this path; 1 is an unstable local. Each is a row with a
1177+
location and a recorded reason, following the pattern B2 set for unresolved
1178+
producer sites, and neither population is filtered out of the printed listing
1179+
any more: the classified rows and the unattributable ones now get the same
1180+
`--top` budget in their own blocks, because a table that showed only what the
1181+
grammar resolved read as a complete census of the slot's readers -- and a
1182+
single location-ordered list would have buried the classified rows under the
1183+
602, which is the same concealment spelled differently.
1184+
- What the report establishes is **syntactic use, not data flow**. It does not
1185+
prove the value came from a registered producer, that the branch is reachable,
1186+
or that a vocabulary with no rows has no reader — the computed-key population
1187+
is precisely why that last claim cannot be made from a name-keyed scan.
11611188
- Exposed through the existing report surface as
1162-
`scripts/generate_semantic_inventory.py --report --consumer-evidence`, opt-in
1163-
because it is extra work: the per-site scan costs 3.35-3.65s over three
1164-
isolated runs on the full tree, against roughly 217s for the ranking that
1165-
`--report` already prints. The drift smoke does not call it, and the only
1166-
pull-request path that reaches `--report` is a pytest fixture repository of
1167-
three files, so the per-PR cost is unchanged.
1189+
`scripts/generate_semantic_inventory.py --report --consumer-evidence`, which
1190+
refuses to run without `--report` because it is advisory evidence, not a
1191+
check. The per-site scan costs 4.1-4.4s over three runs on the full tree,
1192+
against roughly 217s for the ranking that `--report` already prints. The drift
1193+
smoke does not call it.
1194+
- **What this leaves B5 worth.** One vocabulary, 111 attributed rows, 41 of them
1195+
unknown, and no second vocabulary can be covered until an M1/M3 migration
1196+
declares a slot for it. The honest reading is that the machinery is ahead of
1197+
the registry it reads; the value arrives when the first migration needs it,
1198+
not before.
11681199
- Not addressed here: the reach is two roots rather than the tree; TypeScript is
11691200
counted but not parsed; and following a local is one hop, so a value moving
11701201
through two aliases is unknown rather than traced. Widening any of the three
@@ -1418,7 +1449,7 @@ result on the current tree; what changes is what the invariants claim.
14181449
| 2026-09-16 | Q9: compute the full inventory on demand; retire the committed census | Implementation for [maintainer feedback](https://github.com/huangruiteng/loopx/pull/4360#issuecomment-5692062394); PR review pending | Committed snapshot with post-merge regeneration; diff-only scan rejected | 1, I6, 3, 5, 9, 10, 12 |
14191450
| 2026-09-16 | B2: bind one unrenamed re-export hop in the Python producer scanner | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447) B2; PR review pending | Require every consumer to import the owner module (fragile; failed silently in M2); unbounded multi-hop resolution rejected | 5, Appendix A |
14201451
| 2026-09-16 | B1 rename invariance: add the name-keyed divergence advisory; state the limit it does not close | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447) B1; PR review pending | Keying the budget on value sets (rejected: `CONFIDENCE_LEVELS` and `EDGE_CASE_COMPLEXITIES` share `high/low/medium` with different meanings); a committed name ledger (rejected at M0: Q9 retired the committed census). The advisory lists surviving forks by name; it was first described as catching a one-sided rename, which measurement disproved, so both mirrors state the limit as it behaves | 9 |
1421-
| 2026-09-17 | B5: report consumer roles per site from registry-anchored AST evidence, and state the unknown share instead of classifying everything | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447) B5; PR review pending | Register consumers in the registry (rejected: the tracking issue forbids blanket consumer registration, and a declared list is a claim rather than evidence); make the report a merge gate (rejected: F3 is the advisory lane, and a 78.9% unknown share cannot gate anything); report only the sites the grammar resolves (rejected: the per-vocabulary tables would read as complete, so an unrecognized mention and a computed-key read are rows with reasons) | 9, Appendix A, Appendix B |
1452+
| 2026-09-17 | B5 (optional): report consumer roles per site only for vocabularies that declare `literal_scan.field`, and state the unknown share instead of classifying everything | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447) B5; PR review pending | Fall back to the vocabulary id when no slot is declared (rejected on review: it analysed 25 of 26 vocabularies against a field name nobody declared, and every row downstream inherited the guess; they are now named under `missing_slot_identity` and not analysed); register consumers in the registry (rejected: the tracking issue forbids blanket consumer registration, and a declared list is a claim rather than evidence); make the report a merge gate (rejected: F3 is the advisory lane, and a 90.2% unknown share cannot gate anything); report only the sites the grammar resolves (rejected: the tables would read as complete, so an unrecognized mention and a computed-key read are printed rows with reasons) | 9, Appendix A, Appendix B, Appendix C |
14221453
| 2026-09-17 | B0: state schema validation, implementation stage, evidence status and blocking behaviour separately for I2/I11-I14 and the enforcement lanes; require each formal invariant id exactly once | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447) B0; PR review pending | Rename the `blocking_next` lane to match its behaviour (rejected: the lane name is the milestone that owns the check, and renaming it would lose that and collapse the two readings the other way); add a `blocks_today` boolean to `formal_model` (rejected: it would be one more declared field a reader could mistake for a measurement, and the fact is a property of the smoke's `main()`, which no registry edit can change); leave the lane gloss and note the gap in the ledger only (rejected: the gloss is the sentence a reviewer quotes) | 2, 5, 11, Appendix A, Appendix B |
14231454
| 2026-09-17 | Bound F1/F2 to the kernel tier and the scan reach, restate F4 as scope enumeration completeness, and give every obligation a derived `domain` | Implementation, Refs [#4447](https://github.com/huangruiteng/loopx/issues/4447); **kernel-maintainer approval required, not yet given** | Leave the unconditional statements and record the gap in prose only (rejected: the statement was stronger than `validate_production`'s own docstring); restate F4 as per-context value-set disjointness (rejected: refuted by the repo's own data, since `scope_declarations` exists to permit legitimate same-name reuse); widen the scan so the unconditional claim becomes true (rejected: a separate change with its own risk) | 5, 9, Appendix B, Appendix C |
14241455

@@ -1448,6 +1479,7 @@ result on the current tree; what changes is what the invariants claim.
14481479
| E21 | F1/F2 were unconditional but verified over one tier | `3ca868193` | `check_producers`' skip predicate, and the producer scan roots, read from the tree | 6 of 26 vocabularies declare `producers`, exactly the `tier: kernel` ones; the 20 skipped are all `cross_runtime`; the scan reaches 432 of 1203 tracked `loopx/**/*.{py,ts}` files (35.9%), the uncovered bulk being capabilities 285, other control-plane 192, extensions 83 | Counts from the registry and the tracked tree; the reach denominator moves with any new module, so it is reported, not pinned |
14491480
| E22 | Fifteen reported unresolved sites can never become evidence | `3ca868193` | smoke report `unresolved_producer_blockers` | 41 unresolved sites, of which `argument_name_only` 10 and `annotation_only` 5 are a field-named keyword argument and a bare declaration; the other 26 are dynamic or interprocedural | Label-keyed; the two labels are code-owned in the scanner, so the floor moves only by a code edit |
14501481
| E23 | F4 as written could not be violated | `3ca868193` | read `check_scope_declarations` against the F4 statement | Scope is declared and never inferred, so `conflict := collision ∧ scope_overlap` is a definition; what is enforced is that a declaration names every defining module exactly once, over 1 declaration and 4 contexts | Judgement from reading the check; value-set disjointness across contexts is deliberately *not* the property, because `SOURCE_SURFACES` legitimately reuses one name in four contexts (E19) |
1482+
| E26 | B5 consumer evidence covers 1 of 26 registered vocabularies | `d8e7af141` | `scripts/generate_semantic_inventory.py --report --consumer-evidence`, cross-read against the drift smoke's `literal_scan_fields` coverage | 1 vocabulary declares `literal_scan.field` (`effective_action`) and is analysed; 25 are reported `missing_slot_identity` and are not analysed. 713 rows over 485 of 1213 tracked sources: 0 `read`, 61 `interpret`, 9 `pass_through`, 643 `unknown` (90.2%); of the 111 rows attributed to `effective_action`, 41 are unknown (36.9%). The removed vocabulary-id fallback had reported 921 rows at a 78.9% unknown share | Coverage is a registry property, not a code one: it moves only when a vocabulary declares a slot. Rows are syntactic use over a two-root reach, never data flow, and the scan is advisory — it refuses to run without `--report` |
14511483
| E13 | The conflict budget mostly measured local naming | `1dc6ad8d8` | `MODULE_LOCAL_CONVENTION` applied to `conflicting_values` and `same_runtime_forks` names | 16 of 18 conflicts and 7 of 25 forks are module-local conventions; the semantic subsets are 2 and 18 | Classification is a name pattern, documented in the scanner and pinned by a fixture test |
14521484

14531485
## Appendix D: Rejected or superseded alternatives

0 commit comments

Comments
 (0)