Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces ADR 0001, which proposes a standard for defining a Claim's substance and canonicalizing it into a content-addressed identity (fingerprint) using JCS (RFC 8785) or RDFC-1.0, along with an updated ADR index. The review feedback highlights several technical areas for refinement: ensuring determinism for multivalued array fields by specifying a sorting requirement, avoiding timezone-shifting bugs by keeping pure date fields in YYYY-MM-DD format, clarifying how floating-point precision for quantity is handled under JCS, and resolving the logical tension between including usesMethodology in the substance definition while simultaneously questioning if its changes should mint a new identity.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 268e072ed0
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
@blushi — short version: you're right on every substantive point, and on D2 you were right for a bigger reason than either of us had. I went looking for a rebuttal to your D2 comment and found the argument that kills D2-b instead. Two corrections to things I believed going in: what production does today is deterministic The one place I'd still shape the outcome is migration: I don't think this forces re-minting every existing hash. Proposing a dual-anchor path in the revised ADR — existing Also going in off your review: D4's DoS guard becomes mandatory rather than conditional (and it matters more now that RDFC-1.0 is actually on the serving path), the spike gets published to Still worth a call with @JeancarloBarrios, but the question is now narrower and better: not "which algorithm", but what RDF encoding profile we both freeze so the projections agree. I'd like to get that scheduled this week — the Second Muse engagement starts mid-August and involves standing up a claims-engine instance, so this gets exercised by a real client soon and it's expensive to reverse afterward. |
Ten of the review comments, implemented. The two open questions (hash wire format, whether PROV-O should block) and the ConsentDirective split (pending @glandua's sign-off) are deliberately NOT in this commit. :213 — claims/attestations become claimRefs/attestationRefs, multivalued uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined objects during the consent rewrite without re-justifying it. A Claim's identity and lifecycle (verificationStatus, contentHash, supersession) evolve after emission, so embedding one couples this envelope's lifecycle to the claim's. Claim/Attestation imports dropped — a uriorcurie range doesn't need them. :112 — observedLocalDate removed. It was a pure function of observedAt and sourceTimezone, so storing it only created something that could disagree with its own instant. The derivation is now stated normatively on sourceTimezone. :83 — snake_case convention on sourceSystem and recordType, enforced with a pattern rather than described in prose. This immediately caught the very inconsistency reported: the biocultural example's "field-signing-app" failed validation and is now "field_signing_app". :211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0. :236 — PROV-O, narrowly. Added the prov prefix and emittedAt -> prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent: wasGeneratedBy relates an Entity to an Activity rather than a string field, and prov:Agent is a class not a property, so those mappings would produce something that looks PROV-aligned and isn't. Real alignment means modelling the sensor run as a prov:Activity — proposed as its own ADR. docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56. docs:131 — Pydantic section removed; binding is the consumer's concern and belongs in the consuming repo. Fixed the header line that referenced it. Codex P2 (raw-data residency) — the descriptions on rawContentInline and rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No such rule exists; the only rule requires rawContentHash when processingState is COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so the constraint is not expressible on this class without hoisting a mirror field (two sources of truth for one policy — worse than an honest gap). Descriptions now say enforcement is external, a new doc section states the consumer obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml is committed as a negative fixture: a SOVEREIGN record carrying inline raw that passes validation. Kept so the gap stays visible and so the fixture starts failing if a future LinkML or SHACL shape ever expresses it. Verified: both positive examples validate against the edited schema; the negative fixture validates (which is the point); linkml-lint reports no new problems — the one error it emits ("id is not a uri") is pre-existing and repo-wide, present on unmodified OutputRecord.yaml and on Claim.yaml as merged on main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Pushed D2 reversed. You argued for D2-a because I've recorded the two arguments of mine that failed, rather than quietly dropping them, so the reversal is auditable:
Your Your Two new sections from things your review implied:
The spike is published (koi-research#2) and the ADR now links permalinks instead of the scratch path you couldn't open. Its contemporaneous write-up still argues for D2-b; I left it as-run rather than back-editing, since rewriting evidence after the conclusion changed is what makes a spike untrustworthy. Findings 1 and 3 are algorithm-independent; finding 2 gets more load-bearing under D2-a, which is why D4 is now a hard release gate rather than a recommendation. What's left is narrower and it's @JeancarloBarrios's call. D2-a makes the RDF encoding profile load-bearing — namespace, Still open from you: the |
What kind of data was anchored as raw already on mainnet? |
|
Adding @blushi as a ratifier on this one. The Deciders line currently reads Darren + Shawn + JC. The D2 reversal on 31 July turned on the @DarrenZal — could you add her to Deciders next time you touch the file? I've requested her review here in the meantime. Two notes for the deciding call:
Separately, and not a nitpick: the reversal write-up is the right standard for these. Quoting the proto, naming the constraint neither option had been checked against, and recording the two corrections to the prior reasoning is what makes this ratifiable rather than merely decided. Also worth noting that the PR description still describes D2-b as the recommendation — anyone reading the summary rather than the file will get the wrong picture, so it's worth a quick edit. |
…assumption; pin all citations Addresses @blushi's 2026-08-03 review on regen-network#56. - Remove "Raised in review, and it is the stronger position." (D1) — process metadata, not decision rationale. - D2: RDFC-1.0 is now presented as the data module's own name for the algorithm, not a preference. GRAPH_CANONICALIZATION_ALGORITHM_RDFC_1_0 is the only non-zero value in the v2 canonicalization registry, and production targets v2 (derive_ledger_iri sends the v2-only file_extension field). Notes the v1 rename from URDNA2015 at the same wire value, and that the shared byte proves nothing about which spec ran. Adds the live divergence: the deployed graph path runs pyld URDNA2015 and stamps _GRAPH_CANON_URDNA2015 = 1. - Withdraw the claim that JC's work already assumes the lifecycle-record shape. Marie is closer to that codebase and reports no corresponding shapes exist. Booked as Open Question 7 (numbered 7 so refs to 1-6 stay valid), and the Context section now says the ADR does not describe JC's implementation. - D8 rewritten from dual anchoring to a straight cutover. Mainnet query found 8 Raw anchors from the claims engine, all test/demo, and zero executed MsgAttest. Method and bounds stated inline. Algorithm-column bullet promoted to D9, which stands on its own. - Every code citation pinned to a 40-hex commit SHA (27 links, all verified to resolve). No relative or branch links remain.
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses the lean I gave her, so the reasoning is in the doc under "Hash wire format" and summarised here. I said I'd probably go bare hex, to match the implementation and the ledger. Three findings support that, and I re-derived all three: - koi-processor emits bare lowercase hex — ledger_anchor.py:96, `hashlib.blake2b(..., digest_size=32).hexdigest()`. - `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented the token in these examples. - The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto), never a prefixed string. What flipped it: - A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB. `koi_memories.content_hash` is SHA-256 (migrations/004:28) and `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same shape, same length, different algorithm, one codebase. - ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the missing discriminator as a latent defect in its own words: "the current schema cannot express which canonicalization produced a stored hash." Shipping an undiscriminated hash in a WIRE CONTRACT repeats that defect where it is hardest to fix — the ledger and the claims tables can add an adjacent column later because they own both sides of the read. A contract consumed by parties we don't control cannot. So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash. Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the ledger's DigestAlgorithm registry (today exactly one non-zero member, DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum gains one. It does NOT name the ContentHash kind or a canonicalization. For rawContentHash that is complete — a raw payload is not RDF, anchors as ContentHash.Raw, and the proto says Raw "does not specify a deterministic, canonical encoding". For recordContentHash it is NOT complete: this envelope does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same question ADR 0001 is deciding for Claim. Flagged in the slot description rather than quietly decided. Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks string equality and dedup on a value whose whole purpose is content-addressed identity. Rejected multihash/multibase: properly self-describing, but needs a codec table on both sides, produces values nothing in our stack emits, and diverges from the ledger's registry for no gain we can spend today. Cost, stated rather than hidden: koi-processor must prefix at the contract boundary. One line, and the right direction — the standard sets the wire format and the implementation adapts. Not touched: Claim.contentHash and Attestation.contentHash on main have the same gap (range: string, no pattern). Same class as Entity having no identifier slot; belongs in the identity-keys follow-up, not in a PR editing merged schemas. Verified: `make -C schema lint` clean; all three examples pass `linkml-validate -C OutputRecord`; and the pattern was tested against what it must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest all fail with "does not match". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1. What kind of data was anchored as raw on mainnetYou asked about the chain, so I ran a chain-wide census rather than answering only for us. Over the tx index, 2026-02-24 → 2026-08-05: 950 Chain-wide, Raw is mostly documents and images, not JSON: pdf 259, jpeg 58, jpg 50, png 24, json 19, webp 5, heic 2, bin 1. And it is concentrated in one account:
So the registry account runs both paths side by side: Graph for ecocredit metadata IRIs, Raw for what looks like document attachments — 259 PDFs. I had been about to write you "the registry's own data is already all-Graph." That is true of the metadata IRIs and misleading about the account, so I've dropped it. Ours is 8 of 418 — 1.9%, and all eight are test or demo JSON: 4 claims (2026-03-11 ×2, 03-30, 04-23) and 4 attestation blobs (2026-03-11), every one That eight is exhaustive, by a cleaner argument than the per-IRI sweep I'd have offered otherwise: the prod account reports Scope and method, since a "not found" means nothing unless the query is known to work. REST One correction to something I nearly told you: 2. Two accounts I can't attribute — do you know whose they are?
3. The strongest argument for D2-a turns out to be on-chain alreadySame census: there are four So the case for D2-a is not "we have nothing to lose by switching." It is that Graph attestation with this exact algorithm is live working practice on this chain, and every claim our engine has anchored is structurally excluded from it. That is now the lead evidence in D2. Your point from a few months ago — if we want data attested it shouldn't be anchored as raw JSON — is the thing the chain is already doing. The v2 enum argument stands alongside it, with the v1 rename spelled out: 4. D8 still collapses, for a different reason than I first wroteI'm dropping the dual-anchor machinery, but the honest reason is corpus size, not corpus absence: re-anchoring eight throwaway fixtures is not a migration. My first pass at this justified it with "there is no production Raw corpus on mainnet." That was wrong by roughly a factor of 52 and it's withdrawn — 418 Raw anchors, 363 of them made after our last one, most recently 2026-08-05. What I can't answer, and D8 now says so rather than assuming it away: dual-anchoring also buys forward interoperability during a transition. "Our corpus is small" answers is re-minting expensive; it does not answer does anyone else resolve our Raw IRIs. Hence §2. Booked as Open Question 8 rather than decided. The one bullet inside old-D8 worth keeping is now D9: 5. The rest of your review
On the read-through askFair, and I'd rather give you a mechanism than a promise. Posted late — this was written on 6 August and then sat unsent while I was away 7–10 August. I've now pushed the branch: Your four comments from today land after all of the above — I'm reading them now and will answer each one directly, including the |
|
@LinuxIsCool @JeancarloBarrios @glandua — putting this to async sign-off rather than booking another call. Straight about why: Gregory offered Thursday or Friday afternoon on 2026-08-05, JC said both worked, and I never sent the invite. That one's mine. Rather than chase a slot now, I think a written record is the better artifact for a decision like this anyway — but the week of silence wasn't a scheduling problem, it was me. What changed, and whyD2-b (JCS / RFC 8785) is withdrawn. The forcing constraint is on-chain, not aesthetic:
// content_hashes are the content hashes for anchored data. Only RDF graph
// data can be signed as its data model is intended to specifically convey
// semantic meaning.
repeated ContentHash.Graph content_hashes = 2;Our production anchoring emits hash_json = {
"raw": {
"hash": hash_b64,
"digest_algorithm": 1, # DIGEST_ALGORITHM_BLAKE2B_256
"file_extension": "json", # no leading dot
}
}…over a canonical form that is To be precise about what "settled" means here: D2 is not settled because the spike succeeded — the spike self-reports PARTIAL FAILURE, and its 5-field substance set is genuinely defective (two ledger-distinct impact claims collapsed to one fingerprint). D2 is settled because the Independent corroboration, from reading JC's engine rather than reasoning about it: Proposal: split this ADR in twoRight now this document is partly decided and partly blocked on questions only JC can answer, which makes it un-mergeable in either direction. I'd like to split it:
The seam is clean and nothing else moves. The narrowed 0001 still cites the spike evidence and the MsgAttest constraint — those are why D2 was decided, not open questions. Marie — this probably invalidates part of your review, since your threads were written against the broad document. Rather than assume they carry over, could you say which still apply to the narrowed text? I'd expect a fresh pass on it. Two things that gate "Accepted"
The decision rule I'm proposingDeciders are Darren, Shawn, JC per the ADR's own Deciders line. Please approve or object in-thread by Monday 2026-08-17. If either Shawn or JC is silent at that deadline, I'm proposing Gregory decides as sponsor — he owns the contract commitment this gates. To be explicit: this rule is itself part of what's being signed off, and objecting to the process is a valid response. Nobody has authorised me to make Gregory the tiebreaker — he's the sponsor and owns the commitment, which is an argument, not authority. If you'd rather we did this differently, say so and we will. I don't want silence later characterised as consent nobody gave. One governance question, asked not assumed@glandua asked that @blushi be added to the Deciders line. She has done effectively all of the review on this ADR, so the substance of the request is obviously right — but Gregory isn't himself a decider, and Marie hasn't been asked whether she wants the role or its ongoing workload. So: Shawn, JC — do you agree? Marie — do you want it? I'd rather ask than edit the line and call it done. |
|
Re-stating my replies above in the format @davefortson proposed on 2026-08-04, because I answered the threads before applying the protocol and the threads are harder to scan without it. One correction to my own framing first: only the PR description has actually changed so far. The ADR file has not. So most of these are Deferred with a date, not Addressed — I'd rather say that than claim credit for edits that don't exist yet.
Two
The split itself now has a design-intention issue per protocol 5 — filed late, which is the wrong order: #60. Push back there rather than here if the seam looks wrong. |
|
One gap in the sign-off rule above that I should have caught when I wrote it: I can't merge this repo. My permission on Related, and worth stating because it has been quietly shaping this PR: the same permission means I can't request reviewers here either. #55, #56 and #58 all show zero requested reviewers, and I'd been reading that as an oversight on our side. It isn't — the button 403s for me. The only mechanism I have is @-mentioning people in a comment, which is what I've been doing. That also puts @blushi's review in a different light. She hasn't been the only human reviewing these because nobody else was asked; she's been reviewing voluntarily, because nobody can be formally asked. If we want review load spread beyond her, that's an access change rather than a process change. Two small asks that follow from it:
|
|
D7's ordering rule is not implementable for So for a claim with two co-benefits there is nothing guaranteed to be present to sort on, and two engines can order them differently and produce different fingerprints from the same claim. That is the determinism property D7 exists to provide. This matters now because D7 is one of the decisions proposed for
I lean (1), because it keeps the fix inside the canonicalization layer where D7 lives and doesn't reach into the schema. But (3) is the honest minimum if we want D7 Accepted on Monday without deciding (1) properly. Found while drafting the split and reviewing it adversarially before opening a PR — noting that because it's the pre-work check @davefortson's protocol 3 asks for, and this is what it caught. |
@DarrenZal I was only interested in what was anchored as raw as part of the current context, ie through the koi-processor claims engine only, not all the raw anchored content, sorry if that wasn't clear. If it's only testing data, then it's probably fine to ignore them. BTW testing data should have been anchored on testnet, not on mainnet. |
Github comments applied to changed text are simply marked as "outdated" so this should be easy to check comments, on the contrary, that still apply, although I can take a fresh pass when the changes are actually pushed.
I'm fine being a decider, I've already put some significant work into the reviewing process and claims engine is one of my main focus in fact. On another note, I added you to the collaborator of this repo, with write access to ease processes. Also, this ADR feels broader than this repo’s scope. If regen-data-standards is meant to define RDF schemas and vocabularies, then canonicalization, hashing, ledger anchoring, migration policy, KOI/JC engine parity, DoS guards, etc. seem like application/interoperability concerns rather than schema-standard concerns. |
|
Thank you for the write access, and for taking the decider role. I'll add you to the Deciders line now that you've said yes — that was the only thing I was waiting on. Taking your three points in order. The raw census. My mistake, and a useful one to have made: you asked a narrow question and I answered a chain-wide one. Scoped to the koi-processor claims engine, every raw anchor we have made is a test or demo fixture rather than a real claim, so yes, ignorable. I'm deliberately not quoting a count here: local and team-prod databases disagree and I haven't reconciled them. And you're right that they should have gone to testnet. Worth saying plainly: that was a deliberate team decision I went along with, not an accident — "mainnet is our testnet" was agreed on the 2026-03-10 call. It was the wrong call and it left permanent test records on mainnet. I'd rather we reverse it than repeat it. Outdated threads. Also right, and I gave you work GitHub already does. I'll check the outdated markers myself once the changes are pushed rather than asking you to audit your own review. Scope — and I think you may be right. This is the more interesting one, and it's a bigger challenge than the split I proposed, so I'd rather sit with it than reflexively defend. Your cut is: this repo defines RDF shapes (classes, predicates, ranges, cardinality, vocabularies, encoding conventions where they affect generated RDF), and canonicalization / hashing / anchoring / migration policy / engine parity / DoS guards are application-and-interop concerns that belong in a separate spec. Reading ADR 0001 against that line, most of it falls on your side of it. D2 (algorithm), D3 (lexical normalization), D4 (DoS cap), D5 (digest), D8 (migration) are all "how bytes are produced and put on a ledger" — none of them constrain what RDF we generate. The parts that genuinely are schema concerns are narrower: which fields are identity-bearing, whether That is close to the split I proposed but cut along a different axis, and yours is the better axis: I was splitting on settled vs open, which is a scheduling property. You're splitting on what kind of thing it is, which survives contact with time. One complication worth putting in front of you before either of us commits. I posted the split proposal yesterday claiming canonicalization is cleanly separable, then had to correct it (issue #60): JC's engine doesn't hash a single graph. Its preimage is four parts — canonical Claim IRI, content RDF dataset, assertor IRI, So: where would you want the identity protocol to live? A new repo, I'd suggest we don't merge anything on Monday until this is settled. The sign-off deadline was to stop the decision drifting, not to force it past a good question. |
|
@JeancarloBarrios @blushi, could you give us a read on these three seams before tomorrow's 9:30 AM PDT call, if possible? I'm not proposing another ADR or schema edit today. I want us to leave the call with a clear owner and an async ratification path, without reopening the whole design.
A short answer is enough. I can turn it into the concrete follow-up after the call. |
Marie's review: this repo is about data standards, not how the resulting data is used within regen ledger or koi-processor. Keeps the standards compatibility / provenance checklist and the ADR<->schema alignment rules; moves ledger constraints, the source tables and the worked example to koi-processor (gaiaaiagent/koi-processor#54). Also replaces the two repo-relative docs/adr links, which 404 on main because that file only exists on the regen-network#56 branch -- the unresolved Codex review comment from 2026-08-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
…ecklist Dave's 2026-08-04 process proposal asked for a small shared "agent context" file: non-negotiable design constraints, links to authoritative docs, and a bootstrap point for agent workflows that humans can read as a checklist. Ships with the ADR PR because the ADR is what violated the constraint: D2-b recommended JCS over a typed JSON projection, which MsgAttest can never accept. Every constraint carries a citation and a consequence. Constraints beyond the known MsgAttest/Raw pair, all verified in code: - v2 no longer validates algorithm identifiers on-chain. digest_algorithm, canonicalization_algorithm and merkle_tree are uint32, not enums; the validators reject only zero, and the strict membership checks DigestAlgorithm.Validate / GraphCanonicalizationAlgorithm.Validate have zero non-test call sites. A wrong algorithm identifier is accepted silently. Client-side correctness is the only control. - The IRI encodes the hash type (IriPrefixRaw=0 / IriPrefixGraph=1), so Raw->Graph migration mints a NEW IRI rather than re-labelling one. This is why ADR 0001 D8 dual-anchors instead of re-minting. - Graph IRIs must end .rdf; Raw extensions are 2-6 lowercase-or-numeric chars. Reading the IRI suffix is a zero-cost attestability check. - The gap is narrower than "nothing is attestable": claims anchor Raw and are not attestable, but attestation records already go through generate_graph_iri -> broadcast_attest on a graph IRI. - Live divergence: that deployed graph path canonicalizes with URDNA2015 via pyld while ADR 0001 D2-a specifies RDFC-1.0. Both stamp wire byte 1 and the chain does not check it, so a fork would be silent. pyld is also absent from koi-processor/requirements.txt. Ledger citations pinned to regen-ledger 451c3a3f, verified against the remote at that SHA; koi-processor citations to regen-prod c08c0a7e, the deployed branch (line numbers differ from feature branches). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bare 'make -C schema' runs only gen-taxonomy (the first target), not the full build. Verified against schema/Makefile: targets are gen-taxonomy, lint, gen-doc, gen-rdf, clean-rdf, update-graph, all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…assumption; pin all citations Addresses @blushi's 2026-08-03 review on regen-network#56. - Remove "Raised in review, and it is the stronger position." (D1) — process metadata, not decision rationale. - D2: RDFC-1.0 is now presented as the data module's own name for the algorithm, not a preference. GRAPH_CANONICALIZATION_ALGORITHM_RDFC_1_0 is the only non-zero value in the v2 canonicalization registry, and production targets v2 (derive_ledger_iri sends the v2-only file_extension field). Notes the v1 rename from URDNA2015 at the same wire value, and that the shared byte proves nothing about which spec ran. Adds the live divergence: the deployed graph path runs pyld URDNA2015 and stamps _GRAPH_CANON_URDNA2015 = 1. - Withdraw the claim that JC's work already assumes the lifecycle-record shape. Marie is closer to that codebase and reports no corresponding shapes exist. Booked as Open Question 7 (numbered 7 so refs to 1-6 stay valid), and the Context section now says the ADR does not describe JC's implementation. - D8 rewritten from dual anchoring to a straight cutover. Mainnet query found 8 Raw anchors from the claims engine, all test/demo, and zero executed MsgAttest. Method and bounds stated inline. Algorithm-column bullet promoted to D9, which stands on its own. - Every code citation pinned to a 40-hex commit SHA (27 links, all verified to resolve). No relative or branch links remain.
…e MsgAttest to D2 The previous revision's D8 rested on "there is no production Raw corpus on mainnet" and on the claim that a chain-wide Raw census is impossible from a query node. Both were false. A census over the tx index (2026-02-24 → 2026-08-05) returns 950 MsgAnchor: 418 Raw, 532 Graph. Ours is 8 of 418 (1.9%), 363 Raw anchors post-date our last, and the most recent is 2026-08-05. Raw is also the registry service account's live document path (399 anchors, 259 of them PDFs), so "the registry's data is already all-Graph" holds only for ecocredit metadata IRIs. D8 keeps the cutover and still drops the dual-anchor machinery, but now on corpus SIZE — re-anchoring eight throwaway fixtures is not a migration — rather than corpus absence. Every our-account statement is scoped explicitly to regen15eexs…, and exhaustiveness now rests on the sequence argument (sequence: 9, nine indexed txs = 8 MsgAnchor + 1 MsgSend) instead of a per-IRI sweep that could never prove it. That also resolves the previous UNVERIFIED note about an unexplained ninth transaction. D8 no longer asserts away what it cannot answer: dual-anchoring also buys forward interoperability during a transition, and nothing gathered speaks to whether any counterparty dereferences our Raw IRIs. Booked as Open Question 8, together with the two unattributed mainnet accounts anchoring Raw JSON in our shape whose 11 hashes match neither claims database. D2 gains the positive case it was missing: four MsgAttest exist on Regen mainnet (2026-07-13 ×2, 2026-07-19, 2026-08-03), all carrying canonicalization_algorithm: 1. Graph attestation with the algorithm D2-a selects is already working practice on this chain, which is a stronger argument than "we have nothing to lose." Also corrected: the two census transports do not both return 950 — REST returns 950 and Tendermint tx_search returns 815 against a node pruned to height 27420001, and exactly 815 of the 950 sit at or above that height. They reconcile on their overlap; the bounds section says so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on schema refs Three changes, all from reading ybird-labs/claims directly on 2026-09-08. Deciders: adds Marie Gauthier, per @glandua's 2026-08-10 request. Q2 is answered and the two halves come apart. Algorithm: JC does use RDFC-1.0 canonical N-Quads / UTF-8 / SHA-256 on both branches, arrived at independently of D2-a. Substance: not equivalent, and it differs between his own branches. main (8cbf0e5d) hashes Claim IRI + content + assertor + asserted_at; the unmerged design branch (217fafdd) hashes declared schema references + canonical content, derives the ClaimIRI from the fingerprint, and removes provenance from the claim value entirely. main is the merged default, so a casual reader gets the older model. Q9 (new) is the consequence: this ADR's substance set contains no schema declarations, so two implementations that both do RDFC-1.0 correctly still produce different fingerprints. That divergence is invisible at the algorithm layer and surfaces only as claims that fail to match -- the most likely source of a silent cross-engine fork, and covered nowhere else here. Which branch JC considers current is not established; asked by email today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
The crates/ tree is byte-identical between main (8cbf0e5d) and the design branch (217fafdd): same tree hash 1d556cc5, empty diff. AssertionProvenance and the four-part fingerprint are present on both. The new model lives in the design doc, AGENTS.md, README.md and spike/ -- a separate cargo package the root workspace excludes. That relocates the open item. It is not "which branch is current"; it is that the branch's own AGENTS.md forbids the four types its crate still defines, so merging it would not move the implementation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
Applies the five mechanical suggestions from the 2026-09-09 review: - D7 renumbered to D2, with the two dependent references updated (the hasCoBenefits row and the co-benefits question). D2-D6 moved to the implementation companion, so a lone D7 read as a gap. - The `claim_type` -> `hasClaimType` mapping sentence removed; it describes an external implementation and does not belong in the standards repo. - The four "moved to the implementation proposal" entries dropped (JC implementation, digest suite, raw-anchor migration, declared schema references). Recording what left the document is not the document's job. One consequence worth flagging rather than burying: the line above that list promises "historical question numbers are retained here so existing review links remain interpretable", and markdown renumbers an ordered list from its first item, so deleting entries 2, 6, 8 and 9 would have silently renumbered the survivors 1-5 and broken exactly that promise. The remaining numbers are now explicit labels, so the promise stays true. Say if you would rather drop the promise instead and keep a plain list. Marie's structural point -- an ADR with pending decisions is not an ADR -- is answered on the thread as a proposal, not applied here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YLntBrd61M5EUcngh2Lv8m
Marie's call on the thread: "yes let's drop the promise and keep a plain list." Removes the sentence promising that historical question numbers are retained so existing review links stay interpretable, and with it the Q1/Q3/Q4/Q5/Q7 labels that only existed to keep that promise true after four entries were deleted. The section heading loses "and review continuity" for the same reason -- it was naming the promise, not the content. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
3f2bf5d to
8691a3b
Compare
Ten of the review comments, implemented. The two open questions (hash wire format, whether PROV-O should block) and the ConsentDirective split (pending @glandua's sign-off) are deliberately NOT in this commit. :213 — claims/attestations become claimRefs/attestationRefs, multivalued uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined objects during the consent rewrite without re-justifying it. A Claim's identity and lifecycle (verificationStatus, contentHash, supersession) evolve after emission, so embedding one couples this envelope's lifecycle to the claim's. Claim/Attestation imports dropped — a uriorcurie range doesn't need them. :112 — observedLocalDate removed. It was a pure function of observedAt and sourceTimezone, so storing it only created something that could disagree with its own instant. The derivation is now stated normatively on sourceTimezone. :83 — snake_case convention on sourceSystem and recordType, enforced with a pattern rather than described in prose. This immediately caught the very inconsistency reported: the biocultural example's "field-signing-app" failed validation and is now "field_signing_app". :211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0. :236 — PROV-O, narrowly. Added the prov prefix and emittedAt -> prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent: wasGeneratedBy relates an Entity to an Activity rather than a string field, and prov:Agent is a class not a property, so those mappings would produce something that looks PROV-aligned and isn't. Real alignment means modelling the sensor run as a prov:Activity — proposed as its own ADR. docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56. docs:131 — Pydantic section removed; binding is the consumer's concern and belongs in the consuming repo. Fixed the header line that referenced it. Codex P2 (raw-data residency) — the descriptions on rawContentInline and rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No such rule exists; the only rule requires rawContentHash when processingState is COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so the constraint is not expressible on this class without hoisting a mirror field (two sources of truth for one policy — worse than an honest gap). Descriptions now say enforcement is external, a new doc section states the consumer obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml is committed as a negative fixture: a SOVEREIGN record carrying inline raw that passes validation. Kept so the gap stays visible and so the fixture starts failing if a future LinkML or SHACL shape ever expresses it. Verified: both positive examples validate against the edited schema; the negative fixture validates (which is the point); linkml-lint reports no new problems — the one error it emits ("id is not a uri") is pre-existing and repo-wide, present on unmodified OutputRecord.yaml and on Claim.yaml as merged on main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses the lean I gave her, so the reasoning is in the doc under "Hash wire format" and summarised here. I said I'd probably go bare hex, to match the implementation and the ledger. Three findings support that, and I re-derived all three: - koi-processor emits bare lowercase hex — ledger_anchor.py:96, `hashlib.blake2b(..., digest_size=32).hexdigest()`. - `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented the token in these examples. - The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto), never a prefixed string. What flipped it: - A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB. `koi_memories.content_hash` is SHA-256 (migrations/004:28) and `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same shape, same length, different algorithm, one codebase. - ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the missing discriminator as a latent defect in its own words: "the current schema cannot express which canonicalization produced a stored hash." Shipping an undiscriminated hash in a WIRE CONTRACT repeats that defect where it is hardest to fix — the ledger and the claims tables can add an adjacent column later because they own both sides of the read. A contract consumed by parties we don't control cannot. So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash. Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the ledger's DigestAlgorithm registry (today exactly one non-zero member, DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum gains one. It does NOT name the ContentHash kind or a canonicalization. For rawContentHash that is complete — a raw payload is not RDF, anchors as ContentHash.Raw, and the proto says Raw "does not specify a deterministic, canonical encoding". For recordContentHash it is NOT complete: this envelope does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same question ADR 0001 is deciding for Claim. Flagged in the slot description rather than quietly decided. Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks string equality and dedup on a value whose whole purpose is content-addressed identity. Rejected multihash/multibase: properly self-describing, but needs a codec table on both sides, produces values nothing in our stack emits, and diverges from the ledger's registry for no gain we can spend today. Cost, stated rather than hidden: koi-processor must prefix at the contract boundary. One line, and the right direction — the standard sets the wire format and the implementation adapts. Not touched: Claim.contentHash and Attestation.contentHash on main have the same gap (range: string, no pattern). Same class as Entity having no identifier slot; belongs in the identity-keys follow-up, not in a PR editing merged schemas. Verified: `make -C schema lint` clean; all three examples pass `linkml-validate -C OutputRecord`; and the pattern was tested against what it must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest all fail with "does not match". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Marie's review: this repo is about data standards, not how the resulting data is used within regen ledger or koi-processor. Keeps the standards compatibility / provenance checklist and the ADR<->schema alignment rules; moves ledger constraints, the source tables and the worked example to koi-processor (gaiaaiagent/koi-processor#54). Also replaces the two repo-relative docs/adr links, which 404 on main because that file only exists on the regen-network#56 branch -- the unresolved Codex review comment from 2026-08-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
Ten of the review comments, implemented. The two open questions (hash wire format, whether PROV-O should block) and the ConsentDirective split (pending @glandua's sign-off) are deliberately NOT in this commit. :213 — claims/attestations become claimRefs/attestationRefs, multivalued uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined objects during the consent rewrite without re-justifying it. A Claim's identity and lifecycle (verificationStatus, contentHash, supersession) evolve after emission, so embedding one couples this envelope's lifecycle to the claim's. Claim/Attestation imports dropped — a uriorcurie range doesn't need them. :112 — observedLocalDate removed. It was a pure function of observedAt and sourceTimezone, so storing it only created something that could disagree with its own instant. The derivation is now stated normatively on sourceTimezone. :83 — snake_case convention on sourceSystem and recordType, enforced with a pattern rather than described in prose. This immediately caught the very inconsistency reported: the biocultural example's "field-signing-app" failed validation and is now "field_signing_app". :211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0. :236 — PROV-O, narrowly. Added the prov prefix and emittedAt -> prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent: wasGeneratedBy relates an Entity to an Activity rather than a string field, and prov:Agent is a class not a property, so those mappings would produce something that looks PROV-aligned and isn't. Real alignment means modelling the sensor run as a prov:Activity — proposed as its own ADR. docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56. docs:131 — Pydantic section removed; binding is the consumer's concern and belongs in the consuming repo. Fixed the header line that referenced it. Codex P2 (raw-data residency) — the descriptions on rawContentInline and rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No such rule exists; the only rule requires rawContentHash when processingState is COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so the constraint is not expressible on this class without hoisting a mirror field (two sources of truth for one policy — worse than an honest gap). Descriptions now say enforcement is external, a new doc section states the consumer obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml is committed as a negative fixture: a SOVEREIGN record carrying inline raw that passes validation. Kept so the gap stays visible and so the fixture starts failing if a future LinkML or SHACL shape ever expresses it. Verified: both positive examples validate against the edited schema; the negative fixture validates (which is the point); linkml-lint reports no new problems — the one error it emits ("id is not a uri") is pre-existing and repo-wide, present on unmodified OutputRecord.yaml and on Claim.yaml as merged on main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses the lean I gave her, so the reasoning is in the doc under "Hash wire format" and summarised here. I said I'd probably go bare hex, to match the implementation and the ledger. Three findings support that, and I re-derived all three: - koi-processor emits bare lowercase hex — ledger_anchor.py:96, `hashlib.blake2b(..., digest_size=32).hexdigest()`. - `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented the token in these examples. - The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto), never a prefixed string. What flipped it: - A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB. `koi_memories.content_hash` is SHA-256 (migrations/004:28) and `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same shape, same length, different algorithm, one codebase. - ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the missing discriminator as a latent defect in its own words: "the current schema cannot express which canonicalization produced a stored hash." Shipping an undiscriminated hash in a WIRE CONTRACT repeats that defect where it is hardest to fix — the ledger and the claims tables can add an adjacent column later because they own both sides of the read. A contract consumed by parties we don't control cannot. So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash. Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the ledger's DigestAlgorithm registry (today exactly one non-zero member, DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum gains one. It does NOT name the ContentHash kind or a canonicalization. For rawContentHash that is complete — a raw payload is not RDF, anchors as ContentHash.Raw, and the proto says Raw "does not specify a deterministic, canonical encoding". For recordContentHash it is NOT complete: this envelope does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same question ADR 0001 is deciding for Claim. Flagged in the slot description rather than quietly decided. Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks string equality and dedup on a value whose whole purpose is content-addressed identity. Rejected multihash/multibase: properly self-describing, but needs a codec table on both sides, produces values nothing in our stack emits, and diverges from the ledger's registry for no gain we can spend today. Cost, stated rather than hidden: koi-processor must prefix at the contract boundary. One line, and the right direction — the standard sets the wire format and the implementation adapts. Not touched: Claim.contentHash and Attestation.contentHash on main have the same gap (range: string, no pattern). Same class as Entity having no identifier slot; belongs in the identity-keys follow-up, not in a PR editing merged schemas. Verified: `make -C schema lint` clean; all three examples pass `linkml-validate -C OutputRecord`; and the pattern was tested against what it must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest all fail with "does not match". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Carries forward the July 2026 agreement to remove verificationStatus, contentHash and dataIri from Claim.yaml, as agreed rather than proposed. Places impact, co-benefit, quantity, credit-class and methodology slots in specialized claim schemas, with co-benefit collection semantics decided in that scope. Namespace leaves the open decisions and methodology closes only conditionally. Claim period, hasOperator and the PROV-O alignment (input: regen-network#82) stay open. Status stays Proposed.
|
@blushi, this push implements your #70 corrections, as committed in my follow-up.
Status stays Proposed. A 2026-09-23 revision-history entry records the change. |
* docs: land standards-agent-context.md on main Adds the shared agent-context / non-negotiables file proposed as protocol 4 of the 2026-08-04 standards-workflow discussion. It currently exists only on the ADR 0001 branch, which means the file meant to be the shared checklist is invisible to everyone who would use it. Landing it independently so it is available before the next piece of work starts, rather than arriving with an ADR. Contents are unchanged from the ADR branch: the ledger non-negotiables (MsgAttest accepts only ContentHash.Graph; Raw disclaims canonical encoding), where the authoritative sources live, a pre-PR checklist, the review tag vocabulary, and a worked example of the error the file exists to prevent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: narrow standards-agent-context to schema/ADR scope per review Marie's review: this repo is about data standards, not how the resulting data is used within regen ledger or koi-processor. Keeps the standards compatibility / provenance checklist and the ADR<->schema alignment rules; moves ledger constraints, the source tables and the worked example to koi-processor (gaiaaiagent/koi-processor#54). Also replaces the two repo-relative docs/adr links, which 404 on main because that file only exists on the #56 branch -- the unresolved Codex review comment from 2026-08-13. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW * docs: finish standards context scope review * docs: point at standards-agent-context.md from the root README Marie's review: the doc covers more than schemas, so a schema/README pointer would file the whole thing under the narrowest of its four sections. Sections 1 and 2 are schema-specific, but section 3 is a repo-level source index and section 4 (review tag vocabulary, author response convention) governs any PR in this repo. Root README it is. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(agent-context): say explicitly that the review-tag vocabulary covers only two cases Addresses Marie's 2026-09-14 review note on §4: the table lists two tags and did not say that this is the whole vocabulary on purpose. Now states that it covers only the two failure classes that have cost review time here, that untagged comments are fine, and that new tags are added by PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: rename standards-agent-context.md to AGENTS.md AGENTS.md is a cross-vendor convention rather than one vendor's filename, so the objection that blocked this rename on 2026-09-10 does not hold. Claude Code reads it from the repository root from v2.1.277, and it is the name Cursor and GitHub Copilot already read, so the file is now picked up without anyone remembering to point at it. This repository has no CLAUDE.md, which is the condition for Claude Code to auto-load AGENTS.md. The scope note now states which sections are schema and ADR specific (1 to 3) and which apply to every pull request here (4), since a file at this name will be read for changes of any kind. The root README pointer from b6abb27 is kept and repointed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ic Claim hasClaimant is asserted, in-content attribution of whoever takes responsibility for the assertion (review on regen-network#56, regen-network#82 section 3.1). Its PROV subproperty declaration stays with the PROV-O alignment. hasOperator leaves the generic Claim: the operator is described on the relevant domain activity in the specialized claim schema, as responsibility for that activity rather than attribution of the Claim (regen-network#82 section 3.2). Claimant and operator leave the open decisions. Also drops the Context sentence describing earlier versions of the ADR, which the revision history already records.
|
Proposes the base RDF shape and provenance boundaries for Claim, incorporating the current Claim, LinkML and attestation stack. The ADR now gives explicit proposed choices for claimant attribution, typed subjects, assertedAt, claim period, content revision links, activity ownership and specialization, with links to the relevant WP0 and WP1 decisions.
The ADR remains Proposed. Field placement in the schema and adoption of the recommendations remain subject to the team’s review; this PR changes documentation only.
Validation: Markdown fence/link checks, git diff --check, and an independent review against the current schema stack. The old canonicalization-focused filename and README link are replaced with docs/adr/0001-claim-rdf-shape-and-provenance-boundaries.md.