Skip to content

docs: propose Claim RDF shape and provenance boundaries - #56

Open
DarrenZal wants to merge 15 commits into
regen-network:mainfrom
DarrenZal:adr/claim-substance-canonicalization
Open

DarrenZal wants to merge 15 commits into
regen-network:mainfrom
DarrenZal:adr/claim-substance-canonicalization

Conversation

@DarrenZal

@DarrenZal DarrenZal commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

Proposes the base RDF shape and provenance boundaries for Claim, incorporating the current Claim, LinkML and attestation stack. The ADR now gives explicit proposed choices for claimant attribution, typed subjects, assertedAt, claim period, content revision links, activity ownership and specialization, with links to the relevant WP0 and WP1 decisions.

The ADR remains Proposed. Field placement in the schema and adoption of the recommendations remain subject to the team’s review; this PR changes documentation only.

Validation: Markdown fence/link checks, git diff --check, and an independent review against the current schema stack. The old canonicalization-focused filename and README link are replaced with docs/adr/0001-claim-rdf-shape-and-provenance-boundaries.md.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces ADR 0001, which proposes a standard for defining a Claim's substance and canonicalizing it into a content-addressed identity (fingerprint) using JCS (RFC 8785) or RDFC-1.0, along with an updated ADR index. The review feedback highlights several technical areas for refinement: ensuring determinism for multivalued array fields by specifying a sorting requirement, avoiding timezone-shifting bugs by keeping pure date fields in YYYY-MM-DD format, clarifying how floating-point precision for quantity is handled under JCS, and resolving the logical tension between including usesMethodology in the substance definition while simultaneously questioning if its changes should mint a new identity.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 268e072ed0

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
@DarrenZal

Copy link
Copy Markdown
Contributor Author

@blushi — short version: you're right on every substantive point, and on D2 you were right for a bigger reason than either of us had.

I went looking for a rebuttal to your D2 comment and found the argument that kills D2-b instead. MsgAttest.content_hashes is typed repeated ContentHash.Graph, not ContentHash — so a claim anchored as ContentHash.Raw can never be attested on-chain. The attestation layer is the next thing we build, so D2-b doesn't just sit off the semantic path, it forecloses it. I'm withdrawing D2-b and revising the ADR to recommend D2-a.

Two corrections to things I believed going in: what production does today is deterministic json.dumps, not RFC 8785 JCS, so I can't claim JCS is the incumbent; and ContentHash.Raw's own doc comment explicitly disclaims canonical encoding, which means our current anchoring is weaker than we've been treating it, not that it justified D2-b.

The one place I'd still shape the outcome is migration: I don't think this forces re-minting every existing hash. Proposing a dual-anchor path in the revised ADR — existing Raw anchors stay as historical payload-integrity proofs, new claims get a Graph/RDFC-1.0 semantic fingerprint, and we add algorithm/version columns so each claim records which scheme minted it.

Also going in off your review: D4's DoS guard becomes mandatory rather than conditional (and it matters more now that RDFC-1.0 is actually on the serving path), the spike gets published to regen-network/koi-research so the evidence is citable, and I'll settle the set-vs-sequence declaration for hasCoBenefits, which the schema currently leaves open regardless of algorithm.

Still worth a call with @JeancarloBarrios, but the question is now narrower and better: not "which algorithm", but what RDF encoding profile we both freeze so the projections agree. I'd like to get that scheduled this week — the Second Muse engagement starts mid-August and involves standing up a claims-engine instance, so this gets exercised by a real client soon and it's expensive to reverse afterward.

DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Jul 31, 2026
Ten of the review comments, implemented. The two open questions (hash wire
format, whether PROV-O should block) and the ConsentDirective split (pending
@glandua's sign-off) are deliberately NOT in this commit.

:213 — claims/attestations become claimRefs/attestationRefs, multivalued
  uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this
  contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined
  objects during the consent rewrite without re-justifying it. A Claim's identity
  and lifecycle (verificationStatus, contentHash, supersession) evolve after
  emission, so embedding one couples this envelope's lifecycle to the claim's.
  Claim/Attestation imports dropped — a uriorcurie range doesn't need them.

:112 — observedLocalDate removed. It was a pure function of observedAt and
  sourceTimezone, so storing it only created something that could disagree with
  its own instant. The derivation is now stated normatively on sourceTimezone.

:83  — snake_case convention on sourceSystem and recordType, enforced with a
  pattern rather than described in prose. This immediately caught the very
  inconsistency reported: the biocultural example's "field-signing-app" failed
  validation and is now "field_signing_app".

:211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0.

:236 — PROV-O, narrowly. Added the prov prefix and emittedAt ->
  prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing
  sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent:
  wasGeneratedBy relates an Entity to an Activity rather than a string field, and
  prov:Agent is a class not a property, so those mappings would produce something
  that looks PROV-aligned and isn't. Real alignment means modelling the sensor run
  as a prov:Activity — proposed as its own ADR.

docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56.
docs:131 — Pydantic section removed; binding is the consumer's concern and
  belongs in the consuming repo. Fixed the header line that referenced it.

Codex P2 (raw-data residency) — the descriptions on rawContentInline and
  rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No
  such rule exists; the only rule requires rawContentHash when processingState is
  COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so
  the constraint is not expressible on this class without hoisting a mirror field
  (two sources of truth for one policy — worse than an honest gap). Descriptions
  now say enforcement is external, a new doc section states the consumer
  obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml
  is committed as a negative fixture: a SOVEREIGN record carrying inline raw that
  passes validation. Kept so the gap stays visible and so the fixture starts
  failing if a future LinkML or SHACL shape ever expresses it.

Verified: both positive examples validate against the edited schema; the negative
fixture validates (which is the point); linkml-lint reports no new problems — the
one error it emits ("id is not a uri") is pre-existing and repo-wide, present on
unmodified OutputRecord.yaml and on Claim.yaml as merged on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@DarrenZal

Copy link
Copy Markdown
Contributor Author

Pushed 0b1d27a — the ADR is rewritten to D2-a, plus the changes that came out of the rest of your review.

D2 reversed. You argued for D2-a because ContentHash.Graph declares only RDFC-1.0. Chasing that premise turned up a stronger reason neither of us had named: MsgAttest.content_hashes is typed repeated ContentHash.Graph, not ContentHash, and the CLI enforces it again. We anchor ContentHash.Raw today — so claims as we anchor them cannot be attested on-chain at all. Attestation is the next layer we build, so D2-b wasn't just off the semantic path, it foreclosed one we're actively heading for.

I've recorded the two arguments of mine that failed, rather than quietly dropping them, so the reversal is auditable:

  • "JCS is effectively the incumbent" — false. Production runs json.dumps(sort_keys=True, ensure_ascii=True, ...): deterministic, but not RFC 8785 (JCS wants UTF-8 with minimal escaping, UTF-16 code-unit key order, ECMAScript number serialization). And it hashes claim_rid and entity_uri, so it isn't even a clean substance projection.
  • ContentHash.Raw's own doc comment disclaims deterministic canonical encoding. I'd read it as permission; it's the ledger declining to make the guarantee.

Your :70 point, taken further than agreement. You said the mutable fields should leave Claim.yaml altogether rather than just be excluded from substance. That's right, and it's free: koi-processor already keeps verification in its own column with an insert-only claim_state_log, and ledger_anchor.py already omits it and content_hash from the hash. The schema is out of step with the implementation, not the reverse. D1 now says remove, not exclude — an excluded-but-declared slot means every implementation has to remember to exclude it and nothing catches the one that forgets. It also shrinks the sh:closed surface blocking koi-processor #30.

Your :101 point produced the better version of D6: the fingerprint is an index over the claim, not a property of it. Canonicalization becomes total, with no exclusion list to get wrong.

Two new sections from things your review implied:

  • D7 — collection ordering. hasCoBenefits is multivalued + inlined_as_list with no ordering declaration, while other schemas here use list_elements_ordered. Under any algorithm set-vs-sequence hashes differently, so it's now declared: substance collections are sets, sorted on canonical id.
  • D8 — dual-anchor migration. Going D2-a does not mean re-minting every existing hash. Existing Raw anchors stay as historical payload-integrity proofs; new claims mint Graph/RDFC-1.0 identities; and we add algorithm/profile columns, which 064 and 070 lack today — the current schema literally cannot express which canonicalization produced a stored hash.

The spike is published (koi-research#2) and the ADR now links permalinks instead of the scratch path you couldn't open. Its contemporaneous write-up still argues for D2-b; I left it as-run rather than back-editing, since rewriting evidence after the conclusion changed is what makes a spike untrustworthy. Findings 1 and 3 are algorithm-independent; finding 2 gets more load-bearing under D2-a, which is why D4 is now a hard release gate rather than a recommendation.

What's left is narrower and it's @JeancarloBarrios's call. D2-a makes the RDF encoding profile load-bearing — namespace, claim_type as literal or IRI, compact vs expanded IRIs. That's the remaining place a false fork can hide, and it's the one question the code can't answer for us. Second question for him: does his engine already produce RDFC-1.0 canonical N-Quads over an equivalent substance set? If so the adapter is thin.

Still open from you: the b2s256: vs bare-hex wire format on #55, and whether PROV-O should block that PR.

@blushi

blushi commented Aug 3, 2026

Copy link
Copy Markdown
Member

The one place I'd still shape the outcome is migration: I don't think this forces re-minting every existing hash. Proposing a dual-anchor path in the revised ADR — existing Raw anchors stay as historical payload-integrity proofs, new claims get a Graph/RDFC-1.0 semantic fingerprint, and we add algorithm/version columns so each claim records which scheme minted it.

What kind of data was anchored as raw already on mainnet?

Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
@glandua
glandua requested a review from blushi August 10, 2026 19:32
@glandua

glandua commented Aug 10, 2026

Copy link
Copy Markdown
Collaborator

Adding @blushi as a ratifier on this one.

The Deciders line currently reads Darren + Shawn + JC. The D2 reversal on 31 July turned on the MsgAttest / ContentHash.Graph constraint, and that came out of Marie's review. Someone whose objection changed the central decision of a decision record should be on the list that ratifies it.

@DarrenZal — could you add her to Deciders next time you touch the file? I've requested her review here in the meantime.

Two notes for the deciding call:

  • JC's two questions are the gate. D2-a makes the RDF encoding profile load-bearing, and open feat: update design #1 and STANDARDS-3 Initial set of terms, definitions, and a dummy icon  #2 are both addressed to @JeancarloBarrios. The ADR also reasons in places from what the ybird kernel already assumes, and Marie has flagged at least one place where that reading may not hold. Better to have him in before merge than to discover it after.
  • One substance-scope question may still be open. D1 excludes verificationStatus, contentHash, dataIri and supersedes from the projection, with clear reasoning. Marie's position in review was stronger — that they shouldn't be fields on Claim at all. Whether "excluded from substance" is sufficient, or mutable lifecycle state needs to sit on a separate object, is a real fork and I don't think it's been settled either way. Worth deciding explicitly rather than carrying it along with D2.

Separately, and not a nitpick: the reversal write-up is the right standard for these. Quoting the proto, naming the constraint neither option had been checked against, and recording the two corrections to the prior reasoning is what makes this ratifiable rather than merely decided. Also worth noting that the PR description still describes D2-b as the recommendation — anyone reading the summary rather than the file will get the wrong picture, so it's worth a quick edit.

Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Aug 11, 2026
…assumption; pin all citations

Addresses @blushi's 2026-08-03 review on regen-network#56.

- Remove "Raised in review, and it is the stronger position." (D1) — process
  metadata, not decision rationale.
- D2: RDFC-1.0 is now presented as the data module's own name for the algorithm,
  not a preference. GRAPH_CANONICALIZATION_ALGORITHM_RDFC_1_0 is the only
  non-zero value in the v2 canonicalization registry, and production targets v2
  (derive_ledger_iri sends the v2-only file_extension field). Notes the v1
  rename from URDNA2015 at the same wire value, and that the shared byte proves
  nothing about which spec ran. Adds the live divergence: the deployed graph
  path runs pyld URDNA2015 and stamps _GRAPH_CANON_URDNA2015 = 1.
- Withdraw the claim that JC's work already assumes the lifecycle-record shape.
  Marie is closer to that codebase and reports no corresponding shapes exist.
  Booked as Open Question 7 (numbered 7 so refs to 1-6 stay valid), and the
  Context section now says the ADR does not describe JC's implementation.
- D8 rewritten from dual anchoring to a straight cutover. Mainnet query found
  8 Raw anchors from the claims engine, all test/demo, and zero executed
  MsgAttest. Method and bounds stated inline. Algorithm-column bullet promoted
  to D9, which stands on its own.
- Every code citation pinned to a 40-hex commit SHA (27 links, all verified to
  resolve). No relative or branch links remain.
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Aug 11, 2026
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses
the lean I gave her, so the reasoning is in the doc under "Hash wire format" and
summarised here.

I said I'd probably go bare hex, to match the implementation and the ledger.
Three findings support that, and I re-derived all three:

- koi-processor emits bare lowercase hex — ledger_anchor.py:96,
  `hashlib.blake2b(..., digest_size=32).hexdigest()`.
- `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented
  the token in these examples.
- The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are
  `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto),
  never a prefixed string.

What flipped it:

- A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB.
  `koi_memories.content_hash` is SHA-256 (migrations/004:28) and
  `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same
  shape, same length, different algorithm, one codebase.
- ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the
  missing discriminator as a latent defect in its own words: "the current schema
  cannot express which canonicalization produced a stored hash." Shipping an
  undiscriminated hash in a WIRE CONTRACT repeats that defect where it is
  hardest to fix — the ledger and the claims tables can add an adjacent column
  later because they own both sides of the read. A contract consumed by parties
  we don't control cannot.

So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash.

Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the
ledger's DigestAlgorithm registry (today exactly one non-zero member,
DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum
gains one. It does NOT name the ContentHash kind or a canonicalization. For
rawContentHash that is complete — a raw payload is not RDF, anchors as
ContentHash.Raw, and the proto says Raw "does not specify a deterministic,
canonical encoding". For recordContentHash it is NOT complete: this envelope
does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same
question ADR 0001 is deciding for Claim. Flagged in the slot description rather
than quietly decided.

Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is
lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks
string equality and dedup on a value whose whole purpose is content-addressed
identity.

Rejected multihash/multibase: properly self-describing, but needs a codec table
on both sides, produces values nothing in our stack emits, and diverges from the
ledger's registry for no gain we can spend today.

Cost, stated rather than hidden: koi-processor must prefix at the contract
boundary. One line, and the right direction — the standard sets the wire format
and the implementation adapts.

Not touched: Claim.contentHash and Attestation.contentHash on main have the same
gap (range: string, no pattern). Same class as Entity having no identifier slot;
belongs in the identity-keys follow-up, not in a PR editing merged schemas.

Verified: `make -C schema lint` clean; all three examples pass
`linkml-validate -C OutputRecord`; and the pattern was tested against what it
must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest
all fail with "does not match".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@DarrenZal

Copy link
Copy Markdown
Contributor Author

1. What kind of data was anchored as raw on mainnet

You asked about the chain, so I ran a chain-wide census rather than answering only for us. Over the tx index, 2026-02-24 → 2026-08-05: 950 MsgAnchor — 418 Raw, 532 Graph.

Chain-wide, Raw is mostly documents and images, not JSON: pdf 259, jpeg 58, jpg 50, png 24, json 19, webp 5, heic 2, bin 1. And it is concentrated in one account:

Raw Sender
399 regen1ugps77ux3l9cqe7g3pjdqfs6lgz2uq7y5ky2xl also sends 528 of the 532 Graph anchors — the registry service account
10 regen1wkc7x2465hhmr57fc300yxu84p78u3ystue5p5 all json
8 regen15eexs5vt9klzf304v2fczfh2823lwgz8g4apt9 the claims engine — ours
1 regen1jfheyvsah5wqfyawmedme43te056z8gzdnpf3j json

So the registry account runs both paths side by side: Graph for ecocredit metadata IRIs, Raw for what looks like document attachments — 259 PDFs. I had been about to write you "the registry's own data is already all-Graph." That is true of the metadata IRIs and misleading about the account, so I've dropped it.

Ours is 8 of 418 — 1.9%, and all eight are test or demo JSON: 4 claims (2026-03-11 ×2, 03-30, 04-23) and 4 attestation blobs (2026-03-11), every one "raw": {…,"file_extension":"json"}, "graph": null, every hash matching the prod DB byte for byte. Two are explicit fixtures — claimant orn:personal-koi.entity:claims-engine-test-org (a local dev KOI namespace, not a registry entity), statements reading "Test organization restored 50 hectares… (run 1773215949)". One demo org cec_demo_001, one seed sample about Regen itself. Zero third-party ecological claims. We have also never executed a MsgAttest — AttestationsByAttestor is total: 0 for both our keys, and the laptop key has pub_key: null / sequence: 0, so it has never signed anything.

That eight is exhaustive, by a cleaner argument than the per-IRI sweep I'd have offered otherwise: the prod account reports sequence: 9 — lifetime signed transactions — and a tx-index query for that sender returns exactly 9: 8 × MsgAnchor + 1 × MsgSend. Nine signed, nine found, eight anchors. (That also resolves the unexplained ninth transaction I would have flagged as UNVERIFIED. It was the MsgSend.)

Scope and method, since a "not found" means nothing unless the query is known to work. REST cosmos/tx/v1beta1/txs filtered on message.action='/regen.data.v2.MsgAnchor' with pagination.count_total=true → 950, against a node whose index reaches back to height 25498036 (2026-02-09). Tendermint tx_search against a node pruned to 27420001 returns 815 — and exactly 815 of the 950 sit at or above that height, so the two transports agree on their overlap. This is a lower bound, not the whole chain: anything anchored before 2026-02-09 is outside both indexes. Claim-level detail (statements, hashes) is AnchorByIRI per stored IRI, validated first against a positive control — a real C01 batch IRI that returns a full record.

One correction to something I nearly told you: x/data v2 having no list RPC does not mean a chain-wide Raw census is impossible. The no-list-RPC part is true — all 11 query RPCs are keyed lookups — but the tx index settles it, and you probably already knew that.

2. Two accounts I can't attribute — do you know whose they are?

regen1wkc7x2465hhmr57fc300yxu84p78u3ystue5p5 (10 Raw JSON) and regen1jfheyvsah5wqfyawmedme43te056z8gzdnpf3j (1 Raw JSON, 4 Graph, and all four mainnet MsgAttest) are anchoring raw JSON in roughly our shape. None of their 11 hashes matches either of our claim databases, so they are not ours. JC / ybird-labs would be my guess and I'm deliberately not asserting it. If you know whose keys those are, it bears directly on §4.

3. The strongest argument for D2-a turns out to be on-chain already

Same census: there are four MsgAttest on Regen mainnet — 2026-07-13 ×2, 2026-07-19, and 2026-08-03 — all from regen1jfhey…, and every one carries canonicalization_algorithm: 1. RDFC-1.0, the algorithm D2-a selects, three days ago.

So the case for D2-a is not "we have nothing to lose by switching." It is that Graph attestation with this exact algorithm is live working practice on this chain, and every claim our engine has anchored is structurally excluded from it. That is now the lead evidence in D2. Your point from a few months ago — if we want data attested it shouldn't be anchored as raw JSON — is the thing the chain is already doing.

The v2 enum argument stands alongside it, with the v1 rename spelled out: GRAPH_CANONICALIZATION_ALGORITHM_URDNA2015 = 1 (v1) and ..._RDFC_1_0 = 1 (v2) are the same enum slot, so the shared byte 1 is a trap rather than a reassurance — two implementations can stamp it and still emit different N-Quads, and v2 doesn't check (ContentHash_Graph.Validate rejects only == 0; the field is a plain uint32). We already have that divergence live: generate_graph_iri runs pyld URDNA2015 and _content_hash_graph_to_iri stamps _GRAPH_CANON_URDNA2015 = 1. That path serves attestation records rather than claims, but adopting D2-a means reconciling it — follow-up (e).

4. D8 still collapses, for a different reason than I first wrote

I'm dropping the dual-anchor machinery, but the honest reason is corpus size, not corpus absence: re-anchoring eight throwaway fixtures is not a migration. My first pass at this justified it with "there is no production Raw corpus on mainnet." That was wrong by roughly a factor of 52 and it's withdrawn — 418 Raw anchors, 363 of them made after our last one, most recently 2026-08-05.

What I can't answer, and D8 now says so rather than assuming it away: dual-anchoring also buys forward interoperability during a transition. "Our corpus is small" answers is re-minting expensive; it does not answer does anyone else resolve our Raw IRIs. Hence §2. Booked as Open Question 8 rather than decided.

The one bullet inside old-D8 worth keeping is now D9: 064_claims_engine.sql and 070_data_iri.sql have no algorithm/profile/version column — grep -inE 'algorithm|profile|canonical|rdfc|urdna|jcs' over both returns exactly one hit, a comment on line 41 of 064. That is about telling future schemes apart, so it stands on its own once the migration framing goes.

5. The rest of your review

  • JC — "given he's already assuming this shape" was an assertion I had no basis for; withdrawn. The Context section now states that this ADR does not describe JC's implementation, and the companion-record question is Open Question 7, to settle with him rather than about him.
  • Line 76 — "Raised in review, and it is the stronger position." removed; the paragraph reads fine without it.
  • Links — schema/src/Claim.yaml pinned to 0a4ba12a, and I applied the point across the document rather than to that one line: 28 pinned permalinks, zero relative, zero branch, every one resolved through the contents API. There's a "Pinned sources" line under the header naming the three commits so bare backtick references still resolve.

On the read-through ask

Fair, and I'd rather give you a mechanism than a promise. standards-agent-context.md lands with this branch at the repo root: §1 is the ledger constraints with pinned citations, §3 is a pre-PR checklist where every item is mechanically checkable, §6 walks the D2-b failure as the worked example. §1 and §2 above are what came out of running an adversarial pass over my own draft before posting — that pass is what caught the "no Raw corpus" line, and it's the part of this I'd point at rather than the checklist. Tell me what it's still missing against what you actually keep having to catch.


Posted late — this was written on 6 August and then sat unsent while I was away 7–10 August. I've now pushed the branch: 0b1d27a..b3c8a3f, which adds the chain-wide anchor census behind D8, re-grounds D2 on the v2 enum, withdraws the JC assumption, pins the citations, and adds standards-agent-context.md. Everything you reviewed this morning was against the older commit, for which I'm sorry.

Your four comments from today land after all of the above — I'm reading them now and will answer each one directly, including the hasClaimant / AssertionProvenance placement question, which looks like the substantive one.

@DarrenZal

DarrenZal commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor Author

@LinuxIsCool @JeancarloBarrios @glandua — putting this to async sign-off rather than booking another call.

Straight about why: Gregory offered Thursday or Friday afternoon on 2026-08-05, JC said both worked, and I never sent the invite. That one's mine. Rather than chase a slot now, I think a written record is the better artifact for a decision like this anyway — but the week of silence wasn't a scheduling problem, it was me.

What changed, and why

D2-b (JCS / RFC 8785) is withdrawn. The forcing constraint is on-chain, not aesthetic:

MsgAttest.content_hashes is typed repeated ContentHash.Graph — identically in both regen/data/v1/tx.proto:88-91 and v2/tx.proto:93-96, and the proto comment says why in so many words:

// content_hashes are the content hashes for anchored data. Only RDF graph
// data can be signed as its data model is intended to specifically convey
// semantic meaning.
repeated ContentHash.Graph content_hashes = 2;

Our production anchoring emits ContentHash.Raw. The exact lines, on the public repo so you can open them directly — gaiaaiagent/koi-processor api/ledger_anchor.py L196-202 @ eaf1f77:

hash_json = {
    "raw": {
        "hash": hash_b64,
        "digest_algorithm": 1,  # DIGEST_ALGORITHM_BLAKE2B_256
        "file_extension": "json",  # no leading dot
    }
}

…over a canonical form that is json.dumps(obj, sort_keys=True, ensure_ascii=True, separators=(',',':')) with BLAKE2b-256. Claims as we anchor them today can never be attested on-chain. JCS would have locked that in.

To be precise about what "settled" means here: D2 is not settled because the spike succeeded — the spike self-reports PARTIAL FAILURE, and its 5-field substance set is genuinely defective (two ledger-distinct impact claims collapsed to one fingerprint). D2 is settled because the ContentHash.Graph type removes JCS from the option set regardless of anyone's preference. The spike's failures are inputs to D1 and to the hard gates, not to D2.

Independent corroboration, from reading JC's engine rather than reasoning about it: crates/claims/src/domain/fingerprint.rs pins the suite as ClaimValueRdfc10CanonicalNQuadsUtf8Sha256V1 — RDFC-1.0 over canonical N-Quads. Two engines converging on the same algorithm from opposite directions is about as good as this evidence gets, and it gives us a concrete cross-engine parity target rather than a hypothetical one.

Proposal: split this ADR in two

Right now this document is partly decided and partly blocked on questions only JC can answer, which makes it un-mergeable in either direction. I'd like to split it:

  • ADR 0001 narrows to the D2 / RDFC-1.0 canonicalization decision and merges as Accepted. This is the part that's actually forced, and it's what "claims engine v1.0 architectural clarity" needs to mean for the landscape-fund engagement.
  • A second ADR carries the claimant/provenance placement and open questions 1/2/7 and stays Proposed for as long as it needs to.

The seam is clean and nothing else moves. The narrowed 0001 still cites the spike evidence and the MsgAttest constraint — those are why D2 was decided, not open questions.

Marie — this probably invalidates part of your review, since your threads were written against the broad document. Rather than assume they carry over, could you say which still apply to the narrowed text? I'd expect a fresh pass on it.

Two things that gate "Accepted"

  1. The evidence isn't durable yet. This ADR cites the canonicalization spike via pinned permalinks, but that spike lives only on unmerged regen-network/koi-research PR STANDARDS-3 Initial set of terms, definitions, and a dummy icon  #2 — open 12 days, no reviewer ever assigned. An Accepted ADR whose evidence base is an unmerged branch isn't independently reviewable. Can someone review it, or give me merge rights?
  2. D1 is contested and I now think the objection is right. See my replies to Marie's threads — JC's implementation separates ClaimContent from AssertionProvenance, and D1 currently puts hasClaimant inside the substance graph. That belongs in the second ADR, not in the one we accept.

The decision rule I'm proposing

Deciders are Darren, Shawn, JC per the ADR's own Deciders line. Please approve or object in-thread by Monday 2026-08-17.

If either Shawn or JC is silent at that deadline, I'm proposing Gregory decides as sponsor — he owns the contract commitment this gates.

To be explicit: this rule is itself part of what's being signed off, and objecting to the process is a valid response. Nobody has authorised me to make Gregory the tiebreaker — he's the sponsor and owns the commitment, which is an argument, not authority. If you'd rather we did this differently, say so and we will. I don't want silence later characterised as consent nobody gave.

One governance question, asked not assumed

@glandua asked that @blushi be added to the Deciders line. She has done effectively all of the review on this ADR, so the substance of the request is obviously right — but Gregory isn't himself a decider, and Marie hasn't been asked whether she wants the role or its ongoing workload. So: Shawn, JC — do you agree? Marie — do you want it?

I'd rather ask than edit the line and call it done.

@DarrenZal

Copy link
Copy Markdown
Contributor Author

Re-stating my replies above in the format @davefortson proposed on 2026-08-04, because I answered the threads before applying the protocol and the threads are harder to scan without it. One correction to my own framing first: only the PR description has actually changed so far. The ADR file has not. So most of these are Deferred with a date, not Addressed — I'd rather say that than claim credit for edits that don't exist yet.

Thread Tag Disposition
r3759575656 — hasClaimant placement [CONSISTENCY] Conceded, deferred. ADR ⇄ JC's implementation mismatch is real: AssertionProvenance is a sibling of ClaimContent, so claimant is committed to but not in the canonicalized graph. D1 is wrong as written. Moves to the second ADR — before Mon 2026-08-17.
r3758768380 — JSON-LD expansion / namespace [CONSISTENCY] Deferred, same window. Open Question 1 narrows to a profile-version pin; claim_type becomes rfs:hasClaimType explicitly.
r3758092612 — already-answered question still listed [DOC-ACCESS] Deferred, same window. Removing the stale entry.
r3758938825 — describes something not built [DOC-ACCESS] Deferred, same window. Re-casting present tense as prospective.
PR description [DOC-ACCESS] Addressed. Was still recommending D2-b/JCS — the opposite of the file — since Gregory flagged it on 08-10. Now states D2-a, marks D1 contested, and replaces the dead local path RegenAI/scratch/jc-koi-canon-spike/RESULT.md with the koi-research PR #2 pointer.

Two [DOC-ACCESS] items I'm raising against myself, since both are instances of the exact problem the protocol exists to stop:

  • standards-agent-context.md is stranded. I offered it to Marie as the mechanism answering her read-through ask, but it exists only on this PR's branch — it is not on regen-network/regen-data-standards main. Protocol 4 says keep it in the repo; right now nobody can bootstrap from it or use it as a checklist. It should land independently of this ADR rather than riding on it.
  • This ADR's evidence base is on an unmerged branch. The canonicalization spike is regen-network/koi-research PR STANDARDS-3 Initial set of terms, definitions, and a dummy icon  #2, open 12 days with no reviewer ever assigned. An Accepted ADR citing evidence nobody can reach from main is the same failure in a different repo.

The split itself now has a design-intention issue per protocol 5 — filed late, which is the wrong order: #60. Push back there rather than here if the seam looks wrong.

@DarrenZal

Copy link
Copy Markdown
Contributor Author

One gap in the sign-off rule above that I should have caught when I wrote it: I can't merge this repo. My permission on regen-network/regen-data-standards is push=false, so approval by Monday doesn't by itself land anything. Someone with push has to execute it. @glandua, since you're already the tiebreaker in that rule, I'd suggest you — but say if it should be someone else.

Related, and worth stating because it has been quietly shaping this PR: the same permission means I can't request reviewers here either. #55, #56 and #58 all show zero requested reviewers, and I'd been reading that as an oversight on our side. It isn't — the button 403s for me. The only mechanism I have is @-mentioning people in a comment, which is what I've been doing.

That also puts @blushi's review in a different light. She hasn't been the only human reviewing these because nobody else was asked; she's been reviewing voluntarily, because nobody can be formally asked. If we want review load spread beyond her, that's an access change rather than a process change.

Two small asks that follow from it:

  • Someone with push on this repo picks up the merge for whatever gets approved Monday.
  • Same for #61 (standards-agent-context.md onto main), which is deliberately independent of this ADR and shouldn't wait on it.

@DarrenZal

Copy link
Copy Markdown
Contributor Author

[CONSISTENCY] — raising this against my own ADR, before Monday rather than after.

D7's ordering rule is not implementable for hasCoBenefits as the schema stands. D7 says multivalued elements are canonicalized by sorting on the element's canonical identifier. But hasCoBenefits is inlined: true with range: Impact, and Impact.yaml @ 0a4ba12a has no identifier slot: its only required: true attribute is hasImpactType, and name is optional (the schema says it "will be inferred from the ImpactType if not provided").

So for a claim with two co-benefits there is nothing guaranteed to be present to sort on, and two engines can order them differently and produce different fingerprints from the same claim. That is the determinism property D7 exists to provide.

This matters now because D7 is one of the decisions proposed for Accepted at Monday's sign-off. Three options as I see them, and I don't think it's my call alone:

  1. Define a deterministic fallback ordering for identifier-less inlined objects — e.g. sort on the canonical N-Quads serialization of the element itself, which is well-defined under RDFC-1.0 and needs no schema change.
  2. Add a required identifier to Impact (schema change, wider blast radius than this ADR).
  3. Scope D7 to multivalued elements that do carry identifiers, and record hasCoBenefits as a named gap rather than implying it is covered.

I lean (1), because it keeps the fix inside the canonicalization layer where D7 lives and doesn't reach into the schema. But (3) is the honest minimum if we want D7 Accepted on Monday without deciding (1) properly.

Found while drafting the split and reviewing it adversarially before opening a PR — noting that because it's the pre-work check @davefortson's protocol 3 asks for, and this is what it caught.

@blushi

blushi commented Aug 13, 2026 •

Copy link
Copy Markdown
Member

1. What kind of data was anchored as raw on mainnet

You asked about the chain, so I ran a chain-wide census rather than answering only for us. Over the tx index, 2026-02-24 → 2026-08-05: 950 MsgAnchor — 418 Raw, 532 Graph.

Chain-wide, Raw is mostly documents and images, not JSON: pdf 259, jpeg 58, jpg 50, png 24, json 19, webp 5, heic 2, bin 1.

@DarrenZal I was only interested in what was anchored as raw as part of the current context, ie through the koi-processor claims engine only, not all the raw anchored content, sorry if that wasn't clear. If it's only testing data, then it's probably fine to ignore them. BTW testing data should have been anchored on testnet, not on mainnet.

@blushi

blushi commented Aug 13, 2026 •

Copy link
Copy Markdown
Member

Marie — this probably invalidates part of your review, since your threads were written against the broad document. Rather than assume they carry over, could you say which still apply to the narrowed text? I'd expect a fresh pass on it.

Github comments applied to changed text are simply marked as "outdated" so this should be easy to check comments, on the contrary, that still apply, although I can take a fresh pass when the changes are actually pushed.

@glandua asked that @blushi be added to the Deciders line. She has done effectively all of the review on this ADR, so the substance of the request is obviously right — but Gregory isn't himself a decider, and Marie hasn't been asked whether she wants the role or its ongoing workload. So: Shawn, JC — do you agree? Marie — do you want it?

I'd rather ask than edit the line and call it done.

I'm fine being a decider, I've already put some significant work into the reviewing process and claims engine is one of my main focus in fact.

On another note, I added you to the collaborator of this repo, with write access to ease processes.

Also, this ADR feels broader than this repo’s scope. If regen-data-standards is meant to define RDF schemas and vocabularies, then canonicalization, hashing, ledger anchoring, migration policy, KOI/JC engine parity, DoS guards, etc. seem like application/interoperability concerns rather than schema-standard concerns.
I’d keep only the RDF-shape decisions here: classes, predicates, ranges, required fields, cardinality, controlled vocabularies, and maybe RDF encoding conventions where they directly affect generated RDF. The content-addressed identity protocol should probably live in a separate spec or implementation repo and be linked from here.

@DarrenZal

Copy link
Copy Markdown
Contributor Author

Thank you for the write access, and for taking the decider role. I'll add you to the Deciders line now that you've said yes — that was the only thing I was waiting on.

Taking your three points in order.

The raw census. My mistake, and a useful one to have made: you asked a narrow question and I answered a chain-wide one. Scoped to the koi-processor claims engine, every raw anchor we have made is a test or demo fixture rather than a real claim, so yes, ignorable. I'm deliberately not quoting a count here: local and team-prod databases disagree and I haven't reconciled them. And you're right that they should have gone to testnet. Worth saying plainly: that was a deliberate team decision I went along with, not an accident — "mainnet is our testnet" was agreed on the 2026-03-10 call. It was the wrong call and it left permanent test records on mainnet. I'd rather we reverse it than repeat it.

Outdated threads. Also right, and I gave you work GitHub already does. I'll check the outdated markers myself once the changes are pushed rather than asking you to audit your own review.

Scope — and I think you may be right. This is the more interesting one, and it's a bigger challenge than the split I proposed, so I'd rather sit with it than reflexively defend.

Your cut is: this repo defines RDF shapes (classes, predicates, ranges, cardinality, vocabularies, encoding conventions where they affect generated RDF), and canonicalization / hashing / anchoring / migration policy / engine parity / DoS guards are application-and-interop concerns that belong in a separate spec.

Reading ADR 0001 against that line, most of it falls on your side of it. D2 (algorithm), D3 (lexical normalization), D4 (DoS cap), D5 (digest), D8 (migration) are all "how bytes are produced and put on a ledger" — none of them constrain what RDF we generate. The parts that genuinely are schema concerns are narrower: which fields are identity-bearing, whether hasClaimant is substance or provenance, whether lifecycle fields belong on Claim at all. Those are shape questions and they belong here.

That is close to the split I proposed but cut along a different axis, and yours is the better axis: I was splitting on settled vs open, which is a scheduling property. You're splitting on what kind of thing it is, which survives contact with time.

One complication worth putting in front of you before either of us commits. I posted the split proposal yesterday claiming canonicalization is cleanly separable, then had to correct it (issue #60): JC's engine doesn't hash a single graph. Its preimage is four parts — canonical Claim IRI, content RDF dataset, assertor IRI, asserted_at — per fingerprint.rs. Only the second is RDF. So the composition rule takes the assertor as an input, which is exactly the shape question that stays in this repo. The two documents would have a real seam between them, not a clean cut.

So: where would you want the identity protocol to live? A new repo, koi-processor alongside the implementation, or somewhere in koi-gov? I don't have a view worth defending and you know this repo's boundaries better than I do. Happy to do the restructuring once you say where.

I'd suggest we don't merge anything on Monday until this is settled. The sign-off deadline was to stop the decision drifting, not to force it past a good question.

@DarrenZal

Copy link
Copy Markdown
Contributor Author

@JeancarloBarrios @blushi, could you give us a read on these three seams before tomorrow's 9:30 AM PDT call, if possible?

I'm not proposing another ADR or schema edit today. I want us to leave the call with a clear owner and an async ratification path, without reopening the whole design.

  1. What exact versioned RDF encoding profile should regen-data-standards own so two engines produce the same graph? I think that needs to cover predicate IRIs, whether claim type is an IRI or literal, expansion rules, and normalization.
  2. JC, can you confirm the current four-part fingerprint preimage: canonical Claim IRI, content RDF dataset, assertor IRI, and asserted_at? Could we agree on one cross-engine test vector that makes the canonicalized and concatenated pieces explicit?
  3. Marie, does the repo boundary you proposed still look right: RDF shape and directly related encoding conventions here, with the content-addressed identity protocol elsewhere? If so, where should that protocol live, and should claimant/assertor and lifecycle state be represented in the RDF shape, a separate protocol record, or both?

A short answer is enough. I can turn it into the concrete follow-up after the call.

DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 8, 2026
Marie's review: this repo is about data standards, not how the resulting
data is used within regen ledger or koi-processor. Keeps the standards
compatibility / provenance checklist and the ADR<->schema alignment rules;
moves ledger constraints, the source tables and the worked example to
koi-processor (gaiaaiagent/koi-processor#54).

Also replaces the two repo-relative docs/adr links, which 404 on main
because that file only exists on the regen-network#56 branch -- the unresolved Codex
review comment from 2026-08-13.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
DarrenZal and others added 10 commits September 23, 2026 17:26
…ecklist

Dave's 2026-08-04 process proposal asked for a small shared "agent context"
file: non-negotiable design constraints, links to authoritative docs, and a
bootstrap point for agent workflows that humans can read as a checklist.

Ships with the ADR PR because the ADR is what violated the constraint: D2-b
recommended JCS over a typed JSON projection, which MsgAttest can never
accept. Every constraint carries a citation and a consequence.

Constraints beyond the known MsgAttest/Raw pair, all verified in code:

- v2 no longer validates algorithm identifiers on-chain. digest_algorithm,
  canonicalization_algorithm and merkle_tree are uint32, not enums; the
  validators reject only zero, and the strict membership checks
  DigestAlgorithm.Validate / GraphCanonicalizationAlgorithm.Validate have
  zero non-test call sites. A wrong algorithm identifier is accepted
  silently. Client-side correctness is the only control.
- The IRI encodes the hash type (IriPrefixRaw=0 / IriPrefixGraph=1), so
  Raw->Graph migration mints a NEW IRI rather than re-labelling one. This
  is why ADR 0001 D8 dual-anchors instead of re-minting.
- Graph IRIs must end .rdf; Raw extensions are 2-6 lowercase-or-numeric
  chars. Reading the IRI suffix is a zero-cost attestability check.
- The gap is narrower than "nothing is attestable": claims anchor Raw and
  are not attestable, but attestation records already go through
  generate_graph_iri -> broadcast_attest on a graph IRI.
- Live divergence: that deployed graph path canonicalizes with URDNA2015
  via pyld while ADR 0001 D2-a specifies RDFC-1.0. Both stamp wire byte 1
  and the chain does not check it, so a fork would be silent. pyld is also
  absent from koi-processor/requirements.txt.

Ledger citations pinned to regen-ledger 451c3a3f, verified against the
remote at that SHA; koi-processor citations to regen-prod c08c0a7e, the
deployed branch (line numbers differ from feature branches).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bare 'make -C schema' runs only gen-taxonomy (the first target), not the
full build. Verified against schema/Makefile: targets are gen-taxonomy,
lint, gen-doc, gen-rdf, clean-rdf, update-graph, all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…assumption; pin all citations

Addresses @blushi's 2026-08-03 review on regen-network#56.

- Remove "Raised in review, and it is the stronger position." (D1) — process
  metadata, not decision rationale.
- D2: RDFC-1.0 is now presented as the data module's own name for the algorithm,
  not a preference. GRAPH_CANONICALIZATION_ALGORITHM_RDFC_1_0 is the only
  non-zero value in the v2 canonicalization registry, and production targets v2
  (derive_ledger_iri sends the v2-only file_extension field). Notes the v1
  rename from URDNA2015 at the same wire value, and that the shared byte proves
  nothing about which spec ran. Adds the live divergence: the deployed graph
  path runs pyld URDNA2015 and stamps _GRAPH_CANON_URDNA2015 = 1.
- Withdraw the claim that JC's work already assumes the lifecycle-record shape.
  Marie is closer to that codebase and reports no corresponding shapes exist.
  Booked as Open Question 7 (numbered 7 so refs to 1-6 stay valid), and the
  Context section now says the ADR does not describe JC's implementation.
- D8 rewritten from dual anchoring to a straight cutover. Mainnet query found
  8 Raw anchors from the claims engine, all test/demo, and zero executed
  MsgAttest. Method and bounds stated inline. Algorithm-column bullet promoted
  to D9, which stands on its own.
- Every code citation pinned to a 40-hex commit SHA (27 links, all verified to
  resolve). No relative or branch links remain.
…e MsgAttest to D2

The previous revision's D8 rested on "there is no production Raw corpus on
mainnet" and on the claim that a chain-wide Raw census is impossible from a
query node. Both were false. A census over the tx index (2026-02-24 →
2026-08-05) returns 950 MsgAnchor: 418 Raw, 532 Graph. Ours is 8 of 418
(1.9%), 363 Raw anchors post-date our last, and the most recent is
2026-08-05. Raw is also the registry service account's live document path
(399 anchors, 259 of them PDFs), so "the registry's data is already
all-Graph" holds only for ecocredit metadata IRIs.

D8 keeps the cutover and still drops the dual-anchor machinery, but now on
corpus SIZE — re-anchoring eight throwaway fixtures is not a migration —
rather than corpus absence. Every our-account statement is scoped
explicitly to regen15eexs…, and exhaustiveness now rests on the sequence
argument (sequence: 9, nine indexed txs = 8 MsgAnchor + 1 MsgSend) instead
of a per-IRI sweep that could never prove it. That also resolves the
previous UNVERIFIED note about an unexplained ninth transaction.

D8 no longer asserts away what it cannot answer: dual-anchoring also buys
forward interoperability during a transition, and nothing gathered speaks
to whether any counterparty dereferences our Raw IRIs. Booked as Open
Question 8, together with the two unattributed mainnet accounts anchoring
Raw JSON in our shape whose 11 hashes match neither claims database.

D2 gains the positive case it was missing: four MsgAttest exist on Regen
mainnet (2026-07-13 ×2, 2026-07-19, 2026-08-03), all carrying
canonicalization_algorithm: 1. Graph attestation with the algorithm D2-a
selects is already working practice on this chain, which is a stronger
argument than "we have nothing to lose."

Also corrected: the two census transports do not both return 950 — REST
returns 950 and Tendermint tx_search returns 815 against a node pruned to
height 27420001, and exactly 815 of the 950 sit at or above that height.
They reconcile on their overlap; the bounds section says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on schema refs

Three changes, all from reading ybird-labs/claims directly on 2026-09-08.

Deciders: adds Marie Gauthier, per @glandua's 2026-08-10 request.

Q2 is answered and the two halves come apart. Algorithm: JC does use
RDFC-1.0 canonical N-Quads / UTF-8 / SHA-256 on both branches, arrived at
independently of D2-a. Substance: not equivalent, and it differs between
his own branches. main (8cbf0e5d) hashes Claim IRI + content + assertor +
asserted_at; the unmerged design branch (217fafdd) hashes declared schema
references + canonical content, derives the ClaimIRI from the fingerprint,
and removes provenance from the claim value entirely. main is the merged
default, so a casual reader gets the older model.

Q9 (new) is the consequence: this ADR's substance set contains no schema
declarations, so two implementations that both do RDFC-1.0 correctly still
produce different fingerprints. That divergence is invisible at the
algorithm layer and surfaces only as claims that fail to match -- the most
likely source of a silent cross-engine fork, and covered nowhere else here.

Which branch JC considers current is not established; asked by email today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
The crates/ tree is byte-identical between main (8cbf0e5d) and the design
branch (217fafdd): same tree hash 1d556cc5, empty diff. AssertionProvenance
and the four-part fingerprint are present on both. The new model lives in
the design doc, AGENTS.md, README.md and spike/ -- a separate cargo package
the root workspace excludes.

That relocates the open item. It is not "which branch is current"; it is
that the branch's own AGENTS.md forbids the four types its crate still
defines, so merging it would not move the implementation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
Applies the five mechanical suggestions from the 2026-09-09 review:

- D7 renumbered to D2, with the two dependent references updated (the
  hasCoBenefits row and the co-benefits question). D2-D6 moved to the
  implementation companion, so a lone D7 read as a gap.
- The `claim_type` -> `hasClaimType` mapping sentence removed; it describes an
  external implementation and does not belong in the standards repo.
- The four "moved to the implementation proposal" entries dropped (JC
  implementation, digest suite, raw-anchor migration, declared schema
  references). Recording what left the document is not the document's job.

One consequence worth flagging rather than burying: the line above that list
promises "historical question numbers are retained here so existing review
links remain interpretable", and markdown renumbers an ordered list from its
first item, so deleting entries 2, 6, 8 and 9 would have silently renumbered
the survivors 1-5 and broken exactly that promise. The remaining numbers are
now explicit labels, so the promise stays true. Say if you would rather drop
the promise instead and keep a plain list.

Marie's structural point -- an ADR with pending decisions is not an ADR -- is
answered on the thread as a proposal, not applied here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YLntBrd61M5EUcngh2Lv8m
Marie's call on the thread: "yes let's drop the promise and keep a plain list."

Removes the sentence promising that historical question numbers are retained so
existing review links stay interpretable, and with it the Q1/Q3/Q4/Q5/Q7 labels
that only existed to keep that promise true after four entries were deleted.
The section heading loses "and review continuity" for the same reason -- it was
naming the promise, not the content.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@DarrenZal
DarrenZal force-pushed the adr/claim-substance-canonicalization branch from 3f2bf5d to 8691a3b Compare September 24, 2026 00:26
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 24, 2026
Ten of the review comments, implemented. The two open questions (hash wire
format, whether PROV-O should block) and the ConsentDirective split (pending
@glandua's sign-off) are deliberately NOT in this commit.

:213 — claims/attestations become claimRefs/attestationRefs, multivalued
  uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this
  contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined
  objects during the consent rewrite without re-justifying it. A Claim's identity
  and lifecycle (verificationStatus, contentHash, supersession) evolve after
  emission, so embedding one couples this envelope's lifecycle to the claim's.
  Claim/Attestation imports dropped — a uriorcurie range doesn't need them.

:112 — observedLocalDate removed. It was a pure function of observedAt and
  sourceTimezone, so storing it only created something that could disagree with
  its own instant. The derivation is now stated normatively on sourceTimezone.

:83  — snake_case convention on sourceSystem and recordType, enforced with a
  pattern rather than described in prose. This immediately caught the very
  inconsistency reported: the biocultural example's "field-signing-app" failed
  validation and is now "field_signing_app".

:211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0.

:236 — PROV-O, narrowly. Added the prov prefix and emittedAt ->
  prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing
  sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent:
  wasGeneratedBy relates an Entity to an Activity rather than a string field, and
  prov:Agent is a class not a property, so those mappings would produce something
  that looks PROV-aligned and isn't. Real alignment means modelling the sensor run
  as a prov:Activity — proposed as its own ADR.

docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56.
docs:131 — Pydantic section removed; binding is the consumer's concern and
  belongs in the consuming repo. Fixed the header line that referenced it.

Codex P2 (raw-data residency) — the descriptions on rawContentInline and
  rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No
  such rule exists; the only rule requires rawContentHash when processingState is
  COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so
  the constraint is not expressible on this class without hoisting a mirror field
  (two sources of truth for one policy — worse than an honest gap). Descriptions
  now say enforcement is external, a new doc section states the consumer
  obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml
  is committed as a negative fixture: a SOVEREIGN record carrying inline raw that
  passes validation. Kept so the gap stays visible and so the fixture starts
  failing if a future LinkML or SHACL shape ever expresses it.

Verified: both positive examples validate against the edited schema; the negative
fixture validates (which is the point); linkml-lint reports no new problems — the
one error it emits ("id is not a uri") is pre-existing and repo-wide, present on
unmodified OutputRecord.yaml and on Claim.yaml as merged on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 24, 2026
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses
the lean I gave her, so the reasoning is in the doc under "Hash wire format" and
summarised here.

I said I'd probably go bare hex, to match the implementation and the ledger.
Three findings support that, and I re-derived all three:

- koi-processor emits bare lowercase hex — ledger_anchor.py:96,
  `hashlib.blake2b(..., digest_size=32).hexdigest()`.
- `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented
  the token in these examples.
- The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are
  `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto),
  never a prefixed string.

What flipped it:

- A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB.
  `koi_memories.content_hash` is SHA-256 (migrations/004:28) and
  `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same
  shape, same length, different algorithm, one codebase.
- ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the
  missing discriminator as a latent defect in its own words: "the current schema
  cannot express which canonicalization produced a stored hash." Shipping an
  undiscriminated hash in a WIRE CONTRACT repeats that defect where it is
  hardest to fix — the ledger and the claims tables can add an adjacent column
  later because they own both sides of the read. A contract consumed by parties
  we don't control cannot.

So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash.

Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the
ledger's DigestAlgorithm registry (today exactly one non-zero member,
DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum
gains one. It does NOT name the ContentHash kind or a canonicalization. For
rawContentHash that is complete — a raw payload is not RDF, anchors as
ContentHash.Raw, and the proto says Raw "does not specify a deterministic,
canonical encoding". For recordContentHash it is NOT complete: this envelope
does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same
question ADR 0001 is deciding for Claim. Flagged in the slot description rather
than quietly decided.

Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is
lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks
string equality and dedup on a value whose whole purpose is content-addressed
identity.

Rejected multihash/multibase: properly self-describing, but needs a codec table
on both sides, produces values nothing in our stack emits, and diverges from the
ledger's registry for no gain we can spend today.

Cost, stated rather than hidden: koi-processor must prefix at the contract
boundary. One line, and the right direction — the standard sets the wire format
and the implementation adapts.

Not touched: Claim.contentHash and Attestation.contentHash on main have the same
gap (range: string, no pattern). Same class as Entity having no identifier slot;
belongs in the identity-keys follow-up, not in a PR editing merged schemas.

Verified: `make -C schema lint` clean; all three examples pass
`linkml-validate -C OutputRecord`; and the pattern was tested against what it
must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest
all fail with "does not match".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 24, 2026
Marie's review: this repo is about data standards, not how the resulting
data is used within regen ledger or koi-processor. Keeps the standards
compatibility / provenance checklist and the ADR<->schema alignment rules;
moves ledger constraints, the source tables and the worked example to
koi-processor (gaiaaiagent/koi-processor#54).

Also replaces the two repo-relative docs/adr links, which 404 on main
because that file only exists on the regen-network#56 branch -- the unresolved Codex
review comment from 2026-08-13.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 24, 2026
Ten of the review comments, implemented. The two open questions (hash wire
format, whether PROV-O should block) and the ConsentDirective split (pending
@glandua's sign-off) are deliberately NOT in this commit.

:213 — claims/attestations become claimRefs/attestationRefs, multivalued
  uriorcurie. This is a regression fix, not a redesign: the July-3 draft of this
  contract used uriorcurie refs throughout and regen-network#55 replaced them with inlined
  objects during the consent rewrite without re-justifying it. A Claim's identity
  and lifecycle (verificationStatus, contentHash, supersession) evolve after
  emission, so embedding one couples this envelope's lifecycle to the claim's.
  Claim/Attestation imports dropped — a uriorcurie range doesn't need them.

:112 — observedLocalDate removed. It was a pure function of observedAt and
  sourceTimezone, so storing it only created something that could disagree with
  its own instant. The derivation is now stated normatively on sourceTimezone.

:83  — snake_case convention on sourceSystem and recordType, enforced with a
  pattern rather than described in prose. This immediately caught the very
  inconsistency reported: the biocultural example's "field-signing-app" failed
  validation and is now "field_signing_app".

:211, :370 — priorityScore bounded 0.0..1.0; segmentCount minimum 0.

:236 — PROV-O, narrowly. Added the prov prefix and emittedAt ->
  prov:generatedAtTime as an exact_mapping. Deliberately NOT re-pointing
  sensorId/sensorVersion at prov:wasGeneratedBy or consentedBy at prov:Agent:
  wasGeneratedBy relates an Entity to an Activity rather than a string field, and
  prov:Agent is a class not a property, so those mappings would produce something
  that looks PROV-aligned and isn't. Real alignment means modelling the sensor run
  as a prov:Activity — proposed as its own ADR.

docs:112 — the bare "that ADR" now names and links ADR 0001 / PR regen-network#56.
docs:131 — Pydantic section removed; binding is the consumer's concern and
  belongs in the consuming repo. Fixed the header line that referenced it.

Codex P2 (raw-data residency) — the descriptions on rawContentInline and
  rawDataStaysAtSource both claimed enforcement "by an OutputRecord rule". No
  such rule exists; the only rule requires rawContentHash when processingState is
  COMPLETE. A LinkML rule cannot traverse into consent.rawDataStaysAtSource, so
  the constraint is not expressible on this class without hoisting a mirror field
  (two sources of truth for one policy — worse than an honest gap). Descriptions
  now say enforcement is external, a new doc section states the consumer
  obligation, and schema/examples/output-record.INVALID-sovereign-inline-raw.yaml
  is committed as a negative fixture: a SOVEREIGN record carrying inline raw that
  passes validation. Kept so the gap stays visible and so the fixture starts
  failing if a future LinkML or SHACL shape ever expresses it.

Verified: both positive examples validate against the edited schema; the negative
fixture validates (which is the point); linkml-lint reports no new problems — the
one error it emits ("id is not a uri") is pre-existing and repo-wide, present on
unmodified OutputRecord.yaml and on Claim.yaml as merged on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DarrenZal added a commit to DarrenZal/regen-data-standards that referenced this pull request Sep 24, 2026
Settles open question regen-network#1 and my 31 July question to @blushi (:210). It reverses
the lean I gave her, so the reasoning is in the doc under "Hash wire format" and
summarised here.

I said I'd probably go bare hex, to match the implementation and the ledger.
Three findings support that, and I re-derived all three:

- koi-processor emits bare lowercase hex — ledger_anchor.py:96,
  `hashlib.blake2b(..., digest_size=32).hexdigest()`.
- `grep -r b2s256` over koi-processor and regen-ledger: zero hits. I invented
  the token in these examples.
- The ledger carries the algorithm OUT of band: ContentHash.Raw and .Graph are
  `{hash: bytes, digest_algorithm: uint32, …}` (proto/regen/data/v2/types.proto),
  never a prefixed string.

What flipped it:

- A bare 64-hex string is ambiguous, and the ambiguity is already in our own DB.
  `koi_memories.content_hash` is SHA-256 (migrations/004:28) and
  `claims.content_hash` is BLAKE2b-256 (migrations/064:41). Same name, same
  shape, same length, different algorithm, one codebase.
- ADR 0001 D8 (regen-network#56) commits us to two coexisting anchoring schemes and names the
  missing discriminator as a latent defect in its own words: "the current schema
  cannot express which canonicalization produced a stored hash." Shipping an
  undiscriminated hash in a WIRE CONTRACT repeats that defect where it is
  hardest to fix — the ledger and the claims tables can add an adjacent column
  later because they own both sides of the read. A contract consumed by parties
  we don't control cannot.

So: `pattern: "^b2s256:[0-9a-f]{64}$"` on rawContentHash and recordContentHash.

Scoped honestly. The token names the DIGEST ALGORITHM only, drawn from the
ledger's DigestAlgorithm registry (today exactly one non-zero member,
DIGEST_ALGORITHM_BLAKE2B_256 = 1); a new token may only be minted when that enum
gains one. It does NOT name the ContentHash kind or a canonicalization. For
rawContentHash that is complete — a raw payload is not RDF, anchors as
ContentHash.Raw, and the proto says Raw "does not specify a deterministic,
canonical encoding". For recordContentHash it is NOT complete: this envelope
does have an RDF projection, so Raw-over-JSON vs Graph-over-RDFC-1.0 is the same
question ADR 0001 is deciding for Claim. Flagged in the slot description rather
than quietly decided.

Lowercase only, against the bot's suggested `[0-9a-fA-F]`: hexdigest() is
lowercase, and mixed case gives one fingerprint 2^64 spellings, which breaks
string equality and dedup on a value whose whole purpose is content-addressed
identity.

Rejected multihash/multibase: properly self-describing, but needs a codec table
on both sides, produces values nothing in our stack emits, and diverges from the
ledger's registry for no gain we can spend today.

Cost, stated rather than hidden: koi-processor must prefix at the contract
boundary. One line, and the right direction — the standard sets the wire format
and the implementation adapts.

Not touched: Claim.contentHash and Attestation.contentHash on main have the same
gap (range: string, no pattern). Same class as Entity having no identifier slot;
belongs in the identity-keys follow-up, not in a PR editing merged schemas.

Verified: `make -C schema lint` clean; all three examples pass
`linkml-validate -C OutputRecord`; and the pattern was tested against what it
must reject — bare hex, uppercase hex, a `sha256:` token and a 63-char digest
all fail with "does not match".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Carries forward the July 2026 agreement to remove verificationStatus,
contentHash and dataIri from Claim.yaml, as agreed rather than proposed.
Places impact, co-benefit, quantity, credit-class and methodology slots
in specialized claim schemas, with co-benefit collection semantics
decided in that scope. Namespace leaves the open decisions and
methodology closes only conditionally. Claim period, hasOperator and the
PROV-O alignment (input: regen-network#82) stay open. Status stays Proposed.
@DarrenZal

Copy link
Copy Markdown
Contributor Author

@blushi, this push implements your #70 corrections, as committed in my follow-up.

  • verificationStatus, contentHash, dataIri: recorded as agreed in July, to be removed from Claim.yaml, not made optional. The open question is now how external review state is represented, to discuss with Jeancarlo. #71 owns the schema edit.
  • Base Claim versus specialized types: impact, co-benefits, quantity, credit class and methodology now belong to specialized claim schemas. They can still contribute to a specialized claim's identity via claims#1. Claim period and hasOperator stay pending.
  • D2: the general collection rule stays. The hasCoBenefits set-versus-sequence call moves to the specialized-schema scope.
  • Namespace and methodology: namespace leaves the open decisions. Methodology closes conditionally, naming the claim-type content and the supersession relationship as still undefined.
  • WP1-01 — Preliminary research on claims and provenance data models in and outside RDF #67: the PROV-O part stays open, with PR #82 as its input.
  • Six additions: not added.

Status stays Proposed. A 2026-09-23 revision-history entry records the change.

blushi pushed a commit that referenced this pull request Sep 24, 2026
* docs: land standards-agent-context.md on main

Adds the shared agent-context / non-negotiables file proposed as protocol 4
of the 2026-08-04 standards-workflow discussion.

It currently exists only on the ADR 0001 branch, which means the file meant to
be the shared checklist is invisible to everyone who would use it. Landing it
independently so it is available before the next piece of work starts, rather
than arriving with an ADR.

Contents are unchanged from the ADR branch: the ledger non-negotiables
(MsgAttest accepts only ContentHash.Graph; Raw disclaims canonical encoding),
where the authoritative sources live, a pre-PR checklist, the review tag
vocabulary, and a worked example of the error the file exists to prevent.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: narrow standards-agent-context to schema/ADR scope per review

Marie's review: this repo is about data standards, not how the resulting
data is used within regen ledger or koi-processor. Keeps the standards
compatibility / provenance checklist and the ADR<->schema alignment rules;
moves ledger constraints, the source tables and the worked example to
koi-processor (gaiaaiagent/koi-processor#54).

Also replaces the two repo-relative docs/adr links, which 404 on main
because that file only exists on the #56 branch -- the unresolved Codex
review comment from 2026-08-13.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GHFvPJqgz325Ai66joFxoW

* docs: finish standards context scope review

* docs: point at standards-agent-context.md from the root README

Marie's review: the doc covers more than schemas, so a schema/README pointer
would file the whole thing under the narrowest of its four sections. Sections 1
and 2 are schema-specific, but section 3 is a repo-level source index and
section 4 (review tag vocabulary, author response convention) governs any PR in
this repo.

Root README it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(agent-context): say explicitly that the review-tag vocabulary covers only two cases

Addresses Marie's 2026-09-14 review note on §4: the table lists two tags and
did not say that this is the whole vocabulary on purpose. Now states that it
covers only the two failure classes that have cost review time here, that
untagged comments are fine, and that new tags are added by PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: rename standards-agent-context.md to AGENTS.md

AGENTS.md is a cross-vendor convention rather than one vendor's
filename, so the objection that blocked this rename on 2026-09-10
does not hold. Claude Code reads it from the repository root from
v2.1.277, and it is the name Cursor and GitHub Copilot already read,
so the file is now picked up without anyone remembering to point at it.
This repository has no CLAUDE.md, which is the condition for Claude
Code to auto-load AGENTS.md.

The scope note now states which sections are schema and ADR specific
(1 to 3) and which apply to every pull request here (4), since a file
at this name will be read for changes of any kind.

The root README pointer from b6abb27 is kept and repointed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
…ic Claim

hasClaimant is asserted, in-content attribution of whoever takes
responsibility for the assertion (review on regen-network#56, regen-network#82 section 3.1). Its
PROV subproperty declaration stays with the PROV-O alignment.

hasOperator leaves the generic Claim: the operator is described on the
relevant domain activity in the specialized claim schema, as
responsibility for that activity rather than attribution of the Claim
(regen-network#82 section 3.2). Claimant and operator leave the open decisions.

Also drops the Context sentence describing earlier versions of the ADR,
which the revision history already records.
@DarrenZal
DarrenZal requested a review from blushi September 24, 2026 21:56
Comment thread docs/adr/README.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
Comment thread docs/adr/0001-claim-substance-canonicalization.md Outdated
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Low risk] Adds architecture decision record documentation.

The documentation PR appears safe to merge as a proposal, with schema implementation and governance acceptance still explicitly pending.

Summary

The PR adds a proposed ADR for Claim RDF shape, provenance, collection semantics, and lifecycle boundaries, plus an ADR index. It separates these proposals from service identity and anchoring decisions and explicitly leaves schema implementation and governance acceptance pending.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Claim["Claim: immutable asserted content"] --> Subject["Subject"]
  Claim --> Claimant["Claimant attribution"]
  Claim --> Evidence["Cited evidence"]
  Attestation["Separate attestation"] --> Claim
  Service["Service lifecycle and anchoring records"] --> Claim
Loading

Reviews (1) · Last reviewed commit: "docs(adr): specify proposed Claim conten..."

@DarrenZal DarrenZal changed the title ADR 0001: Claim RDF shape and provenance boundaries (Proposed) docs: propose Claim RDF shape and provenance boundaries Oct 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants