Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 49 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,3 +57,52 @@ jobs:

- name: Check dependency hygiene
run: bun run check:deps

rust-validation:
name: Cross-validate vectors against the Rust reference
runs-on: ubuntu-latest

steps:
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7

- name: Setup Bun
uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2
with:
bun-version: latest

- name: Setup Node.js
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7
with:
node-version: '24'

- name: Install dependencies
run: bun install --frozen-lockfile

- name: Check the golden vectors against the working tree
run: bun run test:golden

# The whole corpus (tests/corpus/corpus.ts `allRecipes()`)
- name: Materialise the full corpus
run: bun run vectors:full

# The runner image ships a stable Rust toolchain
- name: Fetch the pinned reference crates
working-directory: tests/rust-validation
run: cargo fetch --locked

- name: Validate the golden vectors against the reference
working-directory: tests/rust-validation
run: cargo run --release --offline --locked -- ../vectors/vectors.json

- name: Validate the full corpus against the reference
working-directory: tests/rust-validation
run: cargo run --release --offline --locked -- ../vectors/full.json

- name: Check that a wrong vector fails the run
working-directory: tests/rust-validation
run: |
if cargo run --release --offline --locked -- mismatch.json; then
echo "the mismatch fixture was accepted" >&2
exit 1
fi
3 changes: 2 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -25,5 +25,6 @@ coverage
.env
.env.*

# Rust cross-validation harness build artifacts
# Rust cross-validation harness build artifacts and the materialised full corpus
tests/rust-validation/target
tests/vectors/full.json
6 changes: 3 additions & 3 deletions .size-limit.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,17 @@
{
"name": "ESM root entry (import *), minified + gzipped",
"path": "dist/index.mjs",
"limit": "21 kB"
"limit": "34 kB"
},
{
"name": "bytewords subpath, minified + gzipped",
"path": "dist/bytewords.mjs",
"limit": "8 kB"
"limit": "21 kB"
},
{
"name": "fountain subpath, minified + gzipped",
"path": "dist/fountain.mjs",
"limit": "18 kB"
"limit": "31 kB"
},
{
"name": "root entry, package's own code (crypto and dcbor external)",
Expand Down
71 changes: 70 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,74 @@
# Changelog

## 1.0.0-beta.3 - 2026-09-14

Decoding follows the reference in full: bc-ur 0.19.2 over the `ur` 0.4.1 and
`minicbor` 0.19.1 crates. Every encoder output is unchanged.

### Changed (breaking)

- **Multipart decoding follows the reference's fountain decoder.** A mixed
part is reduced against the fragments known when it arrives and a buffered
part only when a later simple part comes, so completion can need more
parts than before; a later simple part replaces a fragment derived earlier; the
reassembled message is not checked against the parts' checksum field
(the bytewords checksum of each part still is); and a part with `seqNum` 0
is accepted and counted as the reference's wrapped index, so `result`
then throws `Decoder` "expected item". `FountainDecoder.add` copies the
part's data and never keeps the caller's array.
- **`MultipartDecoder.add` is case-sensitive**, as the reference's `receive`
(only `UR.parse` lower-cases): lower-case a QR payload before `add`.
Every string is validated after completion too, and `add` reports the
reference's outcomes in its order: `InvalidScheme`, `InvalidType`,
`UnexpectedType`, then `Decoder` with the reference's text ("No type
specified" for `ur:test`, "Invalid indices", the bytewords failure before
"Can't decode single-part UR as multi-part", the part codec's messages).
- **`done` and `result`.** `done` is the fountain decoder's completion;
`result` reassembles and decodes the message when first read after
completion and keeps that outcome, throwing `Decoder` ("invalid padding",
"expected item") or `Cbor`. `add` no longer throws for a message that
fails to reassemble or decode.
- **Part CBOR is read as `minicbor` reads it**: integer heads of any width,
byte strings of any length width and trailing bytes are accepted, and
failures carry minicbor's text ("unexpected type u8 at position 0:
expected array", "4294967296 overflows target type at position 1: when
converting u64 to u32", "decode error: invalid CBOR array length", "end of
input bytes"). `encodeFountainPart` truncates the three counters to `u32`
as the reference's `as u32` does.
- **`UR.parse` and `UR.decodeBytes` follow `from_ur_string`**: after the
scheme, the missing slash and the type, the multipart header is checked
("Invalid indices"), then the payload's bytewords, then `NotSinglePart`.
A bytewords failure inside a UR string is `Decoder` with the reference's
reason ("invalid word", "invalid checksum", "invalid length"), no longer
`Bytewords`; `decodeBytewords` alone throws `Bytewords`.
- **The empty UR type is valid**, as the reference's `URType::new` accepts
it: `ur:/…` is produced and read, `URType.isValid("")` is true and
`expectType("")` is `UnexpectedType`.
- **`decodeURWith` accepts only a UR named after the codec's first tag**, as
the reference's `from_ur` (a codec's other tags do not name accepted UR
types), and reports a wrong or invalid type as a dcbor `CborError`
(`Custom`: "expected UR type seed, but found crypto-seed", "invalid UR
type"), as `from_ur` returns a `dcbor::Error`. `TagUnnamed` stays a
`URError`.
- **`canonicalizeByteword` lower-cases ASCII letters only**, as the
reference's `to_ascii_lowercase`; a token with a non-ASCII character
(U+212A KELVIN SIGN, U+0130, U+017F) names no word. `BYTEWORDS` and
`BYTEMOJIS` are frozen.
- **Every argument is checked before any work**, and a fault is `URError`
`InvalidParameter` with `details: { parameter, value }` and the value
rendered exactly: byte arguments must be `Uint8Array`s (a `Buffer` or a
`Uint8Array` from another realm is accepted), string arguments strings,
style strings one of their union, a hand-built `FountainPart` an object
with safe-integer counters, a `u32` checksum and `Uint8Array` data;
`fragmentLength`, `chooseFragments` (`seqNum` and `seqLen` ≥ 1, a `u32`
checksum) and `mixFragments` (non-empty arrays) validate too. Such calls
used to return a value or throw a `TypeError`. `details.value` is
`unknown`.
- **`maxFragmentLength` accepts a `bigint`** up to 2⁶⁴ − 1 in
`MultipartEncoder` and `FountainEncoder`, the reference's `usize`; a
`number` of 2⁵³ or more stays `InvalidParameter`, whose message names the
`bigint` form.

## 1.0.0-beta.2 - 2026-09-12

Compatibility and maintenance updates against `bc-ur-rust` 0.19.2 and
Expand All @@ -20,7 +89,7 @@ below align selected behaviors with the reference.
the `seqNum-seqLen` header is everything between the type and the
**last** slash, each number read as `u16::from_str` reads it (an optional
`+`, digits, at most 65535), and the fountain fields come from the part's
CBOR — the header is no longer compared with them. `UR.parse` accepts the
CBOR - the header is no longer compared with them. `UR.parse` accepts the
`+` too.
- A `maxFragmentLength` of 0 in `MultipartEncoder` / `FountainEncoder`, and
an empty `FountainEncoder` message, report the reference's decoder errors
Expand Down
81 changes: 47 additions & 34 deletions MIGRATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,11 +13,16 @@
- [ ] Bytewords moved to the `/bytewords` subpath; `BytewordsStyle.Minimal` → `"minimal"`.
- [ ] `MultipartDecoder.receive/isComplete/message()` → `add/done/result`;
`MultipartEncoder.currentIndex()/partsCount()` → `index/partCount`, and it is iterable.
Since 1.0.0-beta.2 `add` rejects a single-part UR (`Decoder`, as the reference's
`MultipartDecoder` does) — parse those with `UR.parse` — and no longer compares the
`n-m/` header with the part's CBOR (the reference reads the fountain fields from the CBOR).
- [ ] `decodeBytewords` is case-sensitive since 1.0.0-beta.2 (the reference's `bytewords::decode`);
`UR.parse` and `MultipartDecoder.add` still lower-case a whole UR string.
`add` rejects a single-part UR (`Decoder`, as the reference's `MultipartDecoder` does) —
parse those with `UR.parse` — and does not compare the `n-m/` header with the part's
CBOR (the reference reads the fountain fields from the CBOR). `add` is case-sensitive
(lower-case a QR payload first) and validates every string after completion; `done` is
the fountain decoder's completion and `result` reassembles and decodes on first read,
throwing `Decoder` or `Cbor`.
- [ ] `decodeBytewords` is case-sensitive (the reference's `bytewords::decode`);
`UR.parse` lower-cases a whole UR string, `MultipartDecoder.add` does not.
- [ ] `decodeURWith` accepts only a UR named after the codec's first tag and reports a wrong
type as a dcbor `CborError` (`Custom`), as the reference's `from_ur` does.
- [ ] Catch one `URError` and switch on `code`; the nine error classes,
`Result` and `isError` are gone.
- [ ] `UREncodable`/`URDecodable`/`URCodable` → `ToUR`, `urFor(value)`, `decodeURWith(ur, codec)`.
Expand Down Expand Up @@ -52,7 +57,10 @@
| `isURTypeChar`, `isValidURType`, `validateURType` | `URType.isValid(s)` and the constructor |

The `Cbor` is canonical `@blockchaincommons/dcbor`. `UR.parse` validates in
the reference order: scheme, then type, then payload.
the reference order: scheme, type, the multipart header if there is one
("Invalid indices"), the payload's bytewords (`Decoder` with the reference's
reason), `NotSinglePart`, then the CBOR. The empty type is valid, as the
reference's `URType::new` accepts it.

## 3. Bytewords (`/bytewords` subpath)

Expand All @@ -63,23 +71,23 @@ the reference order: scheme, then type, then payload.
| `encodeToWords(data)` | `identifier(data)` |
| `encodeToMinimalBytewords(data)` | `identifier(data, { style: "minimal" })` |
| `encodeToBytemojis(data)` | `identifier(data, { style: "bytemoji" })` |
| `encodeBytewordsIdentifier(data)` | `shortIdentifier(data)` (4 bytes; `RangeError` otherwise) |
| `encodeBytewordsIdentifier(data)` | `shortIdentifier(data)` (4 bytes; `InvalidParameter` otherwise) |
| `encodeBytemojisIdentifier(data)` | `shortIdentifier(data, { style: "bytemoji" })` |
| `isValidBytemoji`, `canonicalizeByteword`, `BYTEWORDS`, `BYTEMOJIS` | unchanged (tables are `readonly string[]`) |
| `isValidBytemoji`, `canonicalizeByteword`, `BYTEWORDS`, `BYTEMOJIS` | unchanged; the tables are frozen, and `canonicalizeByteword` lower-cases ASCII letters only |
| `bytewords.encode/decode/Style/…` namespace | removed; import the subpath |
| `BYTEWORDS_MAP`, `MINIMAL_BYTEWORDS_MAP` | removed (internal) |

## 4. Multipart

| `@bcts/uniform-resources` | `@blockchaincommons/uniform-resources` |
| --- | --- |
| `new MultipartEncoder(ur, maxLen)` | unchanged (`RangeError` for `maxLen < 1`) |
| `new MultipartEncoder(ur, maxLen)` | unchanged; `maxLen` is a safe integer or a `bigint` up to 2⁶⁴ − 1 (0 is `Decoder`, anything else outside that `InvalidParameter`) |
| `encoder.nextPart()` | unchanged; the encoder is also iterable |
| `encoder.currentIndex()` | `encoder.index` |
| `encoder.partsCount()` | `encoder.partCount` |
| `decoder.receive(part)` | `decoder.add(part)` → `boolean` (made progress) |
| `decoder.isComplete()` | `decoder.done` |
| `decoder.message()` → `UR \| null` | `decoder.result` → `UR \| undefined` |
| `decoder.message()` → `UR \| null` | `decoder.result` → `UR \| undefined` (throws `Decoder` / `Cbor` when the completed message does not reassemble or decode) |
| — | `decoder.progress`, `decoder.reset()` |

```diff
Expand Down Expand Up @@ -108,14 +116,16 @@ One class, `URError`, with a `code` union:
| `CBORError` | `"Cbor"` (`cause`) |
| `URDecodeError`, generic `URError`, bare `Error` from the fountain layer | `"Decoder"` |

| `RangeError` for `maxFragmentLength < 1`, an empty message, a short identifier that is not 4 bytes; `TypeError`/garbage for `NaN` or `1.5` | `"InvalidParameter"` (`details: { parameter, value }`), also for a hand-built `FountainPart` outside `u32` |
| `RangeError` for `maxFragmentLength < 1`, an empty message, a short identifier that is not 4 bytes; `TypeError`/garbage for `NaN`, `1.5` or a wrong type | `"InvalidParameter"` (`details: { parameter, value }`), for every argument outside its domain, checked before any work |
| bare `Error` from `urFor` / `decodeURWith` on an unnamed tag | `"TagUnnamed"` (`details: { tag }`) |

Messages are unchanged where a reference variant exists (`invalid UR
scheme`, `expected UR type X, but found Y`, `Bytewords error (invalid
checksum)`, …). `Result<T>` and `isError` are gone. `details` is a union
discriminated by `code`; `e.details.code === "UnexpectedType"` narrows to
`{ expected, found }`. `URResult<T>` is the non-throwing form
| `UnexpectedTypeError` from `decodableFromUR` | a dcbor `CborError` (`Custom`) from `decodeURWith`, as the reference's `from_ur` |

Messages are the reference's `Display` strings (`invalid UR scheme`,
`expected UR type X, but found Y`, `UR decoder error (invalid checksum)`,
…); a bytewords failure inside a UR string is `Decoder`, `Bytewords` comes
only from `decodeBytewords`. `Result<T>` and `isError` are gone. `details`
is a union discriminated by `code`; `e.details.code === "UnexpectedType"`
narrows to `{ expected, found }`. `URResult<T>` is the non-throwing form
(`URType.tryFrom`).

```ts
Expand All @@ -126,12 +136,12 @@ try {
}
```

Four inputs that decoded before are now rejected (see
[`RUST_DIVERGENCES.md`](./RUST_DIVERGENCES.md)): a multipart message whose
padding is not zero (`Decoder`, as the reference), a part whose fields are
not `u32`s or whose `seqNum` is 0 (`Decoder`), a URL header beyond `u16`
(`Decoder("Invalid indices")`), and the empty UR type (`InvalidType`;
`ur:/…` was never a valid UR).
Inputs the decoders treat as the reference does: a multipart message whose
padding is not zero completes and `result` throws `Decoder`; a part whose
fields are not `u32`s is `Decoder` with minicbor's text; a URL header beyond
`u16` is `Decoder("Invalid indices")`; a part with `seqNum` 0 is accepted
and counted as the reference's wrapped index; an upper-case part is
`InvalidScheme` or `InvalidType`; the empty UR type is valid.

## 6. dcbor bridge

Expand All @@ -144,14 +154,17 @@ not `u32`s or whose `seqNum` is 0 (`Decoder`), a URL header beyond `u16`
| `decodableFromURString(decodable, s)` | `decodeURWith(UR.parse(s), codec)` |
| `URDecodable`, `URCodable`, `isUREncodable`, `isURDecodable`, `isURCodable` | removed |

## 7. Node and TypeScript floors

Node **22.12** and TypeScript **5.7**. The IIFE / global-script build is
gone; use the ESM or CJS entry.

## 8. What did not change

- Every UR string, QR string, bytewords string and multipart part string.
- Decoder acceptance: the same parts complete the same messages (see
`RUST_DIVERGENCES.md` for where that is more than the Rust reference).
- Error messages.
## 7. Decoding follows the reference's `ur` crate

- Completion: a mixed part is reduced against the fragments known when it
arrives, and a buffered part only when a later simple part comes, so a
shuffled or lossy sequence can need more parts than before; the parts a
message needs when they arrive in order are unchanged.
- The reassembled message is returned without a CRC-32 check against the
parts' checksum field; each part's bytewords checksum is still verified.
- Part CBOR is read as `minicbor` reads it (any head width, trailing bytes
ignored), with minicbor's error texts.
- `MultipartDecoder.add` is case-sensitive and validates every string, also
after completion.
- Every UR string, QR string, bytewords string and multipart part string is
unchanged.
21 changes: 13 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,26 +44,31 @@ for (const part of encoder) {
}
decoder.result?.equals(ur); // true

// A scanner that may see either kind dispatches on the `n-m/` header, as a
// caller of the reference does: `MultipartDecoder` rejects a single-part UR.
const isMultipart = (s: string) => /^ur:[^/]+\/\+?\d+-\+?\d+\//i.test(s);
isMultipart("ur:test/lsadaoaxjygonesw"); // false → UR.parse(s)
// A scanner that may see either kind lower-cases the payload (the multipart
// decoder is case-sensitive, as the reference's) and dispatches on the
// `n-m/` header, as a caller of the reference does: `MultipartDecoder`
// rejects a single-part UR.
const scanned = "UR:TEST/LSADAOAXJYGONESW".toLowerCase();
const isMultipart = (s: string) => /^ur:[^/]*\/\+?\d+-\+?\d+\//.test(s);
isMultipart(scanned); // false → UR.parse(scanned)

// Bytewords.
encodeBytewords(new Uint8Array([1, 2, 3, 4, 5]), "standard"); // "acid also apex aqua arch fuel bald nail work"

try {
UR.parse("ur:test/lsadaoaxjygonese");
} catch (e) {
if (URError.isURError(e)) console.log(e.code); // "Bytewords"
if (URError.isURError(e)) console.log(e.code, e.message); // "Decoder", "UR decoder error (invalid checksum)"
}

// Every failure is a URError; branch with `is`, and read `details` by code.
// Failures are URErrors with the reference's codes and messages; an
// argument outside its domain is `InvalidParameter`. Branch with `is`, and
// read `details` by code.
try {
new MultipartEncoder(ur, 1.5);
} catch (e) {
if (URError.isURError(e) && e.is("InvalidParameter")) {
console.log(e.details.parameter, e.message); // "maxFragmentLength", "… must be an integer in [1, …], got 1.5"
console.log(e.details.parameter, e.message); // "maxFragmentLength", "… must be an integer in [1, …] or a bigint in [1, …], got 1.5"
}
}

Expand All @@ -80,13 +85,13 @@ Runnable examples live in the [`examples/`](https://github.com/BlockchainCommons

### Version History

- **1.0.0-beta.3 (September 14, 2026)** - Decoding follows the reference's `ur` crate in full: fountain completion, no CRC-32 check on reassembly, `seqNum` 0, minicbor part CBOR and messages, a case-sensitive `MultipartDecoder` that validates after completion, `UR.parse` in the reference order with `Decoder` for bytewords failures, the empty UR type, `decodeURWith` on the first tag only with a dcbor `CborError` for a wrong type, ASCII-only `canonicalizeByteword`, frozen tables, argument validation on every export, and a `bigint` `maxFragmentLength`.
- **1.0.0-beta.2 (September 12, 2026)** - Decoding follows the reference: `bytewords` decoding is case-sensitive, `MultipartDecoder` rejects single-part URs and reads the fountain fields from the CBOR (the header is parsed as two `u16`s, a leading `+` allowed), and a zero `maxFragmentLength` reports the reference's decoder error.
- **1.0.0-beta.1 (September 9, 2026)** - Initial beta implementation.

### Roadmap

- Continued testing and auditing on the path from beta to a stable **1.0.0** release.
- Continued parity with the Rust reference implementation as it evolves (see [`RUST_DIVERGENCES.md`](./RUST_DIVERGENCES.md)).

### Dependencies

Expand Down
Loading