Skip to content

[Feature] Python Bridge SDK - #71

Merged
iamalwaysuncomfortable merged 125 commits into
masterfrom
feat/bridge-sdk
Sep 25, 2026
Merged

iamalwaysuncomfortable merged 125 commits into
masterfrom
feat/bridge-sdk

Conversation

@iamalwaysuncomfortable

@iamalwaysuncomfortable iamalwaysuncomfortable commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds bridge-sdk/ — the aleo-bridge-sdk Python package (import aleo_bridge), a port of veil's @provablehq/aleo-bridge-sdk 0.1.0 into the web3.py-style verb structure used by our other Python SDKs. It is a standalone package: no dependency on shield-swap journal/profile/stage abstractions.

  • Registry pinned to 2026-08-31.solana-deposits.1 (7 chains, 19 assets, 22 routes; metadata literal-identical to veil).
  • Encodings + protocol modules: Hyperlane (Aleo/EVM/Solana), xReserve (EVM deposits incl. private mint, Aleo burns), Sealance freeze-list proofs, ARC-20 shield/unshield.
  • Connections: Ethereum(rpc_url | w3, private_key | signer), Solana(rpc_url | client, private_key | signer) (own sync JSON-RPC client; solana-py is async-only), Bridge.from_env() / from_profile().
  • Lifecycle verbs on Bridge: quote → execute → wait / get_status → recover → resume / complete, veil's checkpoint format (version 1 allowlist, secrets excluded), timeout ≠ failure, calls are single-use after a broadcast, a lost RPC response never loses the tx id.
  • Agent surface: aleo_bridge.agent.bridge_tools()/dispatch_tool() (confirm-gated writes, secrets redacted), stdio MCP server (aleo_bridge.mcp), generated AGENTS.md (codegen/gen_context.py --check in CI), python -m aleo_bridge.
  • Live suite mirroring veil's test/integration/live/: same gates (BRIDGE_LIVE_FUNDS, BRIDGE_LIVE_STATE_DIR, BRIDGE_LIVE_MAINNET_ACK, BRIDGE_LIVE_MAINNET_CASES, BRIDGE_LIVE_MAINNET_EXECUTE), recover-first state files, one atomic unit per leg, parametrized over every registry route in both directions; scripts/rehearse.py as the operator CLI.

…mported modules

del sys.modules[...] permanently replaced aleo_bridge's module/class objects for
the rest of the pytest session, breaking isinstance checks in any test file that
runs later alphabetically and compares against classes imported at collection
time. Use monkeypatch.delitem so the original modules are restored at teardown.
…d nullifier reads, registry drift) and README
Add checkpoint.py: Checkpoint (frozen dataclass with to_dict/to_json and
from_dict/from_json round-tripping), create_checkpoint(plan, receipt, registry)
reducing a Receipt to the documented recovery allowlist (intent, route,
source/destination transaction ids, delivery verification) while excluding
keys, record plaintext, secretNonce, attestation bodies, payloads, message
hashes, nonces and quote internals, CheckpointStore protocol, and
FileCheckpointStore (one 0600 file per sanitized receipt id, atomic
temp+os.replace writes).

Export Checkpoint/CheckpointStore/FileCheckpointStore/create_checkpoint from
__init__.py. Update the two test_client.py cases that were placeholders for
plan 4's checkpoint wiring (BRIDGE_CHECKPOINT_DIR, Profile.checkpoint_dir) now
that checkpoint.py exists and client.py's lazy imports resolve.
…e JSON-RPC web3 double

Adds aleo_bridge.eth.Ethereum (rpc_url+private_key, w3+signer, or bare w3
read-only/default-account forms), cached chain_id, send_transaction (local
signer or default-account path), wait_for_receipt/get_receipt, and
Ethereum.from_env. web3/eth_account stay lazily imported so the package
still imports without the evm extra.

Adds tests/fakes/fake_web3.py: a real Web3 over a hand-rolled JSON-RPC
provider so tests exercise real signing/calldata against fake transport
state, reusing tests.conftest.FakeAleo rather than a second Aleo fake.

Updates the two plan-1 placeholder assertions in test_client.py that
anticipated this change (Bridge.eth / Bridge.from_env now construct a
real Ethereum connection instead of raising MissingExtraError).
private_burn's record plaintext (input 0) and private_mint's secret nonce
(input 3) could end up in logs via repr(). Print only the input count;
.inputs stays the explicit accessor for callers that need the values.
… ConfigurationError

Receipt.replace() forwarded protocol_state/next_action straight to
dataclasses.replace() unless the caller overrode them, so a replace() copy
aliased the original's dicts and mutating one leaked into the other.
setdefault() a shallow copy of each first.

to_progress()'s missing-routeId check raised a bare ValueError; every other
failure in this package is a BridgeError subclass, so make this one
ConfigurationError too (same message).
_verified_tree() previously fell back to an unverified (possibly empty)
Merkle proof whenever the on-chain freeze_list_root was unreadable, which
would silently hand out a proof against the wrong tree. Raise
ConfigurationError instead and tell the caller to pass merkle_proof=
explicitly.

freeze_list_program()'s import-lookup also narrowed its except clause from
bare Exception to (ProgramNotFound, AleoError), so a transient RPC failure
of another type propagates instead of silently falling back to the static
freeze-list table.

Updated every fake facade setup that expects a proof to seed
usdcx_freezelist.aleo/freeze_list_root[1u8] with the matching root
(EMPTY_TREE_ROOT for an empty list), and added tests for the unreadable-root
and unrelated-exception-propagates cases.
…e 0700

Profile.load_or_create()'s first-write path used a tmp-file-plus-rename,
which is atomic but not exclusive: two processes racing a fresh
$ALEO_BRIDGE_HOME could each "win", the second silently overwriting the
first's already-in-use key. Open profile.json with O_EXCL instead; on
FileExistsError, load whichever profile got there first instead of
clobbering it. The home directory is now created (or healed) at mode 0700,
matching the file's existing 0600.
…tate

burn() silently ignored record=/merkle_proof= for public/public-as-signer
modes instead of rejecting them, and an unknown mode wasn't caught until
deep inside build_burn_inputs (after route/amount/chain work already ran).
Validate mode first, and reject record=/merkle_proof= for any non-private
mode with a ConfigurationError.

Also de-duplicated the Aleo-chain lookup: XReserveModule._aleo_chain_id()
and HyperlaneModule._aleo_chain() were exact copies of Bridge.aleo_chain();
both modules now call the client's version instead of keeping their own.
- units._DECIMAL_RE: \d -> [0-9] (Python's \d matches any Unicode decimal
  digit, not just ASCII).
- encoding.xreserve_deposit_nonce: bound source_domain to uint32 before
  encoding it, consistent with the payload's remote_domain check.
- registry.Asset.matches_address: re.fullmatch instead of re.search, so a
  trailing newline can no longer sneak past a "$"-anchored regex.
- encoding's hex/shape/width checks (hex_to_bytes, bytes32_to_u128_limbs,
  solana_address_to_hyperlane_recipient, xreserve_deposit_payload's
  per-field width check) now raise InvalidRecipientError, and _uint_be
  raises InvalidAmountError, instead of bare ValueError — all BridgeError
  subclasses, same remedy messages. Updated the callers (circle.py,
  hyperlane.py, xreserve.py) that caught the old ValueError type, and the
  tests that asserted on it. xreserve.get_attestation("nope") now raises a
  BridgeError instead of a naked one.
- client.checkpoints_from_env: removed the dead ImportError branch (and its
  "plan 4" wording) now that checkpoint.py always exists.
- client.balance_program: one-line comment recording that arc20_<sym>.aleo's
  balances mapping is verified to be what the warp/xreserve programs'
  mint_public/burn_public spend.
- hyperlane.py: removed the unused Callable import.
…and checkpoint-before-poll ordering

- fake_web3._decode_raw now keeps nonce/gas/gasPrice/maxFeePerGas/maxPriorityFeePerGas from the
  signed raw tx instead of discarding them, so tests can assert send_transaction's exact fee-filling
  arithmetic (gas = estimate * 1.2; EIP-1559 maxFeePerGas = baseFee * 2 + tip).
- FakeRpcProvider gains a legacy=True knob that omits baseFeePerGas from eth_getBlockByNumber,
  exercising eth.py's previously-unreachable legacy gasPrice branch.
- FakeRpcProvider gains receipt_delay/receipt_poll_counts so a receipt can stay pending for N polls
  per hash; test_evm_call's ordering test now proves each checkpoint fires with zero receipt polls
  for its own hash, that polling only starts after, and that approvals fully confirm before the
  main call is broadcast.

No production code changes: eth.py's fee formula and _calls.py's checkpoint-before-poll ordering
were both already correct; the gap was in test strength only.
…nd fix three stale README lines

I5: both literal acknowledgement values in the env-var table (and the one left in
docs/veil-brief.md) become the <see tests/live/config.py> placeholder the rest of the README
already uses. m2: BRIDGE_LIVE_MAINNET_CASES lists the five real case names instead of
'leg-3,leg-5'. m1: pending() is offline, from the checkpoint store. m3: the wait() signature
shows on_error and max_consecutive_errors.
execution_allowed() demands the mainnet acknowledgements only on mainnet — testnet is gated
on BRIDGE_LIVE_FUNDS/BRIDGE_LIVE_STATE_DIR alone, so a recovery run can actually submit the
step a committed transfer is waiting on. --recover now builds its client the way the live
suite does, per environment (helpers.build_bridge, shared with the suite), instead of a bare
Bridge.from_env() that is always the mainnet account.
…f relabelling it 'recover'

lifecycle.resume refuses a private-mint resume with no secret_nonce before any RPC, but the
agent routed that refusal through the write path, so the model was told next: 'recover' —
'the deposit may already be on the wire' — for a transfer that never left the process.
_h_resume now checks it up front like _h_complete and answers with the actionable
configuration error.
… wheel

The bridge smoke piped 'python -m aleo_bridge' into head, which exits 0 even when the wheel
ships no AGENTS.md and the command prints nothing; it now greps the first line for
'# aleo-bridge'. tests/test_package.py asserts that grep exists and that the hatch wheel
config still packages the directory AGENTS.md lives in, with nothing excluding it.
…fee; pending failure shapes aligned

R1: an Aleo-origin xReserve burn pays the withdrawal fee OUT OF the burned amount, so the
delivery baseline execute records is the quote's amount_out (amount_atomic - the route's
withdrawalFeeAtomic), not the full amount. Waiting for the full amount was wrong in both
directions: the real transfer never satisfies it (2.000001 USDCx burned delivered 996501
atomic USDC on testnet), and unrelated inflow to the same address could - a false COMPLETED
that deletes the checkpoint. _delivery_verification takes an expected_atomic override; the
Hyperlane branch keeps the full amount, and a route with no readable fee records no baseline
at all rather than a wrong one. Branch 8's predicate is unchanged.

R2: BridgeStatus.pending is annotated list[Progress | dict] and status()'s docstring
describes the dict failure entries. R3: bridge_pending's failure entries carry
'next': 'failed' like Bridge.pending()'s. R4: the README says where the acknowledgement
strings live rather than claiming they appear nowhere in the repository.
…r Aleo-origin ids

The explorer indexes EVM hashes and Solana signatures; an Aleo at1… id is bech32, so
the base58 decode raised on the mainnet aleo/eth->ethereum/eth return leg after the SDK
had already delivered it (destination balance +1 wei). veil never queries the explorer
for Aleo-origin legs either; the SDK's own delivery check is the verdict.
…ination id

The SDK proves delivery for Aleo-origin legs by the recipient's balance rising (branch 6
for Hyperlane, branch 8 for xReserve) and records no message id; veil's aleo-hyperlane
live test asserts only completed + sourceTxId. The mainnet aleo/eth->ethereum/eth leg
completed with +1 wei on Ethereum and tripped the stricter assertion.
…ansaction detection

A mainnet run on 2026-09-22 lost a WBTC Hyperlane dispatch: the public RPC
suggested a zero priority tip, the dispatch sat unmined, the RPC stopped
reporting it as pending, and the next leg's approval was handed the same nonce
and replaced it. source_status then held the dead hash at SOURCE_CONFIRMING
forever.

- Ethereum.send_transaction floors the EIP-1559 tip at MIN_PRIORITY_FEE_WEI
  (0.1 gwei), overridable with Ethereum(..., min_priority_fee_wei=...), and
  fills the nonce as max(pending count, last nonce broadcast + 1) — reserved
  before the broadcast, so even an ambiguous send keeps its nonce.
- Every EVM broadcast records its nonce in protocol_state['sourceNonce'], which
  create_checkpoint now carries (version stays 1; absence tolerated).
- source_status turns a checkpointed hash the node has forgotten, whose nonce
  the account has moved past, into EXPIRED with sourceError/dropped.
  recover_source stays history-first: a dispatch found in the logs wins, and
  only a scan that ran and matched nothing makes a dropped hash resumable
  (SOURCE_SUBMISSION_PENDING); an unscannable one stays EXPIRED.
The live run's recovery scan died on publicnode with -32602 'block range
extends beyond current head block' for blocks 26034459-26034469 — a range its
own eth_blockNumber had just handed out. The endpoint is load-balanced, so the
node answering eth_getLogs was behind the node that reported the head.

- eth.HEAD_RACE_RE matches only head-block phrasings ('beyond current head',
  'head block', 'exceeds ... head'), so a genuine range-too-large or
  too-many-results error still fails loudly.
- _scan_logs retries such a chunk up to LOG_SCAN_HEAD_RACE_RETRIES (5) times,
  pausing LOG_SCAN_HEAD_RACE_SLEEP_SECONDS (0.5, via the injectable
  EthModule.sleep) and re-reading the head each time. The scan's upper bound
  only ever moves DOWN to the answering node's head, which also clamps the
  initial one. Exhausting the retries still raises BridgeError: a scan that
  cannot finish must never read as an empty history.
- lifecycle._is_transient_error matches the same message, so wait() polls again
  instead of aborting the transfer.
…its checkpoint, and nonces are process-wide

Scoped review of the mainnet fix, all items:

- The nonce high-water mark moves to a module-level dict keyed by (chain id,
  checksummed sender). The live suite builds a fresh Ethereum per phase and per
  leg — exactly the incident's topology — so a per-instance counter reset
  precisely where the replacement happened. Per process, not across processes
  or machines; last_broadcast_nonce stays per instance.
- A dropped verdict now needs TWO eth_getTransactionByHash misses, separated by
  DROPPED_REPROBE_SLEEP_SECONDS, so one lagging backend of a load-balanced
  endpoint cannot declare a live transaction dead. The verdict carries the head
  it was taken at, and recover_source only calls a transfer resumable when the
  history scan ran AND its (clamped) head reached that far; otherwise EXPIRED
  with 'retry recover()'.
- A terminal receipt carrying dropped is kept in the checkpoint store, not
  deleted (lifecycle._keeps_checkpoint, honoured by _persist and _finish); the
  pair is allowlisted into the checkpoint so an offline pending() listing shows
  it as failed-and-explained instead of still confirming.
- The sourceError says what to do: with approvals, recover() then resume(); with
  none, inspect the sender's transactions at that nonce — a fresh execute is
  required. Every EXPIRED one ends with 'the checkpoint is kept'.
- min_priority_fee_wei is a validated property, also read from
  BRIDGE_MIN_PRIORITY_FEE_WEI by from_env.
- HEAD_RACE_RE is word-bounded, so 'max header size' and 'overhead block cache'
  are not head races.
…persedes the dropped record

Re-review of the round-3 hardening:

- _recover_hyperlane/_recover_xreserve take the dropped verdict (both probes,
  latest nonce, head_at_probe) BEFORE the log scan. The chain head is monotone,
  so a scan that runs afterwards covers that head by construction, and the gate
  stays sound: the nonce was already consumed at head_at_probe, so any
  replacement mined at or below it and the scan saw it. Taking the verdict
  afterwards made a chain that simply moved on look like a scan that had
  stopped short.
- _scan_logs returns (logs, covered_head) and the history scans return
  (match, covered_head); the _last_scan_head instance attribute is gone.
- A successful resume seeds its emitter with the checkpoint it started from, so
  the new source transaction's record supersedes the kept dropped one instead of
  leaving a failed ghost beside it; _finish's keep-branch supersedes its
  checkpoint_id the same way when the id differs.
- When the scan really did cover less than the verdict head, the sourceError
  says only: retry recover(); the checkpoint is kept — never 'then resume()'.
- README: how to clear a stale dropped record with store.delete(id).
…ts as dropped

The live WBTC resume run did not reach the dropped verdict. The load-balanced
public RPC still answers eth_getTransactionByHash for the replaced dispatch —
returning the transaction with blockNumber null, a stale mempool view — while
eth_getTransactionReceipt is not-found and the account's latest nonce (85) has
passed the dispatch's (83). The verdict required not-found on both probes, so
the receipt stayed SOURCE_CONFIRMING and wait polled to its ceiling.

A probe now counts as 'gone' when the node either does not know the hash OR
returns it with no block number: both mean it is in no block, and together with
a consumed nonce that is final — it can never be included. A transaction that
comes back WITH a block number is mined and never dropped, however far behind
its receipt read is. The double probe, the nonce read between them and
head_at_probe are unchanged.
…ict once (recover, then resume)

On mainnet the WBTC dispatch was replaced at its nonce; the SDK eventually marked it
EXPIRED/dropped with the advice to recover() then resume(), and the harness raised on
next == 'failed' instead of following it. It now follows exactly once; any other failure
stays final and execute is never called.
…e with the probe's own nonce and sender

Every read behind the verdict can come from a backend that has not imported the
block our own transaction mined in: the single receipt read misses, both probes
serve the transaction from a stale mempool with a null blockNumber, and the
latest nonce read lands on a fresh backend where that very transaction consumed
the nonce. The result was a false EXPIRED whose sourceError tells an operator
that no funds moved, for a transfer in flight.

Two strictly narrowing checks close it. The receipt is re-read once after the
second probe, and a receipt that has appeared withdraws the verdict. Where a
probe served the transaction object, its own nonce and from must match the
checkpoint, or the checkpoint does not describe this hash and no verdict is
issued. Both can only withhold a verdict, never produce one.

The "a consumed nonce is final" claim in the docstring and the README is
qualified as holding for a node that has imported the block that consumed it.
`quote`, `prepare`, `routes` and the agent/MCP tools now take veil's
transfer vocabulary, flattened and snake-cased: keyword-only
`source_chain`, `source_asset`, `destination_chain`, `destination_asset`
and `bridge_protocol` (veil `quote`'s `source: {chain, asset}`,
`destination: {chain, asset}`, `bridgeProtocol`), or `route=` with a
`Route` from `routes()` or its id. The positional `"chain/key"` pair is
gone from the public verbs; `Route.protocol` / `Plan.protocol` stay as
data fields. `destination_asset` and `bridge_protocol` are optional when
the remaining selectors leave one route; an ambiguous match names the
keyword that would separate the candidates.

`Bridge.routes(...)` is new (veil `getRoutes`, scoped to the client's
environment). `Registry.routes` / `find_route` use the same keywords, as
do the internal lookups in eth/hyperlane/xreserve/sol; `recover` rebuilds
a plan from a checkpoint intent through the same path. README, the
codegen quickstart and both AGENTS.md copies regenerated.

707 hermetic tests; pyright unchanged at the pre-existing 33.
@iamalwaysuncomfortable iamalwaysuncomfortable changed the title bridge-sdk: aleo-bridge-sdk Python package (lifecycle verbs, agent tools, live suite) [Feature] Python Bridge SDK Sep 25, 2026
@iamalwaysuncomfortable
iamalwaysuncomfortable merged commit f360c38 into master Sep 25, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant