Skip to content

Persist schedule identities and roll TESLA chains without stale-key reuse #71

Description

@ctfbruce

Problem

The default chain now covers seven days, but exhausted schedules still clamp to the last key and there is no rollover. The dispatcher keeps disclosed keys in memory indexed only by executor ID and epoch, so a reboot/new schedule for the same executor can collide with old immutable entries and historical evidence depends on live memory.

Proposed change

Give each schedule an authenticated identity and durable anchor/disclosure history, with safe rollover and restart rules that never reuse disclosed signing keys.

Acceptance criteria

  • Implement bounded rollover/overlap or stop attribution before exhaustion; use a short test chain to cover the boundary without multi-day tests.
  • Persist sufficient schedule identity/history to distinguish successive boots and chains; restarting with the same seed must not reset authority onto already disclosed keys.
  • Test executor-only restart, dispatcher restart, rollover, duplicate/reordered disclosures and stale key-map updates with retained captures.
  • Expose expiry/unavailable status and export historical verification material; define retention/backup integration without treating a copied seed as a complete recovery design.

Current evidence

confirmed finite-chain/history limitation; source inspection only, no new runtime reproduction. Backlog classification is based on source inspection; this issue does not claim a new runtime reproduction.

Dependencies

  • #69 — Stop using the public TESLA anchor as the first signing key
  • #70 — Specify and version packet-tag algorithms and verification limits

Related work

These are integration points, not prerequisites for starting this issue.

  • #90 — Enroll executor identities and bind them to both control channels
  • #18 — Define retention, export and deletion behavior for measurement data
  • #111 — Create consistent versioned backups of operator state
  • #112 — Restore a verified backup into separate state without replaying measurements

Priority context

Required before accepting untrusted users or advertising the corresponding protected capability; not a blocker for the trusted local TEST alpha.

Validation must use owned local fixtures on supported Linux environments. Record the implementing merge request and relevant test results before closing this issue.


Imported from GitLab issue 70. Originally opened 2026-09-10. Historical GitLab links may require access to the original project.

Activity

  1. added
    bugSomething isn't working
    P2Follow-on improvement or optional expansion after core requirements.
    Source reviewedCurrent source evidence checked; no new runtime reproduction claimed by backlog creation.
    Traffic policy and attributionConsistent traffic policy, aggregate budgets and evidence-backed packet attribution.
    on Sep 24, 2026
  2. vincent10400094 commented on Sep 25, 2026

    @vincent10400094
    Member

    Progress: #203 and #204, merged into dev via #228 (e4a621f). #203 records a new chain on every executor start, so the same seed does not bring back disclosed keys. #204 stops tagging when the chain runs out and keeps dispatcher disclosures per chain anchor. Still open: rolling to a new chain without a restart, a durable dispatcher key store (a dispatcher restart loses every disclosed key), capture-based restart and reorder tests, exposing expiry status, and historical export and retention. Wiping the executor database also resets the generation.

  3. vincent10400094 commented on Sep 29, 2026

    @vincent10400094
    Member

    Plan (agreed 2026-09-29), split into three PRs:
    (a) Durable disclosed-key history plus a dated run lookup: "which runs were active from this IP at time t" instead of the latest N. This comes first, because offline verification of older captures needs it.
    (b) Operator-signed schedule parameters {k_0, t0, I, d}, so an offline verifier doesn't have to trust the dispatcher's word.
    (c) Rollover with an overlap window.
    The disclosure delay becomes configurable (d ≥ 2, default ≈ 15 min) in the fix/tesla-disclosure-delay PR.

  4. vincent10400094 commented on Sep 29, 2026

    @vincent10400094
    Member

    Add to (a): disclose an old chain's tail after an executor restart. With d ≈ 15 min (#343), the last d epochs of a chain can never be verified if the executor restarts, because keys live only in memory and every start builds a new chain.

    What (a) needs for this:

    • Persist the chain state, or a sealed seed, so the tail can be re-derived after a restart.
    • Carry the old anchor and schedule in the hello.
    • Let the dispatcher accept disclosures for non-current chains, into the durable store.

    Until then, #343 logs final_disclosure_at and documents that operators should restart only after it.

  5. ctfbruce commented on Oct 7, 2026

    @ctfbruce
    CollaboratorAuthor

    PR #378, merged as ca27662. All 16 required jobs passed on the exact candidate and merged main.

    With a configured seed, executor schema 8 permits disclosure-only recovery of the immediately previous chain. Restart checks refuse recovery while a prior signer remains, and recovered disclosure/retention timing uses monotonic elapsed time.

    PR #382, merged as 072ab00. All 16 required jobs passed on the exact candidate and merged main. The retained-capture regression checks executor and dispatcher database reopen, duplicate/reordered disclosures and verification using retained history.

    This remains open for authenticated schedule/run metadata and the associated recovery/retention acceptance. Exhaustion already stops attribution, an allowed alternative to seamless rollover. The current tail increment covers only the immediately previous generation and does not claim arbitrary backup rollback or durable acknowledgement of dispatcher key storage.

  6. ctfbruce commented on Oct 7, 2026

    @ctfbruce
    CollaboratorAuthor

    The remaining authenticated-history criteria are implemented in #409. An enrolled executor signs its exact schedule with its TLS identity; the dispatcher binds that proof to the authenticated control peer and stores it immutably. Dated run lookups are separately signed, and exported history verifies after shutdown using independently obtained dispatcher keys and executor certificate pins. Export and trust contract.

    This builds on #203/#204’s fresh generation and exhaustion stop, #378’s disclosure-only recovery of the immediately previous chain, and #382’s retained-capture restart/reordered-disclosure regression. New tests cover proof tampering, enrollment binding, immutable reannouncement and offline export after shutdown.

    Closure uses the issue’s permitted stop at exhaustion alternative. Seamless rollover is not implemented. Recovery requires the configured seed, intact generation history and retired prior signers; it does not support arbitrary backup rollback or promise durable dispatcher acknowledgement of tail storage. Old unsigned history stays unsigned. Independent measured clock acceptance remains open in #72.

    Merged as a9d865d. All 17 checks passed on the exact candidate (guest source check) and the identical merged main (guest source check).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P2Follow-on improvement or optional expansion after core requirements.Source reviewedCurrent source evidence checked; no new runtime reproduction claimed by backlog creation.Traffic policy and attributionConsistent traffic policy, aggregate budgets and evidence-backed packet attribution.bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions