Skip to content

[feat] Distinguish installed, runtime-verified, and workload-verified engine readiness #42

Description

@willsarg

Problem

ARA currently has several different levels of engine evidence, but the user-facing readiness signal does not distinguish them clearly.

Today:

  • ara.engines.is_installed() treats an engine as installed-ready when its isolated interpreter exists and its ARA version/schema stamps match.
  • ara detect uses that cheap result for engine_ready and renders the engine as ready.
  • A fresh or refreshed ara install runs engine_audit.audit_engine(), which can discover that the installed build or runtime is mismatched, but package installation still remains the primary success state.
  • ara doctor --engines exposes richer installation, build, runtime, and workload dimensions.
  • Successful characterization is already bound to methodology and engine fingerprints in durable evidence.

An environment can therefore be present and current according to stamps while its accelerator backend is missing, its device operation fails, or it has never been tested with a model. Calling all of those states ready collapses distinct facts and can overstate what ARA has observed.

ARA should expose three separate concepts:

  1. Installed — the isolated environment exists and its ARA/schema stamps are current.
  2. Runtime-verified — an explicit no-model audit verified the expected backend and a minimal device/runtime operation for the exact engine build and machine.
  3. Workload-verified — a governed characterization succeeded for a model under the exact current methodology and engine fingerprint.

Public starting points

  • ara/engines.py
    • is_installed()
    • engine version/schema stamps and install results
  • ara/engine_env.py
    • isolated environments and bounded no-model probes
  • ara/engine_audit.py
    • audit_engine()
    • installation/build/runtime/workload dimensions
    • engine fingerprints and characterization evidence
  • ara/detect.py
    • construction of Machine.engine_ready
  • ara/serialize.py
    • the durable/machine-readable projection of readiness
  • ara/cli.py
    • render_install(), render_doctor(), detect rendering, and command admission paths
  • ara/db.py
    • characterization evidence and any appropriate home for durable audit authority
  • Tests likely to change:
    • tests/test_engines.py
    • tests/test_engine_audit.py
    • tests/test_detect.py
    • tests/test_serialize.py
    • tests/test_cli.py

Desired behavior

Define a machine-readable readiness model that preserves these distinctions instead of deriving a single ready claim from installation stamps.

A reasonable result shape may expose an explicit state plus supporting facts, for example:

  • absent
  • installed_unverified
  • runtime_verified
  • workload_verified
  • stale or mismatch where appropriate

The exact public shape should be chosen to fit ARA's existing JSON contracts, but it must be possible for callers to distinguish package presence, live runtime verification, and model-workload verification without parsing prose.

Runtime verification is allowed only during an explicit action that already crosses that boundary, such as ara install or ara doctor --engines. Read-only recon must not initialize MLX, CUDA, Vulkan, llama.cpp, or any model.

If runtime verification is persisted for later read-only display, it must be durable evidence rather than a timeless boolean. At minimum it must be bound to:

  • the exact engine fingerprint;
  • the current machine identity;
  • the audit method/schema; and
  • enough timing/provenance information to explain what was observed.

Any engine, package, source, schema, hardware, or audit-method change that invalidates that authority must downgrade the state rather than silently retaining runtime_verified.

Acceptance criteria

  • Installation presence, runtime verification, and workload verification have separate structured representations.
  • ara detect remains read-only and engine-free:
    • it does not import or initialize an installed engine;
    • it does not run a device operation;
    • it does not load or download a model.
  • An installed engine with no current explicit audit is reported as installed but unverified, not simply ready.
  • A failed or mismatched runtime audit cannot be rendered or serialized as runtime-ready.
  • A successful runtime audit is tied to the exact current engine fingerprint and machine identity.
  • A successful characterization is the only route to workload-verified status for that model/runtime cell.
  • Changing the engine fingerprint, ARA/schema version, relevant hardware identity, or audit schema invalidates previously persisted runtime verification.
  • Existing installations without the new evidence migrate safely by becoming installed-unverified; no unobserved verification is manufactured.
  • Text output explains the state and the next relevant action without conflating installation with model measurement.
  • JSON output exposes the distinction without requiring callers to parse display strings.
  • Command admission continues to fail closed wherever current characterization evidence is required.
  • Tests cover at least:
    • absent engine;
    • installed/current but never audited;
    • successful runtime audit;
    • runtime mismatch or probe failure;
    • stale audit after fingerprint change;
    • characterized workload under the exact fingerprint;
    • stale workload evidence;
    • read-only detect proving that no engine probe ran.
  • Documentation describing detect, install, doctor --engines, and characterization uses the new terms consistently.
  • The focused tests and the full 100% statement/branch coverage gate pass.

Constraints and non-goals

  • Preserve the frozen public command tree; this issue changes evidence/state semantics, not command names.
  • Preserve explicit consent boundaries. detect, status, and profile must not begin loading engine runtimes.
  • Preserve the pure-Python, engine-free core and subprocess adapter boundary.
  • Do not treat a scalar runtime audit as proof that a model fits or that a characterization remains valid.
  • Do not require a model download for runtime verification.
  • Avoid an unbounded generic plugin/capability framework. This issue is about truthful state for the engines ARA already ships.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions