Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .changeset/llm-host.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
---
"@azphalt/llm-host": minor
---

New package: the host side of `kind: "llm"` (`spec/llm.md`).
- `llmConsent` gives what a host must show before install.
- `installLlm` verifies the package and refuses a public sandbox. It commits the payload and runner workflow to a private GitHub Actions sandbox in one commit, stores keys as sealed Actions secrets, and runs the one-time setup.
- `runLlm` implements the `github-actions-runner` protocol: dispatch, check-run progress, and the bounded `azphalt-llm-result` artifact, with resume by correlation id.
- `chatLlm` implements `openai-chat`.
- The rolling-delimiter helpers are byte-compatible with the reference runner.
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,7 @@ LICENSE Apache-2.0 (with NOTICE)
| [`@azphalt/registry`](packages/registry) | Verify · index · version · serve · search, the consignment marketplace overlay, and the Stripe Connect charge + onboarding surfaces. |
| [`@azphalt/registry-store-vercel`](packages/registry-store-vercel) | Legacy/alternate Neon + Vercel Blob `RegistryStore` implementation. It remains reusable, but the flagship storefront no longer depends on Vercel. |
| [`@azphalt/repository-client`](packages/repository-client) | Client SDK for the Repository API. |
| [`@azphalt/llm-host`](packages/llm-host) | Host side of `kind: "llm"`: consent, sandbox install, runner and `openai-chat` protocols, rolling delimiters. |
| [`@azphalt/mcp`](packages/mcp) | An MCP server exposing `.azp` verify/inspect/extract to any MCP host. |
| [`@azphalt/web-handoff`](packages/web-handoff) | The web→host install handoff from `spec/web-handoff.md`: build the `azphalt://install` link, attempt it, and fall back when no host claims it. |
| [`create-azphalt`](packages/create-azphalt) | Scaffolder for a new extension package. |
Expand Down
1 change: 1 addition & 0 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,6 +91,7 @@ The native host that embeds the engine and renders the schema is **each app's ow
registry/ verify · index · version · serve · search, plus the consignment overlay
registry-store-vercel/ alternate Neon + Vercel Blob RegistryStore implementation
repository-client/ client SDK for the Repository API
llm-host/ host side of kind: llm (sandbox install, runner, delimiters)
mcp/ an MCP server exposing azp verify/inspect/extract to any MCP host
create-azphalt/ scaffolder for a new extension package
web-handoff/ the web → host install handoff (spec/web-handoff.md)
Expand Down
4 changes: 2 additions & 2 deletions docs/specs/llm.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@ the device**: a hosted OpenAI-compatible endpoint, or open weights the host runs
GitHub Actions sandbox. Modeled on `kind: "mcp"` (mcp-server.md): the package is a signed header plus
a bundled setup script, and the host runs everything **outside the user's device** under consent. The
SDK types (`@azphalt/azdk` `LlmManifest`), the verifier (`validateLlmManifest`, § Verification) and a
reference runner (the first-party `com.hereliesaz.azphalt.llm.*` packages) exist; no conformance
profile does yet.*
reference runner (the first-party `com.hereliesaz.azphalt.llm.*` packages) and a reference host
(`@azphalt/llm-host`) exist; no conformance profile does yet.*

## Why this exists — and why it doesn't break the moat

Expand Down
36 changes: 36 additions & 0 deletions packages/llm-host/package.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
{
"name": "@azphalt/llm-host",
"version": "0.0.0",
"publishConfig": {
"access": "public"
},
"description": "Host side of azphalt kind:\"llm\" packages: consent data, GitHub Actions sandbox install, runner dispatch, and rolling delimiters.",
"license": "Apache-2.0",
"author": "Az (https://hereliesaz.com)",
"type": "module",
"main": "./dist/index.js",
"types": "./dist/index.d.ts",
"exports": {
".": {
"types": "./dist/index.d.ts",
"default": "./dist/index.js"
}
},
"files": [
"dist"
],
"scripts": {
"build": "tsc -p tsconfig.json",
"test": "vitest run",
"typecheck": "tsc -p tsconfig.json --noEmit",
"clean": "rm -rf dist"
},
"dependencies": {
"@azphalt/azdk": "workspace:*",
"@azphalt/azp": "workspace:*",
"blakejs": "^1.2.1",
"fflate": "^0.8.3",
"tweetnacl": "^1.0.3"
},
"sideEffects": false
}
62 changes: 62 additions & 0 deletions packages/llm-host/readme.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# @azphalt/llm-host

The host side of azphalt [`kind: "llm"`](../../spec/llm.md) packages: what to show before install, installing into a private GitHub Actions sandbox, talking to the model over either protocol, and rolling delimiters.

Nothing here runs a package's setup on the device. Setup and the runner run in the sandbox; this library only calls the GitHub REST API and, for `openai-chat`, the model endpoint.

## Install a package

~~~ts
import { llmConsent, installLlm } from "@azphalt/llm-host";
import { readAzp } from "@azphalt/azp";

const consent = llmConsent(readAzp(azpBytes).manifest);
// Show consent.runs / endpointHost / dataHandling / modelLicense / setupTokenPermissions /
// inputs / fetches, and collect the inputs, before going further.

const { install, setup } = await installLlm({
github: { token: setupToken }, // used once for this call; drop it afterwards
owner: "me", repo: "azphalt-llm", // a private repository used only for llm sandboxes
createIfMissing: true,
azp: azpBytes,
inputs: { providerKey: "…" }, // stored as sealed Actions secrets, never committed
onProgress: (seq, message) => console.log(seq, message),
});
// Persist `install`. `setup.status` is "completed" once the sandbox answered a smoke test.
~~~

The package is verified before anything is written. The sandbox must be private. The payload lands in `llm/<package id>/`, the runner workflow in `.github/workflows/azphalt-llm-<package id>.yml`, in one commit, so several packages can share one sandbox.

## Run a task

~~~ts
import { runLlm, newSessionKey, sessionTags, wrap } from "@azphalt/llm-host";

const sessionKey = newSessionKey(); // one per conversation
const tags = await sessionTags(sessionKey, turn); // every tag up to this turn
const result = await runLlm({
github: { token }, install,
task: {
sessionKey, turn,
messages: [{ role: "user", content: `Summarize this issue: ${wrap(untrustedIssueText, tags)}` }],
},
onProgress: (seq, message) => …,
});
~~~

Untrusted material goes only inside `wrap()`. The runner turns each wrapped segment into its own `user` message, so the model never sees a tag, and a result whose text contains any session tag comes back as `failed`. `result.json` is read from the `azphalt-llm-result` artifact with the spec's reference bounds (4 MB zipped, 2 MB JSON); every field is untrusted model output.

After a restart, pass `resume: true` with the same `task.correlationId` to follow the run you already dispatched instead of starting another.

## Call an endpoint directly

For packages that declare `openai-chat`, `chatLlm({ manifest, key, messages, sessionKey, turn })` calls `{baseUrl}/chat/completions` itself, doing the same delimiter translation and output check on the host.

## API

- `llmConsent(manifest)`: the disclosure a host must show before install.
- `installLlm(options)`: verify, commit, seal secrets, and run setup.
- `runLlm(options)`, `findRun(github, install, correlationId)`, `parseResult(bytes)`: the `github-actions-runner` protocol.
- `chatLlm(options)`: the `openai-chat` protocol.
- `newSessionKey()`, `turnTag()`, `sessionTags()`, `wrap()`, `scrub()`, `translate()`, `containsSessionTag()`: rolling delimiters, byte-compatible with the reference runner.
- `sealedBox(message, publicKey)`: libsodium `crypto_box_seal`, as GitHub requires for secrets.
59 changes: 59 additions & 0 deletions packages/llm-host/src/chat.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
/**
* The `openai-chat` protocol (`spec/llm.md` § Protocols): the host calls the endpoint directly and
* is itself the trusted translator of § Rolling delimiters.
*/
import type { Manifest } from "@azphalt/azdk";
import { llmBlock } from "./consent.js";
import { containsSessionTag, sessionTags, translate, type ChatMessage } from "./delimiters.js";

export interface ChatOptions {
manifest: Manifest;
messages: ChatMessage[];
/** The bearer key from the package's `authInput`, when the user gave one. */
key?: string;
sessionKey?: string;
turn?: number;
/** Overrides `endpoint.defaultModel`. */
model?: string;
maxTokens?: number;
temperature?: number;
fetch?: typeof fetch;
signal?: AbortSignal;
}

export interface ChatResult {
text: string;
inputTokens?: number;
outputTokens?: number;
}

export async function chatLlm(opts: ChatOptions): Promise<ChatResult> {
const llm = llmBlock(opts.manifest);
const { endpoint } = llm;
if (!endpoint.protocols.includes("openai-chat") || !endpoint.baseUrl) throw new Error(`${opts.manifest.id} does not offer openai-chat`);
if (endpoint.auth === "required-bearer" && !opts.key) throw new Error(`${opts.manifest.id} needs a key`);
const model = opts.model || endpoint.defaultModel;
if (!model) throw new Error("no model named and the package declares no defaultModel");

const tags = opts.sessionKey ? await sessionTags(opts.sessionKey, opts.turn ?? 0) : [];
const messages = translate(opts.messages, tags);
const res = await (opts.fetch ?? fetch)(endpoint.baseUrl.replace(/\/$/, "") + "/chat/completions", {
method: "POST",
headers: { "content-type": "application/json", ...(opts.key ? { authorization: `Bearer ${opts.key}` } : {}) },
body: JSON.stringify({
model,
messages,
...(opts.maxTokens !== undefined ? { max_tokens: opts.maxTokens } : {}),
...(opts.temperature !== undefined ? { temperature: opts.temperature } : {}),
}),
signal: opts.signal,
});
if (!res.ok) throw new Error(`model endpoint answered HTTP ${res.status}`);
const body = (await res.json()) as {
choices?: { message?: { content?: string | null } }[];
usage?: { prompt_tokens?: number; completion_tokens?: number };
};
const text = body.choices?.[0]?.message?.content ?? "";
if (containsSessionTag(text, tags)) throw new Error("output contained a session tag; rejected");
return { text, inputTokens: body.usage?.prompt_tokens, outputTokens: body.usage?.completion_tokens };
}
59 changes: 59 additions & 0 deletions packages/llm-host/src/consent.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
/**
* What a host MUST show before installing a `kind:"llm"` package (`spec/llm.md` § Discovery and
* § Setup): where prompts go and what the operator does with them, the weights' licence, the setup
* token's permissions, which inputs become sandbox secrets, and every URL setup downloads.
*/
import type { LlmDataHandling, LlmManifest, Manifest } from "@azphalt/azdk";

export interface LlmConsent {
packageId: string;
name: string;
tier: LlmManifest["tier"];
/** `sandbox`: prompts stay in the user's private runner. `third-party`: they go to [endpointHost]. */
runs: "sandbox" | "third-party";
endpointHost?: string;
defaultModel?: string;
dataHandling?: LlmDataHandling;
modelLicense?: { spdx?: string; commercialUse?: boolean; url?: string };
/** Total bytes of weights the sandbox downloads and caches. */
weightsBytes: number;
/** Permissions the one-time setup token needs. */
setupTokenPermissions: string[];
/** Inputs the host prompts for, and whether each becomes a sandbox secret. */
inputs: { id: string; description?: string; optional: boolean; secretName?: string }[];
/** Every URL setup downloads, weights included. */
fetches: string[];
/** The runner job's permission block. */
runPermissions: Record<string, string>;
}

/** Throws unless [manifest] is a `kind:"llm"` package. */
export function llmBlock(manifest: Manifest): LlmManifest {
if (manifest.kind !== "llm" || !manifest.llm) throw new Error(`${manifest.id} is not a kind:"llm" package`);
return manifest.llm;
}

export function llmConsent(manifest: Manifest): LlmConsent {
const llm = llmBlock(manifest);
const secretFor = new Map((llm.setup.secrets ?? []).map((s) => [s.input, s.name]));
return {
packageId: manifest.id,
name: manifest.name,
tier: llm.tier,
runs: llm.tier === "sandbox-weights" ? "sandbox" : "third-party",
endpointHost: llm.endpoint.baseUrl ? new URL(llm.endpoint.baseUrl).host : undefined,
defaultModel: llm.endpoint.defaultModel,
dataHandling: llm.dataHandling,
modelLicense: llm.weights?.modelLicense,
weightsBytes: (llm.weights?.files ?? []).reduce((n, f) => n + (f.byteSize ?? 0), 0),
setupTokenPermissions: llm.setup.requires?.githubToken ?? [],
inputs: (llm.inputs ?? []).map((i) => ({
id: i.id,
description: i.description,
optional: i.optional === true,
secretName: secretFor.get(i.id),
})),
fetches: [...(llm.setup.fetches ?? []).map((f) => f.url), ...(llm.weights?.files ?? []).map((f) => f.remoteUrl)],
runPermissions: { ...(llm.run?.permissions ?? {}) },
};
}
117 changes: 117 additions & 0 deletions packages/llm-host/src/delimiters.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,117 @@
/**
* Rolling delimiters (`spec/llm.md` § Rolling delimiters).
*
* The host wraps each untrusted segment of turn `n` in `⟦tag_n⟧ … ⟦/tag_n⟧`, where
* `tag_n = base32(HMAC-SHA256(sessionKey, "azphalt-llm-turn:" || n))[0:26]`. Material cannot forge a
* tag it cannot compute, so it cannot pose as instructions. A trusted translator — this module for
* `openai-chat`, the sandbox runner for `github-actions-runner` — turns tagged segments into separate
* `user` messages, so tags never reach the model; output containing any session tag is rejected.
*
* Must stay byte-for-byte compatible with the reference runner (`scripts/llm-sandbox/run.py`).
*/

export interface ChatMessage {
role: "system" | "user" | "assistant";
content: string;
}

const B32 = "ABCDEFGHIJKLMNOPQRSTUVWXYZ234567";
// Chat-template control tokens and role markers of the common open-weight families (as run.py).
const NATIVE_MARKERS = /<\|[A-Za-z0-9_]{1,40}\|>|\[\/?INST\]|<<\/?SYS>>|<\/?s>|<start_of_turn>|<end_of_turn>/g;

function base32(bytes: Uint8Array): string {
let out = "";
let bits = 0;
let value = 0;
for (const byte of bytes) {
value = (value << 8) | byte;
bits += 8;
while (bits >= 5) {
out += B32[(value >>> (bits - 5)) & 31];
bits -= 5;
}
}
if (bits > 0) out += B32[(value << (5 - bits)) & 31];
return out;
}

function toBase64Url(bytes: Uint8Array): string {
let binary = "";
for (const b of bytes) binary += String.fromCharCode(b);
return btoa(binary).replace(/\+/g, "-").replace(/\//g, "_").replace(/=+$/, "");
}

function fromBase64Url(text: string): Uint8Array<ArrayBuffer> {
const binary = atob(text.replace(/-/g, "+").replace(/_/g, "/") + "=".repeat((4 - (text.length % 4)) % 4));
return Uint8Array.from(binary, (c) => c.charCodeAt(0));
}

/** A fresh session key: 32 random bytes, unpadded base64url. One per conversation. */
export function newSessionKey(): string {
return toBase64Url(crypto.getRandomValues(new Uint8Array(32)));
}

/** `tag_n` for [sessionKey] (unpadded base64url, at least 16 bytes). */
export async function turnTag(sessionKey: string, n: number): Promise<string> {
const raw = fromBase64Url(sessionKey);
if (raw.length < 16) throw new Error("sessionKey must decode to at least 16 bytes");
if (!Number.isInteger(n) || n < 0) throw new Error("turn must be a non-negative integer");
const key = await crypto.subtle.importKey("raw", raw, { name: "HMAC", hash: "SHA-256" }, false, ["sign"]);
const mac = new Uint8Array(await crypto.subtle.sign("HMAC", key, new TextEncoder().encode(`azphalt-llm-turn:${n}`)));
return base32(mac).slice(0, 26);
}

/** Every tag of the session up to and including [turn], oldest first. */
export async function sessionTags(sessionKey: string, turn: number): Promise<string[]> {
return Promise.all(Array.from({ length: turn + 1 }, (_, n) => turnTag(sessionKey, n)));
}

/** Remove every session tag and native control marker from untrusted text. */
export function scrub(text: string, tags: string[]): string {
let out = text;
for (const tag of tags) out = out.split(`⟦${tag}⟧`).join("").split(`⟦/${tag}⟧`).join("");
return out.replace(NATIVE_MARKERS, "");
}

/** Wrap untrusted [material] for the current turn (the last of [tags]), scrubbing it first. */
export function wrap(material: string, tags: string[]): string {
const tag = tags[tags.length - 1];
if (!tag) throw new Error("wrap needs the session's tags");
return `⟦${tag}⟧${scrub(material, tags)}⟦/${tag}⟧`;
}

/** True when [text] carries any tag of the session: model output that does is rejected. */
export function containsSessionTag(text: string, tags: string[]): boolean {
return tags.some((tag) => text.includes(tag));
}

/**
* The translator: tagged segments of the current turn become their own `user` messages, scrubbed;
* text outside tags keeps its role. Throws on an unterminated segment or a tag that survives.
*/
export function translate(messages: ChatMessage[], tags: string[]): ChatMessage[] {
const out: ChatMessage[] = [];
const add = (role: ChatMessage["role"], content: string) => {
if (content.trim()) out.push({ role, content });
};
for (const m of messages) {
if (!tags.length) {
add(m.role, m.content);
continue;
}
const tag = tags[tags.length - 1];
const opening = `⟦${tag}⟧`;
const closing = `⟦/${tag}⟧`;
let pos = 0;
for (let start = m.content.indexOf(opening, pos); start >= 0; start = m.content.indexOf(opening, pos)) {
const end = m.content.indexOf(closing, start + opening.length);
if (end < 0) throw new Error("unterminated tagged segment");
add(m.role, m.content.slice(pos, start));
add("user", scrub(m.content.slice(start + opening.length, end), tags));
pos = end + closing.length;
}
add(m.role, m.content.slice(pos));
}
if (out.some((m) => containsSessionTag(m.content, tags))) throw new Error("a session tag survived translation");
return out;
}
Loading
Loading