Skip to content

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

jev-context

Pi carries every installed skill's description in the system prompt on every call and hopes the model notices the right one. This extension removes that index and loads the skills a turn needs, chosen by a judge that costs about two tenths of a cent per turn.

jev-context: one session, two context policies

This is an extension for Pi, the terminal coding agent. It replaces Pi's attention-based skill discovery with a scored one, using TypeSafe's Jev, a System One model that returns calibrated probabilities instead of prose. The reasoning model never sees a skill list to notice. A cheap judge decides which skills load, and code enforces the decision. Since 0.3.0 skill routing is the extension's only job (see Removed in 0.3.0).

The problem

A Pi session pays its full context on every model call. Eighteen installed skills means 18 descriptions in the system prompt, read on every turn, noticed unreliably (Pi's own docs admit models often fail to load the skill they need). When the model does notice a skill, it reads the SKILL.md with a tool call, and that call and its output stay in history for the rest of the session.

How it works

The extension hooks three Pi lifecycle events (session_start, input, before_agent_start) and makes one kind of decision: which skills to load for this message.

Skill routing (input, before_agent_start). Jev replaces Pi's native skill discovery and nothing else. When a Jev key is configured, the extension removes Pi's native <skills> index from the system prompt on every turn, identically, so the prompt prefix is the same for the whole session. On each user message it builds a routing state (your latest message verbatim in its own field, plus a digest of the conversation: user turns plus assistant text and thinking, tool I/O excluded, newest-first, capped at 80KB) and scores every not-yet-loaded skill against it, one request per skill in parallel, each carrying the skill's scoring text rather than its description: a short routing card when one is current (see Skill cards below), else the skill's full body. Every skill whose score reaches loadThreshold (0.6) loads in that one message, highest score first: the input is rewritten to /skill:<best> <your text> with each further skill pre-rendered as a byte-identical <skill name="…" location="…"> block ahead of your text, so Pi's own /skill: expansion writes the first block and the rest ride along verbatim in the same user message (the TUI may show only the first as a skill card). A loaded skill is never scored or loaded again, and the stack stays in history like any manual /skill: load. Because the blocks are appended at the end of the context, the cached prefix is untouched. One deterministic shortcut sits ahead of scoring: name a skill next to the word "skill" ("use the spec skill", "herdr skill", "skill: tldraw") and it loads without a Jev request, ahead of any Jev winners: an explicitly named skill is a user invocation, which is also the one way manual-only skills route; set explicitMentions: false to turn the rule off. Skills marked manual-only (disable-model-invocation) are otherwise never auto-loaded. Without a key, Pi's native index stays in the system prompt and nothing is routed.

One rule shapes everything: mechanics in code, judgment in the model. Caching, budget caps, and thresholds are deterministic TypeScript; relevance is Jev's. The only exception is the explicit-mention shortcut above, which the owner chose because naming a skill is already a user decision.

How we use Pi's seams

The input event fires once per user message, before Pi's own skill expansion, which is the only moment routing makes sense (tool-loop iterations don't change what the user wants). Skill loading reuses Pi's native /skill: expansion rather than a parallel mechanism. before_agent_start exposes the system prompt options, so removing the native skill index is one assignment (systemPromptOptions.skills = []) that Pi applies the same way on every turn. Stats are a plain slash command (/skill_stats); /skill:name stays Pi's own command. The extension registers no context handler, so it never modifies the message list the model sees, and it never changes the active tool set.

The extension is one file of erasable TypeScript with zero runtime dependencies. It runs under Node's strip-types mode, and tests run on node:test against a fixture server; no test touches the network. The quality contract (VERIFYING.md: biome with zero warnings, tsgo with erasableSyntaxOnly, pinned-dependency ratchet, scenario-mirrored behavior tests) is enforced by npm run check && npm test.

Skill cards (/index_skills)

Scoring a raw SKILL.md means judging install steps and API detail when the question is "is this relevant now?". The extension can score against a short card instead: what the skill is for, when to use it, and when not to (naming the sibling skill to pick instead). /index_skills writes those cards. It is owner-invoked only — nothing spawns in the background: for each skill it runs one headless Pi child (pi -p --no-session --no-extensions --no-skills --no-context-files --tools read,ls,grep,find --model <model>, working directory the skill's own directory, so it can read the skill's files but is told to modify nothing), gives it the full SKILL.md plus the name: description line of every other skill, and parses back one JSON object {summary, use_cases, when_to_use, when_not_to_use}.

/index_skills                 # every skill whose card is missing or stale
/index_skills tldraw spec     # just these, even if current
/index_skills --force         # all skills, current or not

Cards live in <agentDir>/jev-context-skill-index.json (never in the repo), one entry per skill recording the source SKILL.md path, the sha256 of the content it was written from, the model, and the card. Writes are atomic (temp file + rename) and merge into the existing file: a child that fails, times out, or returns unparseable output leaves that skill's previous card untouched and is logged as SKILL_INDEX_FAILED: skill=… reason=…; each success logs SKILL_INDEX_OK. The command ends with one summary notify (indexed=N failed=M unchanged=K).

At session start the extension loads the file. A skill whose card's sha256 matches its current SKILL.md scores against the rendered card; any other skill scores against the raw body, and one notice per session lists them (run /index_skills), logged as SKILL_INDEX_STALE. ROUTE_DECISION telemetry records which text each candidate carried (scoring_text), so the eval harness can attribute score differences to the text. Cards only ever feed scoring: the block the model receives is always the real SKILL.md through Pi's own expansion.

Relevant config (all optional):

{
  "scoringText": "card",          // "card" (default) or "body" (forces raw bodies, the pre-card format)
  "indexModel": "zai/glm-5.3",    // provider/id for the children; default: the current session model
  "indexConcurrency": 4,           // max children in flight
  "indexTimeoutMs": 180000         // per-child timeout
}

Cache discipline

Prefix caching is where naive context manipulation fails. Insert a message mid-history and everything after the insertion point recomputes; do it every turn and every call pays full price. Our rules: the native skill index is removed identically on every turn, so the system prompt never changes mid-session, and a routed skill enters as part of the new user message, at the end of the context, through Pi's own /skill: expansion. Routing decisions land at the user-turn boundary only, where the prefix was going to change anyway because the user just typed. A routed session keeps the same cache-hit shape as a native one.

Numbers

Everything below comes from the harness in eval/, which replays real Pi session logs through the extension's real router and counts what each arm would have billed. Corpora are owner-local: point it at your own sessions (they live under ~/.pi/agent/sessions/) and see what you get. Ours:

Golden set (12 conversation-heavy sessions, 133 user turns, 1,662 model calls): skill routing alone saved 33.9%. Earlier versions of this README quoted larger figures that also counted tool-schema and pruning savings; those features are gone, and so are the numbers.

Routing quality, from a 12-session labeled eval (6 positive, 6 hard negatives, 19 skills, full-body scoring): labeled-relevant skills scored 0.81-0.92; the 210 irrelevant judgments had median 0.08, p90 0.43, max 0.92. The distributions overlap, which is why a naked cutoff was not enough: the false-positive rate at 0.6 was 7.6%, concentrated in broad-description skills that run hot, which per-skill thresholds then absorb. A naked 0.5 would have shipped false confidence.

Cost and latency. Jev is billed on input only, $0.042 per million tokens, output free. A turn in a long session routes 18 skills in parallel for ~53k input tokens: $0.0022 and 841ms, both off the critical path (the user is reading; the main model's time-to-first-token is longer). A fresh-corpus eval run over the full golden set cost $1.69. Replays against the recorded cache are free and byte-identical.

Live shakedown (this box, real sessions): a whiteboard prompt routed the tldraw skill at 0.98 against a next-best 0.10. A neutral prompt loaded nothing, max score 0.02.

Failure behavior

If Jev is unreachable, over quota, or unconfigured, routing fails static: with no key the session runs in native skill mode (Pi's skill index stays in the system prompt), and a scoring failure passes your message through unchanged, with one loud ui.notify per error class plus a ROUTE_DEGRADED log line. Explicitly named skills still load. The extension never writes to the transcript, so it cannot hold a session hostage.

Telemetry

Every routing decision lands in ~/.pi/agent/jev-context-telemetry.jsonl as structured JSON (ROUTE_DECISION per pass and SKILL_MODE once per session), with scores, scoring text kind, latencies, and token counts. /skill_stats renders aggregates in-session. Telemetry files written by earlier versions may also carry TOOL_SURFACE, PRUNE_JUDGED and PRUNE_EPOCH lines; /skill_stats skips them. Thresholds are config, and the intended workflow is to tune them from your own telemetry after a few days, not to trust ours.

Removed in 0.3.0

Tool surfacing and epoch pruning were removed by owner decision on 2026-10-05. Tool surfacing hid whole tool namespaces from the model (a live tavily_* miss), and Pi 1.0's deferred tool exposure and tool_search cover tool loading natively. Pruning is out of scope for this extension. The context, agent_settled and message_end handlers are gone, and the config keys toolNamespaces, coreTools, toolSurfaceThreshold, pruneThreshold and pruneStateCapBytes are ignored with one warning. context_edit entries that earlier versions already wrote into old sessions stay in those sessions: they are Pi-native entries, and Pi keeps applying them when those sessions are resumed.

Install

pi install git:github.com/Growth-Kinetics/jev-context          # latest
pi install git:github.com/Growth-Kinetics/jev-context@v0.3.0   # pinned

Or manually: symlink extensions/jev-context.ts into ~/.pi/agent/extensions/.

Configure in ~/.pi/agent/jev-context.json (all fields optional):

{
  "apiKeyEnv": "PI_TYPESAFE_JEV",       // or "apiKeyFile": "~/.pi/agent/secrets/jev.key"
  "loadThreshold": 0.6,
  "consoleLog": false,                   // true echoes judgment lines to stderr
  "explicitMentions": true,              // "the spec skill" loads spec deterministically, no Jev call
  "scoringText": "card",                 // score skills against routing cards; "body" forces raw SKILL.md
  "indexModel": "zai/glm-5.3"           // model for /index_skills children; default: session model
}

Naming a skill in your latest message is first-class evidence: the routing state carries that message verbatim in its own field, so one explicit word is not diluted by a long session's digest.

The skill index is Pi's own: whatever skills Pi discovered for the session (its canonical skill roots, packages, --skill flags). The user config lives in Pi's agent dir (PI_CODING_AGENT_DIR, default ~/.pi/agent); a project-level .pi/jev-context.json overrides it.

Data handling

Routing sends the routing state (your latest message plus a digest of user turns and assistant text and thinking; tool calls and outputs are excluded) and each candidate skill's scoring text to the configured TypeSafe endpoint over HTTPS. That is the whole egress surface. The key is never logged, the transcript is never written, and nothing is sent anywhere else. /index_skills spawns local pi child processes, only on the explicit command; their model traffic uses your own Pi model and auth configuration, exactly as if you had run pi -p yourself. If your sessions are sensitive, read extensions/jev-context.ts first; it is one file, and the network path is one function.

Development

npm install --ignore-scripts
npm run check   # biome (zero warnings) + tsgo (erasableSyntaxOnly) + pinned deps + eval gate
npm test        # node --test, no network
JEV_SMOKE_MODEL=zai/glm-5.3 npm run smoke:native   # live, owner-run

npm run smoke:native is a live check, not part of npm run check: it runs the installed pi three times in print mode against one session in a scratch agent dir (this extension, one fixture skill, a loopback fixture Jev server, your copied auth.json/models.json) and asserts from the session file that the system prompt has no skills section, the fixture skill loaded once through Pi's native expansion on the first turn, and every turn got an answer. It prints SMOKE_OK, SMOKE_FAIL: <reason>, or SMOKE_SKIP: reason=… (no pi, no JEV_SMOKE_MODEL, or no source auth.json; JEV_SMOKE_REQUIRE=1 turns a skip into a failure). JEV_SMOKE_AGENT_DIR picks the source agent dir.

VERIFYING.md is the binding contract. eval/README.md documents the harness. Built autonomously by a GOAL loop of coding agents (kimi-coding/k3 and zai/glm-5.3) executing specs against that contract, reviewed by fresh-eyes subagents per milestone; the process artifacts are not in the repo, but the eval harness they built is how you should decide whether to trust any of the above numbers.

License

MIT

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages