Skip to content

PipelineScore v4 launch - #9

Merged
drewmattie-code merged 22 commits into
mainfrom
feat/v4
Oct 1, 2026
Merged

drewmattie-code merged 22 commits into
mainfrom
feat/v4

Conversation

@drewmattie-code

Copy link
Copy Markdown
Owner

Launches the v4 test (71 tasks across code, repo repair, agent tasks, function calling, reasoning, long documents and instructions).

  • CLI: pipelinescore v4 runner, Docker sandbox, v4-submit (refuses aborted runs and runs with provider errors). One answer budget for every provider (16K). Empty instruction answers score 0.
  • API: every board query is split by testpack (v3 vs 4.*); ?testpack=v3|v4 picks one, default stays v3 for existing API clients. Leaderboard rows carry hardware_tag.
  • Web: /leaderboard is now the v4 board; the v3 board moves to /leaderboard/v3 as an archive; /leaderboard/v4 redirects. Every existing v3 page asks for testpack=v3 explicitly. Run pages show v4 suites.

Launch runs (seed launch1, same 16K budget): Qwen 3.8 27B 94.8, Gemma 4 31B 90.3, Muse Glimmer 30B 89.4, MiniMax M2.7 83.2, gpt-oss-20b 70.6. Submitted after the backend deploys.

Native tool calling (OpenAI-compatible and Anthropic wire formats), a
multi-turn agent loop, a Docker sandbox for model-written code (no
network, capped memory/CPU/pids, read-only root, non-root), seeded task
templates so answers can't be memorized, and quality-only scoring with
speed reported separately.

Eight starter tasks, one or two per suite: code, repo repair, agent,
function calling (parallel + abstention), reasoning, long document,
instruction following. test/v4.test.ts proves every grader passes a
scripted correct answer and fails a plausible wrong one across seeds.

Run: ps-bench v4 --provider local|openai|minimax|anthropic --model <id>
code 20 (14 Python, 6 JavaScript), repo repair 4, agent 6, function
calling 14, reasoning 12, long document 5, instruction following 10.
Every task is templated from the run seed and every grader is proven by
a test that passes a correct answer and fails plausible wrong ones
(about 2,300 checks across test/v4-*.test.ts).

Also: MiniMax runs with its 196,608 output-token maximum so reasoning
models are never cut off, a --max-tokens floor for any provider, and a
60-minute request timeout.
…wers

A backend in an error state (seen: Ollama after a Metal compute error)
keeps answering HTTP 200 with an empty choice, zero prompt tokens and no
finish reason. That scored 69 tasks as model failures in a pilot run.
The provider now throws on that shape, and a run stops after 3 provider
errors in a row, marking the result aborted and exiting 2.
Pilot on gpt-oss-20b: its chat format emits one tool call per message,
so single-turn parallel tasks scored it 0.1-0.25 for making every right
call one at a time. Single-turn tool tasks now answer each call with an
ok result for up to 3 follow-up turns and grade the union of calls;
duplicates and extra calls still cost points. Results also flag a raw,
unparsed tool call in the text as a server problem.
Answering each call with a bare ok led gpt-oss-20b to add unrequested
calls (emailing a customer, re-searching) after it had already made the
right ones, which cost tasks it had passed. Grading now keeps the best
score across the cumulative calls after each turn, so credit lands when
the requested calls are complete; a duplicate before finishing still
fails. Up to 6 follow-up turns so four parallel calls can arrive one at a
time.
`ps-bench profile <results...> --out profile.json` merges saved v4
runs into pipelinescore-routing-profile/1: per model and hardware tag,
local vs cloud, per-suite scores with confidence bands, per-task scores,
speed, and example prompts per suite so a router can place live requests
next to the kind of task each score came from. Refuses aborted runs and
mixed testpack versions. v4 results now record a hardware tag
(detected for local models, 'cloud' otherwise, or --hardware-tag).
…n calling

Pilot showed four suites saturating: MiniMax-M2.7 scored 100 on
reasoning, long documents and instructions, and a 4B model matched strong
models on function calling (93 vs 93-96).

- reason: multi-event tanks, knights and knaves, closed-road routing,
  crewed critical paths, inclusion-exclusion traps; 500/500 distinct
  prompts per task, independent solvers agree.
- instruct: 8-11 interacting, precisely defined rules per task
  (acrostics, exact letter and word counts, arithmetic-consistent CSV),
  two all-or-nothing tasks; per-task token caps removed.
- longdoc: 90-104K-char documents, 4-hop chains with transfers and
  corrections, near-name decoys, computed answers.
- fncall: 27-tool menu with near misses, time-zone/date/unit math,
  dependent calls fed realistic tool results (new optional toolResult
  hook), harder abstention and clarification.

Calibration: fncall now 4B 30 / gpt-oss-20b 45 / MiniMax 89;
reasoning gpt-oss 83 -> 50; instructions gpt-oss 92 -> 62; MiniMax still
100 on reasoning, instructions and long documents (it verifies against
every checkable rule, thinking up to 8 minutes a task).
A 504 from the provider during a 40-minute MiniMax task scored 0 in the
calibration run. Retries go 2 -> 4; results list provider_errors; the CLI
warns to rerun those tasks; routing profiles refuse such runs.
Profiles carried only benchmark prompts as suite examples, so a router
learned what the test looks like, not what users type. 12 hand-written
everyday requests per suite now come first. Measured on the Weave fork
(held-out sets): everyday prompts routed to the right job type 60% -> 90%,
benchmark-style 91% -> 85%.
Reads the x-router-model response header; results list served_by per
task, in call order. Needed to measure routed runs.
Routers that pin sessions (Weave keeps a conversation on one model for
prompt caching) treated all 71 independent tasks as one session, so a
task pinned to the local model dragged later tasks with it. Each task now
carries its own Session-Id header, constant across its turns.
Routers cap time-to-first-byte (Weave: 30 s). A long-thinking model sent a
non-streaming request timed out there before its first byte. --stream
reassembles text, split tool-call deltas and usage into the same shape,
and keeps the router's x-router-model header.
Needed to opt out of router UI decorations, e.g. Weave's streamed routing
banner (X-Weave-Routing-Marker: off), which otherwise lands in the answer
text and fails exact-format tasks.
- Every query over submissions takes testpack=v3|v4 (v4 = testpack 4.*); the default comes from
  DEFAULT_TESTPACK and stays v3, so deploying this changes nothing until launch flips the env var.
- ps-bench v4-submit <results...>: posts saved v4 runs (suite means as categories, bands/seed/speed in
  score_detail); refuses aborted runs and runs with provider errors.
- Verified on a local API: default board = v3 only, ?testpack=v4 = v4 only, stats/users follow the board,
  DEFAULT_TESTPACK=v4 flips the default.
- /leaderboard/v4: best run per model + hardware, seven suites, run time beside the score
- Every existing web request now asks for testpack=v3 explicitly, so flipping the API default to v4
  cannot change the current pages; /s/<id> renders v4 suites for v4 runs
- API leaderboard entries carry hardware_tag
- Checked locally against a scratch API (one v3 + one v4 run): v4 board shows only v4, v3 board only v3,
  no overflow at 1440/390, no runtime errors
…ng NameError

Glimmer and Gemma 4 both spent their whole 16K output budget thinking on code-parse-duration-1
and returned no code; the grade read 'NameError' on all 31 cases. Score is unchanged (0).
Rules like 'no bullet ends with punctuation' hold vacuously on empty text, so an empty
answer earned 9-25% on four tasks in the Qwen 3.8 launch run. The test now requires 0.
It was raised to the API ceiling (196,608), so MiniMax passed four launch tasks that needed
16.5K-39.5K tokens while local models were cut off at 16K. One board, one limit (Drew, 10-01).
mkdtemp makes the folder 0700 and the container runs as nobody. On Linux the bind mount keeps
host permissions, so every code task saw 'can't open file' and scored 0. Docker Desktop on macOS
hides it, which is why the launch runs (all on Macs) were fine and CI (Linux) failed.
@cloudflare-workers-and-pages

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Updated (UTC)
🔵 In progress
View logs
pipelinescore 84a173b Oct 01 2026, 04:10 PM

@drewmattie-code
drewmattie-code merged commit ff7c3e8 into main Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant