Repository navigation
PipelineScore v4 launch - #9
Merged
Merged
Conversation
Native tool calling (OpenAI-compatible and Anthropic wire formats), a multi-turn agent loop, a Docker sandbox for model-written code (no network, capped memory/CPU/pids, read-only root, non-root), seeded task templates so answers can't be memorized, and quality-only scoring with speed reported separately. Eight starter tasks, one or two per suite: code, repo repair, agent, function calling (parallel + abstention), reasoning, long document, instruction following. test/v4.test.ts proves every grader passes a scripted correct answer and fails a plausible wrong one across seeds. Run: ps-bench v4 --provider local|openai|minimax|anthropic --model <id>
code 20 (14 Python, 6 JavaScript), repo repair 4, agent 6, function calling 14, reasoning 12, long document 5, instruction following 10. Every task is templated from the run seed and every grader is proven by a test that passes a correct answer and fails plausible wrong ones (about 2,300 checks across test/v4-*.test.ts). Also: MiniMax runs with its 196,608 output-token maximum so reasoning models are never cut off, a --max-tokens floor for any provider, and a 60-minute request timeout.
…wers A backend in an error state (seen: Ollama after a Metal compute error) keeps answering HTTP 200 with an empty choice, zero prompt tokens and no finish reason. That scored 69 tasks as model failures in a pilot run. The provider now throws on that shape, and a run stops after 3 provider errors in a row, marking the result aborted and exiting 2.
Pilot on gpt-oss-20b: its chat format emits one tool call per message, so single-turn parallel tasks scored it 0.1-0.25 for making every right call one at a time. Single-turn tool tasks now answer each call with an ok result for up to 3 follow-up turns and grade the union of calls; duplicates and extra calls still cost points. Results also flag a raw, unparsed tool call in the text as a server problem.
Answering each call with a bare ok led gpt-oss-20b to add unrequested calls (emailing a customer, re-searching) after it had already made the right ones, which cost tasks it had passed. Grading now keeps the best score across the cumulative calls after each turn, so credit lands when the requested calls are complete; a duplicate before finishing still fails. Up to 6 follow-up turns so four parallel calls can arrive one at a time.
`ps-bench profile <results...> --out profile.json` merges saved v4 runs into pipelinescore-routing-profile/1: per model and hardware tag, local vs cloud, per-suite scores with confidence bands, per-task scores, speed, and example prompts per suite so a router can place live requests next to the kind of task each score came from. Refuses aborted runs and mixed testpack versions. v4 results now record a hardware tag (detected for local models, 'cloud' otherwise, or --hardware-tag).
…n calling Pilot showed four suites saturating: MiniMax-M2.7 scored 100 on reasoning, long documents and instructions, and a 4B model matched strong models on function calling (93 vs 93-96). - reason: multi-event tanks, knights and knaves, closed-road routing, crewed critical paths, inclusion-exclusion traps; 500/500 distinct prompts per task, independent solvers agree. - instruct: 8-11 interacting, precisely defined rules per task (acrostics, exact letter and word counts, arithmetic-consistent CSV), two all-or-nothing tasks; per-task token caps removed. - longdoc: 90-104K-char documents, 4-hop chains with transfers and corrections, near-name decoys, computed answers. - fncall: 27-tool menu with near misses, time-zone/date/unit math, dependent calls fed realistic tool results (new optional toolResult hook), harder abstention and clarification. Calibration: fncall now 4B 30 / gpt-oss-20b 45 / MiniMax 89; reasoning gpt-oss 83 -> 50; instructions gpt-oss 92 -> 62; MiniMax still 100 on reasoning, instructions and long documents (it verifies against every checkable rule, thinking up to 8 minutes a task).
A 504 from the provider during a 40-minute MiniMax task scored 0 in the calibration run. Retries go 2 -> 4; results list provider_errors; the CLI warns to rerun those tasks; routing profiles refuse such runs.
Profiles carried only benchmark prompts as suite examples, so a router learned what the test looks like, not what users type. 12 hand-written everyday requests per suite now come first. Measured on the Weave fork (held-out sets): everyday prompts routed to the right job type 60% -> 90%, benchmark-style 91% -> 85%.
Reads the x-router-model response header; results list served_by per task, in call order. Needed to measure routed runs.
Routers that pin sessions (Weave keeps a conversation on one model for prompt caching) treated all 71 independent tasks as one session, so a task pinned to the local model dragged later tasks with it. Each task now carries its own Session-Id header, constant across its turns.
Routers cap time-to-first-byte (Weave: 30 s). A long-thinking model sent a non-streaming request timed out there before its first byte. --stream reassembles text, split tool-call deltas and usage into the same shape, and keeps the router's x-router-model header.
Needed to opt out of router UI decorations, e.g. Weave's streamed routing banner (X-Weave-Routing-Marker: off), which otherwise lands in the answer text and fails exact-format tasks.
- Every query over submissions takes testpack=v3|v4 (v4 = testpack 4.*); the default comes from DEFAULT_TESTPACK and stays v3, so deploying this changes nothing until launch flips the env var. - ps-bench v4-submit <results...>: posts saved v4 runs (suite means as categories, bands/seed/speed in score_detail); refuses aborted runs and runs with provider errors. - Verified on a local API: default board = v3 only, ?testpack=v4 = v4 only, stats/users follow the board, DEFAULT_TESTPACK=v4 flips the default.
- /leaderboard/v4: best run per model + hardware, seven suites, run time beside the score - Every existing web request now asks for testpack=v3 explicitly, so flipping the API default to v4 cannot change the current pages; /s/<id> renders v4 suites for v4 runs - API leaderboard entries carry hardware_tag - Checked locally against a scratch API (one v3 + one v4 run): v4 board shows only v4, v3 board only v3, no overflow at 1440/390, no runtime errors
…ng NameError Glimmer and Gemma 4 both spent their whole 16K output budget thinking on code-parse-duration-1 and returned no code; the grade read 'NameError' on all 31 cases. Score is unchanged (0).
Rules like 'no bullet ends with punctuation' hold vacuously on empty text, so an empty answer earned 9-25% on four tasks in the Qwen 3.8 launch run. The test now requires 0.
It was raised to the API ceiling (196,608), so MiniMax passed four launch tasks that needed 16.5K-39.5K tokens while local models were cut off at 16K. One board, one limit (Drew, 10-01).
…as an archive, /leaderboard/v4 redirects
mkdtemp makes the folder 0700 and the container runs as nobody. On Linux the bind mount keeps host permissions, so every code task saw 'can't open file' and scored 0. Docker Desktop on macOS hides it, which is why the launch runs (all on Macs) were fine and CI (Linux) failed.
Deploying with
|
| Status | Name | Latest Commit | Updated (UTC) |
|---|---|---|---|
| 🔵 In progress View logs |
pipelinescore | 84a173b | Oct 01 2026, 04:10 PM |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Launches the v4 test (71 tasks across code, repo repair, agent tasks, function calling, reasoning, long documents and instructions).
pipelinescore v4runner, Docker sandbox,v4-submit(refuses aborted runs and runs with provider errors). One answer budget for every provider (16K). Empty instruction answers score 0.4.*);?testpack=v3|v4picks one, default stays v3 for existing API clients. Leaderboard rows carryhardware_tag./leaderboardis now the v4 board; the v3 board moves to/leaderboard/v3as an archive;/leaderboard/v4redirects. Every existing v3 page asks fortestpack=v3explicitly. Run pages show v4 suites.Launch runs (seed launch1, same 16K budget): Qwen 3.8 27B 94.8, Gemma 4 31B 90.3, Muse Glimmer 30B 89.4, MiniMax M2.7 83.2, gpt-oss-20b 70.6. Submitted after the backend deploys.