Every model has its own personality when it comes to tool calling. This guide covers what we've learned about each one and how the proxy is tuned to compensate.
| Model | Tier | VRAM/RAM | Speed | Tool Accuracy | Profile |
|---|---|---|---|---|---|
| Qwen3-Coder 30B | Fast | 19GB GPU | 1-15s | Excellent | qwen3-coder.toml |
| Qwen3-Coder 480B | Quality | 290GB hybrid | 22-24s | Perfect | qwen3-coder-480b.toml |
| Qwen3 30B | Fast | 18GB GPU | 10-18s | Good | qwen3.toml |
| Qwen3 235B | Quality | 142GB hybrid | 2-5 min | Very good | qwen3-235b.toml |
| Devstral Small 2 24B | Fast | 15GB GPU | 1-7s | Excellent | devstral-small-2.toml |
| GLM-4.7-Flash | Fast | 19GB GPU | ~2s | Excellent | glm-4.7-flash.toml |
| GPT-OSS 20B | Fast | 12GB GPU | 1-4s | Good (simple), Weak (complex) | gpt-oss.toml |
| GPT-OSS 120B | Quality | 65GB hybrid | 6-17s | Excellent | gpt-oss-120b.toml |
| Llama 4 Maverick | Quality | large hybrid | varies | Moderate | llama4-maverick.toml |
| Llama 4 Scout | Fast | - | - | Not working | llama4-scout.toml |
| Llama 3.3 70B | Standalone | 57GB hybrid | 60-100s | Moderate (with retries) | llama3.3.toml |
The star of the show. Same MoE architecture as regular Qwen3 30B but purpose-built for code and agentic tool use. Runs fully on GPU (~1-15s per tool call).
Why it's great:
- Does not use
<think>tags — no stripping needed - Produces clean, native tool calls on the first attempt most of the time
- 256K native context window
- In testing: 25+ consecutive requests at 100% success rate, 50 messages deep
Quirks:
- Occasionally narrates mid-conversation (just like regular Qwen3), but self-corrects after a single nudge
- Without
accept_text_after_tool_use = true, it will loop forever making tool calls instead of giving a final answer — a coding model that literally can't stop coding - No
<think>tags, no embedded JSON blobs, no hallucinated parameters
Profile tuning:
strip_thinking = false # No think tags to strip
accept_text_after_tool_use = true # Let it give final answers
condense_tools = true # Reduce context pressureThe 480B quality-tier coder model. 290GB, runs hybrid CPU/GPU on port 11435. Tool call accuracy is perfect — every completion validated first attempt with zero retries.
Limitations: Speed. With a single A30 (24GB VRAM), ~1-2 minutes per turn and longer conversations can exceed timeouts. Adding a second GPU (A2 16GB) brought this down to 22-24s per turn — a dramatic improvement from moving more layers off CPU.
Quirks:
- Same clean tool-calling behavior as the 30B
- Best suited for short, high-stakes exchanges where quality > speed
- Escalation disabled — nothing bigger to escalate to
The workhorse. Fast, reliable, takes direction well. Runs fully on GPU (~10-18s per tool call).
Why it works:
- Responds well to error feedback — usually self-corrects within 2-3 retries
- Benefits heavily from system prompt nudging
- Escalates to the 235B when all else fails
Quirks:
- Wraps reasoning in
<think>tags — needsstrip_thinking = true - Occasionally hallucinates plausible-sounding parameters that don't exist in the schema
- Responds well to structured "CRITICAL RULES" system prompts
- Without a system nudge, will happily describe what it would do instead of doing it
Profile tuning:
strip_thinking = true # Clean up internal monologue
condense_tools = true # 85% reduction in tool description size
max_system_tokens = 800 # Hard cap on system promptThe big gun. Runs hybrid CPU/GPU (~2.5 min per tool call with dual GPU, 2-5 min with A30 only). Don't use it as your daily driver — it's the escalation target when the 30B can't figure it out. At 142GB, only 26% fits on GPU even with both cards — the A2 helps but doesn't move the needle as dramatically as it does for smaller quality models.
Why it works:
- Rarely needs retries (
max_retries = 1) - Surprisingly good at recovering from the 30B's mistakes when given the same context
Quirks:
- Same
<think>tag habit as its smaller sibling - Understands what went wrong and corrects course when given the 30B's error context
OpenAI's open-weight model (August 2025). Purpose-built for tool calling with native OpenAI format. Ships pre-quantized in MXFP4. Blazing fast — 1-4 seconds per tool call on warm requests.
Why it's great:
- ~12GB, fits entirely on A30 GPU
- Matches o3-mini on TauBench
- Extremely fast: 1-4s per tool call after warm-up (first request ~30s for model load)
- Simple tool calls (read, glob, grep) work cleanly on first attempt
- Natural 20B -> 120B escalation chain
Quirks:
- Complex nested schemas are its weakness. Tools with deeply nested required fields (like OpenCode's
questiontool with itsquestions[].headerandquestions[].optionsstructure) consistently fail — the model flattens nested fields to the top level or sends wrong types. Escalation to the 120B handles this. - Narration tendency in longer conversations. Starting around message 12+, begins responding with text instead of tool calls on roughly half of requests. The proxy's narration feedback nudge corrects it on the next attempt (typically adds ~5s). Setting
accept_text_after_tool_use = truelets it give final answers without looping. - Weak on multilingual/Chinese tasks.
Profile tuning:
strip_thinking = false # No think tags
condense_tools = true # Reduce context pressure
accept_text_after_tool_use = true # Narrates mid-conversation, let it answer
max_system_tokens = 800 # Keep context tight
# Escalation to 120B covers complex schema failuresThe quality-tier GPT-OSS model. ~65GB in MXFP4, runs hybrid CPU/GPU on port 11435. Zero validation failures in testing — every tool call valid on first attempt, including complex nested schemas that the 20B can't handle.
Why it's great:
- Matches o4-mini on TauBench
- 100% first-attempt tool call accuracy across all tested tools
- Handles complex nested schemas (like OpenCode's
questiontool) that break the 20B - With A30 only: 39-56s per call. With A30 + A2: 6-17s per call (3-4x speedup from pushing past 50% GPU)
- Escalation target for GPT-OSS 20B
Quirks:
- Same narration tendency as the 20B at deeper conversations (~message 10+). Responds with text instead of tool calls. With
accept_text_after_tool_use = true, this is accepted as the final answer without a costly retry. - Speeds up as context builds — went from 56s to 39s over the session as the model warmed up.
Profile tuning:
max_retries = 2 # Quality model, fewer retries needed
request_timeout = 900.0 # Generous timeout for hybrid mode
accept_text_after_tool_use = true # Narrates at deeper conversations
escalation.enabled = false # Nothing bigger to escalate toPurpose-built for agentic coding. Mistral's coding agent model, fine-tuned from Mistral Small 3.1. Dense 24B — not MoE, but still fits comfortably on A30 at ~15GB. Apache 2.0 license. 68% SWE-Bench Verified.
Why it's great:
- 100% first-attempt tool call accuracy in testing — zero retries, zero validation failures, zero param repairs
- Maintained accuracy deep into conversations (40+ messages, 10+ chained tool calls)
- Fast and consistent: 1-7s per tool call, speeds up as KV cache warms
- First model family (Mistral) outside the Qwen/Llama/GLM/GPT-OSS lineup
Quirks:
- Wanders on open-ended tasks without context cap. The 384K native context is too much rope — on broad search tasks (e.g., "find all files that import X"), it can lose focus and start searching for unrelated things. Setting
num_ctx = 65536fixes this. - No
<think>tags, no embedded JSON blobs - No escalation partner available (no larger Mistral model on the server)
Profile tuning:
strip_thinking = false # No think tags
accept_text_after_tool_use = true # Let it give final answers
condense_tools = true # Keep context tight
num_ctx = 65536 # Prevents wandering on open-ended tasks
escalation.enabled = false # No Mistral quality-tier availableFast and clean. Zhipu/Z.ai's open-weight model (MIT license). Different model family from everything else in the lineup — non-Western architecture, strong tau-Bench tool-use scores. Runs fully on GPU (~19GB Q4).
Why it works:
- Clean tool calls on first attempt — passed every test without retries
- Correct tool selection even with multiple tools available
- Proper parallel tool calls when appropriate
- ~2s response time once warm
- 198K native context (we set
num_ctx = 65536for reliable tool use)
Quirks:
- Returns thinking in a separate
reasoningfield (not<think>tags in content) —ChatMessageuses Pydantic's defaultextra = "ignore", so thereasoningfield is dropped during response parsing - Mixes explanatory text alongside tool calls — needs
accept_text_after_tool_use = true strip_thinking = trueis set as a safety net but the real filtering happens via field dropping- Ollama 0.15.1+ required for tool calling quality fixes
Profile tuning:
strip_thinking = true # Safety net (reasoning field auto-dropped)
accept_text_after_tool_use = true # Mixes text with tool calls
condense_tools = true # 3B active params, keep context tight
num_ctx = 65536 # Needs large context for reliable tool useThe storyteller. Maverick really wants to explain itself. It will write a paragraph about what it's going to do, then tack a JSON tool call blob onto the end.
How the proxy handles it:
The _rescue_embedded_tool_calls feature was built specifically for Maverick. It:
- Scans the text response for embedded JSON objects
- Parses them as tool calls
- Strips the JSON from the text, preserving the narration
- Returns both the tool call and the cleaned text
Quirks:
- Mixes text and tool calls in the same response
- Occasionally sends empty arguments
{}on the first try, then fixes after feedback - Will search for
requirements.txtthree times before tryingpyproject.toml - Needs
accept_text_after_tool_use = trueor it loops forever
Not working. Does not produce usable tool calls even with aggressive proxy workarounds (tool_choice_override = "required", exclude_tools, simplified system prompt, reduced num_ctx).
The profile exists for experimentation. Use Maverick instead.
Working. A dense 70B model — no MoE here, so it's heavier on resources (~57GB, runs 59% CPU / 41% GPU on A30). Slower than the MoE models but produces usable tool calls with proxy assistance.
Why it works:
- Responds to error feedback and self-corrects parameter names on retry
- Follows system prompt instructions about tool discovery (glob first, then read)
- The proxy's param name repair fixes minor near-miss names; broader mismatches are corrected via validation feedback and retries
Quirks:
- Guesses parameter names. Often produces near-miss names (e.g.,
questioninstead ofquestions,filePathinstead offile_path). Near-misses are auto-repaired; larger deviations are corrected via validation feedback and retries. - Slow cold starts. First request after loading takes 60-70s; needs
backend_timeout = 300to avoid timeouts. - Assumes Node.js projects. Without explicit guidance in the system suffix, it will try to read
package.jsonbefore anything else — even for Python projects. The profile's system suffix now tells it to glob the root first. - No escalation configured. Runs standalone; no quality-tier Llama model is practical for on-demand escalation given memory constraints.
Profile tuning:
backend_timeout = 300.0 # Cold starts need headroom
strip_thinking = false # No think tags
condense_tools = true # Reduce context pressure
accept_text_after_tool_use = true # Let it give final answers
escalation.enabled = false # No escalation targetEach fast model has a corresponding quality model it escalates to:
Fast (GPU) Quality (Hybrid CPU/GPU)
────────── ───────────────────────
qwen3-coder:30b -> qwen3-coder:480b
qwen3:30b-a3b -> qwen3:235b-a22b
gpt-oss:20b -> gpt-oss:120b
Escalation adds an error summary from the failed attempt so the quality model understands what went wrong.
-
Pull the model on the appropriate Ollama instance:
# Fast model (GPU) ollama pull your-model:small # Quality model (hybrid) OLLAMA_HOST=localhost:11435 ollama pull your-model:large
-
Create a profile in
profiles/:- Start by copying the profile of the most similar existing model
- Set the
patternto match your model name - Name the file so it sorts correctly (more specific before general)
-
Tune the profile:
- Does it use
<think>tags? Setstrip_thinking = true - Does it narrate instead of calling tools? Add a strong
system_suffix - Does it mix text + tool calls? Set
accept_text_after_tool_use = true - Does it struggle with large contexts? Enable
condense_toolsandcondense_system_prompt - Does it have a quality counterpart? Configure
[escalation]
- Does it use
-
Test it:
curl http://localhost:8079/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"your-model:small","messages":[{"role":"user","content":"Read the file README.md"}],"tools":[{"type":"function","function":{"name":"read","description":"Read a file","parameters":{"type":"object","properties":{"path":{"type":"string"}},"required":["path"]}}}]}'
-
Watch the logs for validation errors, rescues, and retries. Adjust the profile accordingly.