52 models, ranked on whether they resolved the customer’s problem without breaking policy — scored deterministically from replayable traces.
| # | Model | Overall▾ | Task | Tool | SOP | Safety | Grounding* | Tokens/success | p50 latency | Coverage |
|---|---|---|---|---|---|---|---|---|---|---|
| oraclereferencecorrect by construction | 100.00 | 1.000 | — | — | 1.000 | — | — | — | — | |
| anthropic/claude-fable-5 | 88.87 | 0.754 | 0.912 | 0.911 | 0.889 | 0.913 | 9,730 | 20.1 s | 100% | |
| moonshotai/kimi-k3 ★ | 86.96 | 0.715 | 0.902 | 0.906 | 0.908 | 0.870 | 5,759 | 18.9 s | 100% | |
| openai/gpt-5.6-luna-pro ★ | 86.15 | 0.860 | 0.867 | 0.841 | 0.947 | 0.826 | 23,211 | 14.0 s | 100% | |
| z-ai/glm-5.2 ★ | 81.11 | 0.435 | 0.811 | 0.910 | 0.927 | 0.923 | 7,481 | 9.1 s | 100% | |
| anthropic/claude-opus-4.6 | 81.03 | 0.425 | 0.808 | 0.913 | 0.889 | 0.937 | 13,549 | 11.3 s | 100% | |
| openai/gpt-5.6-sol ★ | 78.89 | 0.382 | 0.791 | 0.908 | 0.942 | 0.879 | 8,074 | 10.5 s | 100% | |
| google/gemini-3.1-pro-preview | 77.09 | 0.387 | 0.793 | 0.907 | 0.884 | 0.797 | 15,888 | 9.1 s | 100% | |
| deepseek/deepseek-v3.2 | 76.42 | 0.222 | 0.738 | 0.918 | 0.913 | 0.952 | 23,488 | 8.2 s | 100% | |
| google/gemini-3.5-flash | 75.57 | 0.227 | 0.740 | 0.908 | 0.894 | 0.923 | 24,933 | 4.2 s | 100% | |
| anthropic/claude-opus-4.7 | 75.36 | 0.295 | 0.762 | 0.913 | 0.908 | 0.816 | 22,632 | 6.5 s | 100% | |
| anthropic/claude-sonnet-4.5 | 75.16 | 0.188 | 0.727 | 0.913 | 0.899 | 0.947 | 27,767 | 6.3 s | 100% | |
| openai/gpt-5.5 | 75.04 | 0.208 | 0.734 | 0.908 | 0.899 | 0.918 | 21,536 | 5.0 s | 100% | |
| anthropic/claude-haiku-4.5 | 74.86 | 0.188 | 0.727 | 0.915 | 0.903 | 0.927 | 28,448 | 3.7 s | 100% | |
| openai/gpt-5.6-sol-pro ★ | 74.80 | 0.222 | 0.738 | 0.908 | 0.947 | 0.879 | 77,784 | 16.6 s | 100% | |
| minimax/minimax-m2.5 | 74.60 | 0.184 | 0.720 | 0.916 | 0.908 | 0.918 | 29,927 | 8.3 s | 100% | |
| minimax/minimax-m2.7 | 74.16 | 0.237 | 0.731 | 0.907 | 0.903 | 0.855 | 21,625 | 9.1 s | 100% | |
| z-ai/glm-5.1 | 74.06 | 0.193 | 0.724 | 0.907 | 0.889 | 0.899 | 25,293 | 8.4 s | 100% | |
| z-ai/glm-5 | 73.58 | 0.179 | 0.709 | 0.905 | 0.908 | 0.903 | 28,092 | 11.0 s | 100% | |
| openai/gpt-5.3-codex | 73.48 | 0.188 | 0.716 | 0.890 | 0.889 | 0.903 | 24,792 | 4.8 s | 100% | |
| qwen/qwen3.7-max | 73.46 | 0.191 | 0.728 | 0.910 | 0.887 | 0.863 | 31,217 | 13.7 s | 99% | |
| qwen/qwen3.7-plus | 73.21 | 0.193 | 0.729 | 0.905 | 0.899 | 0.845 | 30,124 | 11.3 s | 100% | |
| google/gemma-4-26b-a4b-it | 73.06 | 0.188 | 0.727 | 0.903 | 0.899 | 0.845 | 27,612 | 4.2 s | 100% | |
| z-ai/glm-4.5-air | 72.93 | 0.179 | 0.712 | 0.904 | 0.899 | 0.879 | 29,721 | 10.9 s | 100% | |
| google/gemini-3-flash-preview | 72.78 | 0.227 | 0.740 | 0.908 | 0.899 | 0.783 | 22,370 | 4.1 s | 100% | |
| google/gemini-2.5-flash-lite | 72.77 | 0.188 | 0.721 | 0.901 | 0.899 | 0.831 | 29,145 | 2.9 s | 100% | |
| anthropic/claude-sonnet-4.6 | 72.30 | 0.188 | 0.727 | 0.913 | 0.899 | 0.797 | 28,716 | 6.2 s | 100% | |
| deepseek/deepseek-v4-flash | 72.17 | 0.188 | 0.724 | 0.900 | 0.899 | 0.816 | 25,307 | 11.0 s | 100% | |
| moonshotai/kimi-k2.5 | 71.89 | 0.237 | 0.728 | 0.897 | 0.884 | 0.773 | 20,435 | 16.0 s | 100% | |
| openai/gpt-5.4 | 71.71 | 0.188 | 0.695 | 0.882 | 0.894 | 0.850 | 22,827 | 3.4 s | 100% | |
| google/gemini-3.1-flash-lite-preview | 71.59 | 0.188 | 0.727 | 0.902 | 0.899 | 0.749 | 26,860 | 1.4 s | 100% | |
| google/gemini-3.1-flash-lite | 71.49 | 0.188 | 0.727 | 0.905 | 0.899 | 0.758 | 26,827 | 2.8 s | 100% | |
| google/gemma-4-31b-it | 70.83 | 0.184 | 0.722 | 0.902 | 0.899 | 0.754 | 27,721 | 10.8 s | 100% | |
| openai/gpt-5.4-nano | 70.59 | 0.155 | 0.659 | 0.871 | 0.908 | 0.870 | 27,155 | 2.8 s | 100% | |
| openai/gpt-5-mini | 70.39 | 0.169 | 0.688 | 0.892 | 0.889 | 0.816 | 28,528 | 14.5 s | 100% | |
| qwen/qwen3-235b-a22b-2507 | 70.12 | 0.188 | 0.727 | 0.903 | 0.884 | 0.705 | 24,959 | 4.4 s | 100% | |
| minimax/minimax-m3 | 70.05 | 0.174 | 0.706 | 0.903 | 0.889 | 0.744 | 30,419 | 6.6 s | 100% | |
| deepseek/deepseek-v4-pro | 69.94 | 0.184 | 0.704 | 0.895 | 0.894 | 0.749 | 27,089 | 15.0 s | 100% | |
| stepfun/step-3.7-flash | 68.66 | 0.130 | 0.655 | 0.875 | 0.860 | 0.855 | 28,974 | 4.5 s | 100% | |
| openai/gpt-5.6-luna ★ | 68.41 | 0.618 | 0.627 | 0.721 | 0.942 | 0.599 | 4,438 | 4.8 s | 100% | |
| openai/gpt-4o-mini | 68.14 | 0.169 | 0.708 | 0.898 | 0.899 | 0.647 | 27,919 | 3.8 s | 100% | |
| inclusionai/ling-2.6-flash | 67.78 | 0.126 | 0.671 | 0.877 | 0.874 | 0.773 | 35,606 | 4.2 s | 100% | |
| google/gemini-2.5-flash | 67.20 | 0.188 | 0.724 | 0.901 | 0.899 | 0.556 | 26,434 | 3.4 s | 100% | |
| anthropic/claude-opus-4.8 | 66.18 | 0.082 | 0.480 | 0.906 | 0.884 | 0.829 | 71,567 | 5.7 s | 100% | |
| openai/gpt-5.4-mini | 65.74 | 0.184 | 0.655 | 0.851 | 0.899 | 0.628 | 23,021 | 2.2 s | 100% | |
| moonshotai/kimi-k2.6 | 65.63 | 0.227 | 0.659 | 0.839 | 0.831 | 0.686 | 14,662 | 11.7 s | 100% | |
| openai/gpt-5.6-terra-pro ★ | 63.89 | 0.241 | 0.703 | 0.874 | 0.971 | 0.396 | 69,519 | 9.3 s | 100% | |
| mistralai/mistral-nemo | 61.13 | 0.072 | 0.667 | 0.899 | 0.889 | 0.430 | 67,877 | 3.9 s | 100% | |
| openai/gpt-5.6-terra ★ | 57.18 | 0.232 | 0.564 | 0.774 | 0.976 | 0.372 | 10,753 | 3.6 s | 100% | |
| openai/gpt-oss-120b | 54.01 | 0.097 | 0.510 | 0.806 | 0.865 | 0.387 | 42,502 | 8.9 s | 100% | |
| tencent/hy3-preview | 41.35 | 0.010 | 0.438 | 0.651 | 0.657 | 0.299 | 221,504 | 24.2 s | 100% | |
| xiaomi/mimo-v2.5-pro | 32.44 | 0.048 | 0.309 | 0.439 | 0.444 | 0.377 | 33,156 | 18.4 s | 100% | |
| xiaomi/mimo-v2.5 | 25.38 | 0.048 | 0.238 | 0.349 | 0.357 | 0.275 | 25,891 | 19.5 s | 100% | |
| no-opreferencereplies politely, never acts | 24.89 | 0.000 | — | — | 0.990 | — | — | — | — |
* Provisional. factual_grounding is a literal phrase matcher; its column is dimmed because fine-grained ordering should not be read from it. Tokens/success is total tokens spent divided by scenarios actually resolved (task_success = 1.0) — it spans roughly 23× across the table and is the deployment-relevant figure. p50 latency measures the full JSON action loop (up to 8 sequential provider calls), not single-turn voice latency. The no-op scores 0.990 on safety by never acting — never quote safety without task_success next to it.
A model is ranked only if at least 90% of its trials produced a usable trace. The rule is load-bearing: without it, poolside/laguna-m.1:free ranked first on 7% coverage, and two HTTP-404 models ranked at 0.00 as though they had been evaluated and failed. The full story is in the methodology report.
Unmeasured is a different statement from “scored badly” — these six fall below the 90% coverage floor and are not ranked.
| Model | Trial coverage | Why |
|---|---|---|
| nvidia/nemotron-3-super-120b-a12b:free | 54% | Free-tier rate limiting |
| nvidia/nemotron-3-ultra-550b-a55b:free | 52% | Free-tier rate limiting |
| poolside/laguna-m.1:free | 7% | Free-tier rate limiting |
| nex-agi/nex-n2-pro:free | 0% | HTTP 404, moved behind a paywall |
| openai/gpt-oss-120b:free | 0% | HTTP 404, free tier withdrawn |
| openrouter/owl-alpha | 0% | HTTP 404, no endpoints |
What the numbers do and don’t say.
task_success runs from 0.010 to 0.754 — even the leader fails a quarter of scenarios. The most common failure is completing the customer-facing action, then skipping a required bookkeeping call: thoroughness, not reasoning, which is why this ordering does not mirror reasoning benchmarks.
tokens_per_success runs from 9,730 to 221,504. It charges an agent for the budget it burns on failures, which is how a deployment experiences cost. Two similar-looking overall scores can differ by an order of magnitude here.
Safety ranges 0.357–0.913 in this sweep. Only acting before a guard is satisfied counts against it — wrong arguments are scored under tool correctness instead. The no-op still scores 0.990 by never acting — never quote safety alone.
The table above is generated from published, SHA-256-pinned artifacts. Rebuilding it is offline and needs no API key; scoring a model fresh needs an OpenRouter key.
# Score one model against the track (needs OPENROUTER_API_KEY) python scripts/run_openvoicecs.py score-provider \ --provider openrouter --model google/gemini-2.5-flash-lite \ --track text_to_action --trials 3 --json-trace --output report.json # Rebuild this table from the published reports (offline, no API key) python scripts/build_leaderboard.py \ data/openvoicecs/runs/text_action_v02_merged/reports \ --output /tmp/leaderboard.csv