Patientdesk Labs · OpenVoiceCS-Bench · text_to_action track

OpenVoiceCS-Bench Leaderboard

52 models, ranked on whether they resolved the customer’s problem without breaking policy — scored deterministically from replayable traces.

Methodology report →
52 of 58 models ranked 6 unmeasured 69 × 3 scenarios × trials 1 of 5 tracks swept run text_action_v02_merged July 2026
Before quoting 1 Grounding is provisional — trust the podium, not the mid-table 2 No confidence intervals — a two-point gap is noise 3 One track of five — not a voice-agent ranking Details →
52 models Overall = weighted blend of 8 deterministic metrics no-op floor 24.89 follow-up run click headers to sort
# Model Overall Task Tool SOP Safety Grounding* Tokens/success p50 latency Coverage
oraclereferencecorrect by construction100.001.0001.000
anthropic/claude-fable-588.870.7540.9120.9110.8890.9139,73020.1 s100%
moonshotai/kimi-k3 86.960.7150.9020.9060.9080.8705,75918.9 s100%
openai/gpt-5.6-luna-pro 86.150.8600.8670.8410.9470.82623,21114.0 s100%
z-ai/glm-5.2 81.110.4350.8110.9100.9270.9237,4819.1 s100%
anthropic/claude-opus-4.681.030.4250.8080.9130.8890.93713,54911.3 s100%
openai/gpt-5.6-sol 78.890.3820.7910.9080.9420.8798,07410.5 s100%
google/gemini-3.1-pro-preview77.090.3870.7930.9070.8840.79715,8889.1 s100%
deepseek/deepseek-v3.276.420.2220.7380.9180.9130.95223,4888.2 s100%
google/gemini-3.5-flash75.570.2270.7400.9080.8940.92324,9334.2 s100%
anthropic/claude-opus-4.775.360.2950.7620.9130.9080.81622,6326.5 s100%
anthropic/claude-sonnet-4.575.160.1880.7270.9130.8990.94727,7676.3 s100%
openai/gpt-5.575.040.2080.7340.9080.8990.91821,5365.0 s100%
anthropic/claude-haiku-4.574.860.1880.7270.9150.9030.92728,4483.7 s100%
openai/gpt-5.6-sol-pro 74.800.2220.7380.9080.9470.87977,78416.6 s100%
minimax/minimax-m2.574.600.1840.7200.9160.9080.91829,9278.3 s100%
minimax/minimax-m2.774.160.2370.7310.9070.9030.85521,6259.1 s100%
z-ai/glm-5.174.060.1930.7240.9070.8890.89925,2938.4 s100%
z-ai/glm-573.580.1790.7090.9050.9080.90328,09211.0 s100%
openai/gpt-5.3-codex73.480.1880.7160.8900.8890.90324,7924.8 s100%
qwen/qwen3.7-max73.460.1910.7280.9100.8870.86331,21713.7 s99%
qwen/qwen3.7-plus73.210.1930.7290.9050.8990.84530,12411.3 s100%
google/gemma-4-26b-a4b-it73.060.1880.7270.9030.8990.84527,6124.2 s100%
z-ai/glm-4.5-air72.930.1790.7120.9040.8990.87929,72110.9 s100%
google/gemini-3-flash-preview72.780.2270.7400.9080.8990.78322,3704.1 s100%
google/gemini-2.5-flash-lite72.770.1880.7210.9010.8990.83129,1452.9 s100%
anthropic/claude-sonnet-4.672.300.1880.7270.9130.8990.79728,7166.2 s100%
deepseek/deepseek-v4-flash72.170.1880.7240.9000.8990.81625,30711.0 s100%
moonshotai/kimi-k2.571.890.2370.7280.8970.8840.77320,43516.0 s100%
openai/gpt-5.471.710.1880.6950.8820.8940.85022,8273.4 s100%
google/gemini-3.1-flash-lite-preview71.590.1880.7270.9020.8990.74926,8601.4 s100%
google/gemini-3.1-flash-lite71.490.1880.7270.9050.8990.75826,8272.8 s100%
google/gemma-4-31b-it70.830.1840.7220.9020.8990.75427,72110.8 s100%
openai/gpt-5.4-nano70.590.1550.6590.8710.9080.87027,1552.8 s100%
openai/gpt-5-mini70.390.1690.6880.8920.8890.81628,52814.5 s100%
qwen/qwen3-235b-a22b-250770.120.1880.7270.9030.8840.70524,9594.4 s100%
minimax/minimax-m370.050.1740.7060.9030.8890.74430,4196.6 s100%
deepseek/deepseek-v4-pro69.940.1840.7040.8950.8940.74927,08915.0 s100%
stepfun/step-3.7-flash68.660.1300.6550.8750.8600.85528,9744.5 s100%
openai/gpt-5.6-luna 68.410.6180.6270.7210.9420.5994,4384.8 s100%
openai/gpt-4o-mini68.140.1690.7080.8980.8990.64727,9193.8 s100%
inclusionai/ling-2.6-flash67.780.1260.6710.8770.8740.77335,6064.2 s100%
google/gemini-2.5-flash67.200.1880.7240.9010.8990.55626,4343.4 s100%
anthropic/claude-opus-4.866.180.0820.4800.9060.8840.82971,5675.7 s100%
openai/gpt-5.4-mini65.740.1840.6550.8510.8990.62823,0212.2 s100%
moonshotai/kimi-k2.665.630.2270.6590.8390.8310.68614,66211.7 s100%
openai/gpt-5.6-terra-pro 63.890.2410.7030.8740.9710.39669,5199.3 s100%
mistralai/mistral-nemo61.130.0720.6670.8990.8890.43067,8773.9 s100%
openai/gpt-5.6-terra 57.180.2320.5640.7740.9760.37210,7533.6 s100%
openai/gpt-oss-120b54.010.0970.5100.8060.8650.38742,5028.9 s100%
tencent/hy3-preview41.350.0100.4380.6510.6570.299221,50424.2 s100%
xiaomi/mimo-v2.5-pro32.440.0480.3090.4390.4440.37733,15618.4 s100%
xiaomi/mimo-v2.525.380.0480.2380.3490.3570.27525,89119.5 s100%
no-opreferencereplies politely, never acts24.890.0000.990

* Provisional. factual_grounding is a literal phrase matcher; its column is dimmed because fine-grained ordering should not be read from it. Tokens/success is total tokens spent divided by scenarios actually resolved (task_success = 1.0) — it spans roughly 23× across the table and is the deployment-relevant figure. p50 latency measures the full JSON action loop (up to 8 sequential provider calls), not single-turn voice latency. The no-op scores 0.990 on safety by never acting — never quote safety without task_success next to it.

Unmeasured Models & How To Read The Numbers

A model is ranked only if at least 90% of its trials produced a usable trace. The rule is load-bearing: without it, poolside/laguna-m.1:free ranked first on 7% coverage, and two HTTP-404 models ranked at 0.00 as though they had been evaluated and failed. The full story is in the methodology report.

Six models we failed to measure

Unmeasured is a different statement from “scored badly” — these six fall below the 90% coverage floor and are not ranked.

ModelTrial coverageWhy
nvidia/nemotron-3-super-120b-a12b:free54%Free-tier rate limiting
nvidia/nemotron-3-ultra-550b-a55b:free52%Free-tier rate limiting
poolside/laguna-m.1:free7%Free-tier rate limiting
nex-agi/nex-n2-pro:free0%HTTP 404, moved behind a paywall
openai/gpt-oss-120b:free0%HTTP 404, free tier withdrawn
openrouter/owl-alpha0%HTTP 404, no endpoints
Three notes worth more than the ranking

What the numbers do and don’t say.

Task success is the real discriminator

task_success runs from 0.010 to 0.754 — even the leader fails a quarter of scenarios. The most common failure is completing the customer-facing action, then skipping a required bookkeeping call: thoroughness, not reasoning, which is why this ordering does not mirror reasoning benchmarks.

Token efficiency spans ~23×

tokens_per_success runs from 9,730 to 221,504. It charges an agent for the budget it burns on failures, which is how a deployment experiences cost. Two similar-looking overall scores can differ by an order of magnitude here.

Safety is a don’t-do-harm measure — read it in context

Safety ranges 0.357–0.913 in this sweep. Only acting before a guard is satisfied counts against it — wrong arguments are scored under tool correctness instead. The no-op still scores 0.990 by never acting — never quote safety alone.

Rebuild This Table Yourself

The table above is generated from published, SHA-256-pinned artifacts. Rebuilding it is offline and needs no API key; scoring a model fresh needs an OpenRouter key.

# Score one model against the track (needs OPENROUTER_API_KEY)
python scripts/run_openvoicecs.py score-provider \
  --provider openrouter --model google/gemini-2.5-flash-lite \
  --track text_to_action --trials 3 --json-trace --output report.json

# Rebuild this table from the published reports (offline, no API key)
python scripts/build_leaderboard.py \
  data/openvoicecs/runs/text_action_v02_merged/reports \
  --output /tmp/leaderboard.csv