Voice AITTSSTTAgentsBenchmarks+

Machines talk.
We keep score.

A voice-first AI research lab. Blind panels of human ears judge how agents sound; replayable traces prove whether they resolved the call. We build Turkish speech models and measure them the same way — and our playground is the dental front desk, the toughest room a voice AI can work.

Blind votes
12,245blind votes
Models ranked
52agents ranked
Dental scenarios
483dental scenarios
Open voice corpus
5.008 hopen voice corpus
New · Story and technical report

Antalia 1, our open Turkish voice

Weights, code, a consented voice corpus and the paper — and the listening test it failed.

01 What we measure

Three questions, answered with a measurement, not an opinion.

Everything we publish answers one of them. Pick a question, then a piece of work; the numbers and their caveats are one click away.

03 Sounds human · Release 2026.2

Two voices, one sentence, no names. Pick one.

A vetted panel of paid native speakers hears two AI voices read the same sentence and picks one — blind. Each sitting asks one of four questions and never mixes them: overall preference, sounds human, clear and correct, rhythm and expression. Frozen on 25 July. The boards show overall preference, one language each — note how little they share.

Overall preference · top 6 per language
English2185 votes · 103 listeners
1Google GeminiPuck1138
2Smallestnolan1125
3Cartesiadb6b0ed51124
4Speechifygeffen_321121
5xAIaltair1107
6Speechifydominic_321100
Turkish851 votes · 44 listeners
1Google GeminiPuck1302
2Google GeminiKore1195
3Google GeminiKore1185
4Google GeminiPuck1183
5xAIara1157
6MiniMaxTurkish_CalmWoman1151
95% range score bottom of leader’s range
Blind comparisons
12,245

Frozen on 25 July 2026, English and Turkish.

Paid, vetted listeners
224

Every one cleared onboarding before they could vote. No public vote button.

Voices · companies
72from 16

Four questions, never mixed in one sitting.

The language gap is the biggest single finding. Cartesia’s sonic-3.5 is 3rd in English and 19th in Turkish; ElevenLabs’ eleven_v3 is 37th and 8th — Gemini 3.1 Flash Puck, top of both boards, is the exception. And the panel is not infallible: 9.3% of hidden same-clip controls still got a confident winner picked.

04 Resolves the call · Latest release

Did the agent actually fix it?

Voice agents will be judged by what they do, not just how they sound. OpenVoiceCS-Bench replays every tool call an agent makes against a scenario-local state copy and scores eight deterministic metrics — same trace in, same score out. 52 of 58 attempted models are ranked on the text-to-action track; six are published as unmeasured rather than penalised.

Top 5 of 52 — text_to_actionoverall / 100 · dashed line = do-nothing floor
1claude-fable-5anthropic88.87
2kimi-k3moonshotai86.96
3gpt-5.6-luna-proopenai86.15
4glm-5.2z-ai81.11
5claude-opus-4.6anthropic81.03
69 × 3scenarios × trials per model, replayed deterministically
24.89the do-nothing floor — the dashed line in every bar
~23×spread in tokens burned per resolved case
6models unmeasured — published as such, not scored 0
05 The playground

The dental front desk is our proving ground.

A dental front-desk call is the hardest small domain we know: anxious patients, clinical boundaries, strict privacy — and a measurable outcome at the end. Everything the lab builds gets tested here first. DentesBench scores phone agents on 483 scenarios; the deployment-weighted board counts quality (80%), cost (10%) and latency (10%) — and a fine-tuned open-weights model holds #1. Full methodology →

Top 5 — deployment-weighted ranking · April 2026
#ModelEmpathySafetyAccuracyBrevityToneV2PassLatencyCost / resp
1Gemma 4 31B (OpenRouter)6.89.67.38.76.88.1875%2.4s$0.00006
2GLM-5 Turbo (OpenRouter)7.09.77.78.77.38.0184%3.5s$0.00074
3GPT-5.46.99.68.08.37.07.9286%1.5s$0.00161
4Claude Sonnet 4.67.49.57.78.37.77.8688%2.3s$0.00189
5Claude Opus 4.67.59.67.88.37.87.4191%3.1s$0.00318

The Gemma 4 fine-tuning report covers how the #1 model above was trained. Next up in the playground: dental-domain speech recognition — accents, phone-line noise, and the terminology generic models mangle.

Research posts

Post 01 of 8