A voice-first AI research lab. Blind panels of human ears judge how agents sound; replayable traces prove whether they resolved the call. And our playground is the dental front desk — the toughest room a voice AI can work.
↓ see the evidenceEverything we publish answers one of three questions — with a measurement, not an opinion.
A vetted, paid panel of native speakers votes blind on pairs of AI voices — four separate questions, never mixed, with hidden controls to catch guessing.
Customer-service agents scored deterministically from replayable traces: did the problem get fixed without breaking policy — and can it prove it?
Our playground is the dental front desk — narrow, high-stakes, measurable. We fine-tune open models there and judge them against the giants.
A vetted panel of paid native speakers hears two AI voices read the same sentence and picks one — blind. Each sitting asks one of four questions and never mixes them: overall preference, sounds human, clear and correct, rhythm and expression. 12,245 comparisons, frozen on 25 July. The two boards below show overall preference, one language each — note how little they share. Open the Voice Arena →
The language gap is the biggest single finding: Cartesia’s sonic-3.5 is 3rd in English and 19th in Turkish, ElevenLabs’ eleven_v3 is 37th and 8th — Gemini 3.1 Flash Puck, top of both boards, is the exception. And the panel is not infallible: 9.3% of hidden same-clip controls still got a confident winner picked.
Voice agents will be judged by what they do, not just how they sound. OpenVoiceCS-Bench replays every tool call an agent makes against a scenario-local state copy and scores eight deterministic metrics — same trace in, same score out. 52 of 58 attempted models are ranked on the text-to-action track; six are published as unmeasured rather than penalised. How the scoring works →
A dental front-desk call is the hardest small domain we know: anxious patients, clinical boundaries, strict privacy — and a measurable outcome at the end. Everything the lab builds gets tested here first. DentesBench scores phone agents on 483 scenarios; the deployment-weighted board counts quality (80%), cost (10%), and latency (10%) — and a fine-tuned open-weights model holds #1. Full methodology →
Next up in the playground: dental-domain STT — accents, phone-line noise, and terminology generic models mangle. The Gemma 4 fine-tuning report covers how the #1 model above was trained.
Every number we publish ships with its methodology and its caveats attached.
Open the OpenVoiceCS-Bench Leaderboard Open the Voice Arena