Machines talk.
We keep score.
A voice-first AI research lab. Blind panels of human ears judge how agents sound; replayable traces prove whether they resolved the call. We build Turkish speech models and measure them the same way — and our playground is the dental front desk, the toughest room a voice AI can work.
- Blind votes
- 12,245blind votes
- Models ranked
- 52agents ranked
- Dental scenarios
- 483dental scenarios
- Open voice corpus
- 5.008 hopen voice corpus
Antalia 1, our open Turkish voice
Weights, code, a consented voice corpus and the paper — and the listening test it failed.
Three questions, answered with a measurement, not an opinion.
Everything we publish answers one of them. Pick a question, then a piece of work; the numbers and their caveats are one click away.
We measure other people’s voices. This one is ours.
Antalia 1 is our first Turkish text-to-speech model, and we released all of it: the weights, the code, a five-hour consented voice corpus, the evaluation suite and a technical report. It is no longer developed; the story and the report explain what worked, what did not, and why a native listener could still tell it apart from the real voice.
Antalia 1
Our first Turkish voice, released in full. Development stopped; the report is candid about why.
Two voices, one sentence, no names. Pick one.
A vetted panel of paid native speakers hears two AI voices read the same sentence and picks one — blind. Each sitting asks one of four questions and never mixes them: overall preference, sounds human, clear and correct, rhythm and expression. Frozen on 25 July. The boards show overall preference, one language each — note how little they share.
Frozen on 25 July 2026, English and Turkish.
Every one cleared onboarding before they could vote. No public vote button.
Four questions, never mixed in one sitting.
The language gap is the biggest single finding. Cartesia’s sonic-3.5 is 3rd in English and 19th in Turkish; ElevenLabs’ eleven_v3 is 37th and 8th — Gemini 3.1 Flash Puck, top of both boards, is the exception. And the panel is not infallible: 9.3% of hidden same-clip controls still got a confident winner picked.
Did the agent actually fix it?
Voice agents will be judged by what they do, not just how they sound. OpenVoiceCS-Bench replays every tool call an agent makes against a scenario-local state copy and scores eight deterministic metrics — same trace in, same score out. 52 of 58 attempted models are ranked on the text-to-action track; six are published as unmeasured rather than penalised.
The dental front desk is our proving ground.
A dental front-desk call is the hardest small domain we know: anxious patients, clinical boundaries, strict privacy — and a measurable outcome at the end. Everything the lab builds gets tested here first. DentesBench scores phone agents on 483 scenarios; the deployment-weighted board counts quality (80%), cost (10%) and latency (10%) — and a fine-tuned open-weights model holds #1. Full methodology →
| # | Model | Empathy | Safety | Accuracy | Brevity | Tone | V2 | Pass | Latency | Cost / resp |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemma 4 31B (OpenRouter) | 6.8 | 9.6 | 7.3 | 8.7 | 6.8 | 8.18 | 75% | 2.4s | $0.00006 |
| 2 | GLM-5 Turbo (OpenRouter) | 7.0 | 9.7 | 7.7 | 8.7 | 7.3 | 8.01 | 84% | 3.5s | $0.00074 |
| 3 | GPT-5.4 | 6.9 | 9.6 | 8.0 | 8.3 | 7.0 | 7.92 | 86% | 1.5s | $0.00161 |
| 4 | Claude Sonnet 4.6 | 7.4 | 9.5 | 7.7 | 8.3 | 7.7 | 7.86 | 88% | 2.3s | $0.00189 |
| 5 | Claude Opus 4.6 | 7.5 | 9.6 | 7.8 | 8.3 | 7.8 | 7.41 | 91% | 3.1s | $0.00318 |
The Gemma 4 fine-tuning report covers how the #1 model above was trained. Next up in the playground: dental-domain speech recognition — accents, phone-line noise, and the terminology generic models mangle.
Research posts
Post 01 of 8