Patientdesk Labs · the research arm of Patientdesk.ai

Machines talk.
We keep score.

A voice-first AI research lab. Blind panels of human ears judge how agents sound; replayable traces prove whether they resolved the call. And our playground is the dental front desk — the toughest room a voice AI can work.

see the evidence
What We Do

Three Questions Drive the Lab

Everything we publish answers one of three questions — with a measurement, not an opinion.

01

Does it sound human?

A vetted, paid panel of native speakers votes blind on pairs of AI voices — four separate questions, never mixed, with hidden controls to catch guessing.

12,245comparisons across 72 voices from 16 companies, in English and Turkish
AI Voice Arena →
02

Did it resolve the call?

Customer-service agents scored deterministically from replayable traces: did the problem get fixed without breaking policy — and can it prove it?

52models ranked, 25.38–88.87 against a 24.89 do-nothing floor
OpenVoiceCS-Bench Leaderboard →
03

Can small beat frontier?

Our playground is the dental front desk — narrow, high-stakes, measurable. We fine-tune open models there and judge them against the giants.

#1a fine-tuned Gemma 4 31B tops our deployment-weighted board
Gemma 4 fine-tuning →
Measured, not vibedDeterministic scoring, replayable traces, blind panels — never a judge’s mood.
Caveats in printLimitations publish next to the numbers. Unmeasured is not “scored badly”.
Playground, not nicheDental is the proving ground, not the ceiling — what survives it transfers.
Sounds Human · Release 2026.2

AI Voice Arena

A vetted panel of paid native speakers hears two AI voices read the same sentence and picks one — blind. Each sitting asks one of four questions and never mixes them: overall preference, sounds human, clear and correct, rhythm and expression. 12,245 comparisons, frozen on 25 July. The two boards below show overall preference, one language each — note how little they share. Open the Voice Arena →

English — overall preference2185 votes · 103 listeners
1 Google GeminiPuck 1138
2 Smallestnolan 1125
3 Cartesiadb6b0ed5 1124
4 Speechifygeffen_32 1121
5 xAIaltair 1107
6 Speechifydominic_32 1100
Turkish — overall preference851 votes · 44 listeners
1 Google GeminiPuck 1302
2 Google GeminiKore 1195
3 Google GeminiKore 1185
4 Google GeminiPuck 1183
5 xAIara 1157
6 MiniMaxTurkish_CalmWoman 1151
95% range score bottom of leader’s range

The language gap is the biggest single finding: Cartesia’s sonic-3.5 is 3rd in English and 19th in Turkish, ElevenLabs’ eleven_v3 is 37th and 8th — Gemini 3.1 Flash Puck, top of both boards, is the exception. And the panel is not infallible: 9.3% of hidden same-clip controls still got a confident winner picked.

Open the Arena Read the Report
Resolves the Call · Latest Release

OpenVoiceCS-Bench: Did the Agent Actually Fix It?

Voice agents will be judged by what they do, not just how they sound. OpenVoiceCS-Bench replays every tool call an agent makes against a scenario-local state copy and scores eight deterministic metrics — same trace in, same score out. 52 of 58 attempted models are ranked on the text-to-action track; six are published as unmeasured rather than penalised. How the scoring works →

Top 5 of 52 — text_to_actionoverall / 100
1 claude-fable-5anthropic 88.87
2 kimi-k3moonshotai 86.96
3 gpt-5.6-luna-proopenai 86.15
4 glm-5.2z-ai 81.11
5 claude-opus-4.6anthropic 81.03
69 × 3scenarios × trials per model, replayed deterministically
24.89the do-nothing floor — the dashed line in every bar
~23×spread in tokens burned per resolved case
6models unmeasured — published as such, not scored 0
See All 52 Models Read the Methodology
The Playground

Dental Clinics: Our Proving Ground

A dental front-desk call is the hardest small domain we know: anxious patients, clinical boundaries, strict privacy — and a measurable outcome at the end. Everything the lab builds gets tested here first. DentesBench scores phone agents on 483 scenarios; the deployment-weighted board counts quality (80%), cost (10%), and latency (10%) — and a fine-tuned open-weights model holds #1. Full methodology →

Top 5 — Deployment-Weighted Ranking April 2026
#Model EmpathySafety AccuracyBrevity ToneV2 Pass Latency Cost/resp
1Gemma 4 31B (OpenRouter) 6.89.6 7.38.7 6.88.18 75% 2.4s $0.00006
2GLM-5 Turbo (OpenRouter) 7.09.7 7.78.7 7.38.01 84% 3.5s $0.00074
3GPT-5.4 6.99.6 8.08.3 7.07.92 86% 1.5s $0.00161
4Claude Sonnet 4.6 7.49.5 7.78.3 7.77.86 88% 2.3s $0.00189
5Claude Opus 4.6 7.59.6 7.88.3 7.87.41 91% 3.1s $0.00318

Next up in the playground: dental-domain STT — accents, phone-line noise, and terminology generic models mangle. The Gemma 4 fine-tuning report covers how the #1 model above was trained.

Start with a Leaderboard

Every number we publish ships with its methodology and its caveats attached.

Open the OpenVoiceCS-Bench Leaderboard Open the Voice Arena