Patientdesk Labs is the research arm of Patientdesk.ai. We build benchmarks, fine-tune models, and develop domain-specific tools so that when an AI answers the phone at a dental office, it actually works.
Dental clinic AI has unique requirements that generic models don't meet. Our research focuses on four areas where domain-specific work makes the biggest difference — from the model that reasons, to the voice the patient actually hears.
The first benchmark for evaluating LLMs as dental clinic phone agents. 483 scenarios across 10 categories, scoring empathy, clinical safety, accuracy, brevity, and tone — plus a deployment-weighted leaderboard.
Read the paper →Soul-document-driven fine-tuning of Gemma 4 31B with Opus 4.6 as judge. After one iteration of SFT + DPO: 8.46 on DentesBench — beating every frontier API model including Opus itself.
Read the paper →A blind listening leaderboard for AI voices. People compare two voices reading the same sentence without seeing the brand. 44 English and 28 Turkish voices, 16 companies, 12,245 comparisons — and the same voice can rank 3rd in English and 19th in Turkish.
Open the arena →Adapting speech-to-text models for dental clinic phone audio. Patient calls with accents, background noise, and dental terminology that generic models consistently get wrong — "prophylaxis" shouldn't become "prophy lax is."
Paper coming soonA dental receptionist AI has to be warm without accidentally diagnosing, efficient without being cold, and helpful without overstepping clinical boundaries. No off-the-shelf model gets this right consistently.
People listen to two AI voices reading the same sentence and pick one, without ever seeing who made either. 12,245 comparisons across 44 English and 28 Turkish voices from 16 companies. The two boards below ask the same question in two languages — note how little they share. Open the full arena →
With 4.6× the data of our pilot, English now separates: 20 of the 43 challengers sit clearly behind the leader. The cross-language gap is the story — Speechify tops the English board and finishes 8th of 10 companies in Turkish, while ElevenLabs is 12th of 16 in English and 3rd in Turkish.
Eight models evaluated on 483 dental phone agent scenarios. The v2 score weights quality (80%), cost (10%), and latency (10%) to reflect real deployment constraints. Full methodology →
Read the full DentesBench paper for methodology, results, and what we've learned about the tradeoffs in dental AI.
Read the Paper Visit Patientdesk.ai