At A Glance
What This Report Says In One Minute

People listened to pairs of AI voices reading the same sentence and picked the one they preferred, never seeing who made either. This is round two: 12,245 comparisons, 4.6 times the size of our pilot.

12,245 blind comparisons 224 listeners, 44 English and 28 Turkish voices from 16 companies including ElevenLabs, OpenAI, Resemble AI and Rime. Every voice got enough votes to earn a rank.
English separates — at the bottom Our pilot could not separate the English field at all. With 4.6× the data, 20 of the 43 challengers are clearly behind the leader and the typical error bar halved to 140 points. The top 24 remain a pack we cannot order.
The language gap is the story Speechify tops English at 63.6% and finishes 8th of 10 in Turkish at 37.3%. ElevenLabs goes the other way: 12th of 16 in English, 3rd in Turkish.
We corrected ourselves The pilot suggested the four listening questions measured different things. With more data they largely agree. That finding was small-sample noise and we are retiring it.
Contents
Table Of Contents

What changed since the pilot, how the voting works, what round two found, and what we still cannot tell you.

Summary

Most AI voice companies show you a demo page. Very few tell you how they tested, how confident they are, or how the voice holds up in a language other than English. We needed to pick a voice for a phone agent that talks to real patients, could not find that information anywhere, so we measured it ourselves.

Round 2026.2 covers 44 English voices and 28 Turkish voices from 16 companies. Two voices read the same sentence, a listener picks one, and nobody sees a brand name until voting closes. Every score ships with an error bar, and the full results file is a public download.

Key finding: the company that wins English is near the bottom in Turkish, and the one that is mid-table in English comes third in Turkish. Picking a voice on English evidence and shipping it in another language is a coin flip.

What Changed Since The Pilot

Our July pilot was honest about being too small: 2,675 votes across 40 English voices meant 21 of the 22 ranked entries overlapped the leader, and we said plainly that we could not rank English. Round two is 4.6 times bigger, and the picture is different in three specific ways.

Pilot 2026.1Round 2026.2
Votes counted2,67512,245
Listeners44224
Voices66 from 13 companies72 from 16 companies
Voices with enough votes to rank48 of 66all 72
Typical English error bar340 points140 points
English voices clearly behind the leader1 of 2120 of 43

First, English separates now — though not at the top. The error bars roughly halved, and 20 of the 44 voices are now measurably behind the leading pack instead of statistically tangled with it. The other 24 still overlap each other.

Second, the cross-language finding survived and got stronger. It was the one pilot result we said we believed, and more data has not softened it.

Third, and least comfortably, one pilot finding did not survive. We reported that the four listening questions measured genuinely different things, pointing at a near-zero correlation between “sounds human” and overall preference in Turkish. With four times the data that correlation is 0.89. It was noise. We were measuring our own sample size.

How closely each question tracks overall preferenceEnglish pilotEnglish nowTurkish pilotTurkish now
Sounds human0.460.610.080.89
Clear and correct0.520.760.680.95
Rhythm and expression0.410.760.540.87

Splitting the questions is still worth doing — it is how you find out they agree, and it keeps the option open for a future round where they might not. But we no longer claim they measure different things, and anyone who quoted our pilot on that should stop.

How It Works

The point of the setup is to test voices, not brand loyalty or listener patience. Nothing about the method changed between rounds, which is what makes the two comparable.

Both voices read the same sentence

Every comparison uses identical text, from a fixed bank of 24 sentences — 12 English, 12 Turkish — covering plain statements, questions, lists, names, numbers and dates, abbreviations, long winding sentences, and emotional lines.

Sentence bank2 of 24
English — numbers, dates, money
“Your reservation is confirmed for September eighteenth at 6:45 p.m., and the total is thirty-two dollars.”
Turkish — names and places
“Şişli'de Zeynep'le buluşan Can, ardından metroya binip Yenikapı yönüne doğru ilerledi.”
The two sets are not translations. Each language gets sentences that are hard in that language — Turkish suffix chains and vowel harmony, English abbreviations and borrowed words.

Nobody sees the brand names

The listener's browser gets an audio file and its length, and nothing else. File names are random IDs, so no company or model name appears in the page, the network traffic, or anywhere a curious listener could dig. Buttons say “First voice” and “Second voice”. The map from clip to company lives in a separate file that is only opened after voting closes.

One question at a time, and you have to listen

A listener gets 25 comparisons in a sitting, all asking the same one of four questions. The server rejects any vote where the listener did not play at least 60% of both clips, so a modified browser cannot fake it. The two players cannot run at once. Replays are unlimited. There is no skip — you pick a voice, say they sound the same, or report an audio problem. Nobody sees the same pair twice, and nobody votes more than once on a given question in a given language.

Hidden checks

Three of the 25 comparisons are traps: the same clip played on both sides, at positions 5, 13 and 21. They look completely normal, both sides getting their own download link for the identical file. The honest answers are “about the same” or “audio problem”; picking a winner counts as a miss. Traps are paid like any other comparison, which is what keeps them hidden, and they never count toward anyone's score.

How the scores work

We use a standard head-to-head method called Bradley–Terry, which turns a pile of pairwise results into a strength number per voice. A tie counts as half a win. Scores are centred so the average voice in each language sits at 1000. Error bars come from resampling listeners rather than individual clicks, because one person voting 25 times is not 25 independent opinions. A voice needs 20 votes to earn a rank; in this round every voice cleared that easily.

What We Collected

Voting ran from 20 to 25 July 2026 and closed at 556 completed sittings.

Figure 1
The totals
Only finished sittings count. Nothing was thrown out as broken or duplicated.
12,245votes counted
224listeners
556finished sittings
864audio clips (80 min)
164 English listeners and 58 Turkish, giving 103 and 44 behind each individual leaderboard.
Figure 2
How people answered
All 12,432 usable responses, before the audio-problem reports are set aside.
Picked a winner — 10,714 (86.2%) About the same — 1,531 (12.3%) Audio problem — 187 (1.5%)
The tie rate fell from 15.1% in the pilot to 12.5% of counted votes, which is what you expect when a wider field puts more clearly-unequal pairs in front of people.

English Results

This is what the pilot could not give us, with one caveat. The field is 44 voices, the bottom 20 are now measurably behind, and the top 24 are still a crowd.

Figure 3
English overall preference — top 20 of 44
The dashed line sits at the bottom of the leader's range. A voice is only clearly behind the leader if its whole bar falls to the left of that line. Across the full 44, 20 do.
500750100012501500
1Google Gemini Puck
1138
2Smallest nolan
1125
3Cartesia db6b0ed5
1124
4Speechify geffen_32
1121
5xAI altair
1107
6Speechify dominic_32
1100
7Resemble AI 6e37aa15
1095
8Google Gemini Kore
1092
9Cartesia 630ed21c
1083
10Gradium YTpq7expH9539ERJ
1077
11MiniMax English_Trustworth_Man
1063
12Smallest chelsea
1061
13Cartesia db6b0ed5
1061
14Resemble AI 819fcc57
1060
15ElevenLabs iP95p4xoKVk53GoZ742B
1059
16StepFun lively-girl
1043
17Hume AI Female Meditation Guide
1038
18Gradium LFZvm12tW_z0xfGo
1030
19xAI ara
1029
20Async cca0e076
1023
95% rangescorebottom of leader’s rangeclearly behind
Typical bar is 140 points wide, against 340 in the pilot. First to last spans 522 points.

Google Gemini's gemini-3.1-flash-tts-preview leads at 1138, but its range overlaps every voice down to rank 24. The honest reading is that the English field splits in two: a broad leading pack of 24 voices we cannot order internally, and 20 that are clearly behind it. That is a real gain over the pilot, where only one voice was separable from the leader — but it is separation of the bottom half, not a winner at the top.

RankCompanyModelVoiceScore95% rangeVotes
1Google Geminigemini-3.1-flash-tts-previewPuck1137.81075–120593
2Smallestlightning_v3.1_pronolan1125.41055–120986
3Cartesiasonic-3.5-2026-05-04db6b0ed51123.71069–1184100
4Speechifysimba-3.2geffen_321121.21051–1191107
5xAIgrok-voice-ttsaltair1107.31036–117898
6Speechifysimba-3.2dominic_321099.81036–116198
7Resemble AIchatterbox6e37aa151095.11028–117597
8Google Geminigemini-3.1-flash-tts-previewKore1092.41033–1161116
9Cartesiasonic-3.5-2026-05-04630ed21c1083.21015–114889
10GradiumdefaultYTpq7expH9539ERJ1077.0999–1153112
11MiniMaxspeech-2.8-hdEnglish_Trustworth_Man1062.71005–112788
12Smallestlightning_v3.1_prochelsea1061.41007–1117109
13Cartesiasonic-3db6b0ed51060.5997–1138111
14Resemble AIchatterbox819fcc571060.41003–1121108
15ElevenLabseleven_v3iP95p4xoKVk53GoZ742B1058.7988–114091
16StepFunstep-tts-2lively-girl1042.6971–110898
17Hume AI1Female Meditation Guide1037.5981–1102104
18GradiumdefaultLFZvm12tW_z0xfGo1029.9966–108895
19xAIgrok-voice-ttsara1028.8974–1092108
20Asyncasync_flash_v1.0cca0e0761023.3958–1091106
21Google Geminigemini-2.5-pro-preview-ttsKore1022.4973–1079111
22MiniMaxspeech-2.8-hdEnglish_CalmWoman1021.2951–1091110
23Asyncasync_flash_v1.5e5a67eaf1018.4942–109283
24Asyncasync_flash_v1.0317bf8051007.3931–107690
25OpenAIgpt-4o-mini-ttsnova1005.2936–1072101
26MiniMaxspeech-2.8-turboEnglish_CalmWoman1001.3932–1067114
27Hume AI2Female Meditation Guide992.7931–1061106
28Cartesiasonic-3630ed21c974.5891–104893
29Rimecodamasonry969.1894–103677
30Fish Audios2-proe3cd3841956.8885–1031108
31Asyncasync_flash_v1.50aef6559956.2885–1017114
32Smallestlightning_v3.1liam955.9882–103197
33Smallestlightning_v3.1avery954.7884–102197
34Fish Audios2.1-pro-freec2623f0c949.3884–1015104
35OpenAIgpt-4o-mini-ttsonyx940.4860–102484
36Rimecodaastra936.2862–997107
37ElevenLabseleven_v3Xb7hH8MSUJpSbSDYk0k2915.7842–988101
38Google Geminigemini-2.5-pro-preview-ttsPuck899.6825–97792
39StepFunstep-tts-2magnetic-voiced-male891.2814–96692
40Inworldinworld-tts-2Bianca883.8815–938109
41Inworldinworld-tts-2Callum874.5799–94687
42Inworldinworld-tts-1.5-maxBianca865.5793–936109
43Inworldinworld-tts-1.5-maxCallum863.1775–94384
44Fish Audios2.1-pro-freec5f56a6c616.2399–72686

Turkish Results

Turkish has a smaller field and a much clearer top. 19 of 27 challengers are clearly behind the leader.

Figure 4
Turkish overall preference — all 28 voices
Same construction as Figure 3. Google takes the top four positions outright.
500750100012501500
1Google Gemini Puck
1302
2Google Gemini Kore
1195
3Google Gemini Kore
1185
4Google Gemini Puck
1183
5xAI ara
1157
6MiniMax Turkish_CalmWoman
1151
7ElevenLabs iP95p4xoKVk53GoZ742B
1147
8ElevenLabs Xb7hH8MSUJpSbSDYk0k2
1133
9Resemble AI 64ad5770
1128
10MiniMax Turkish_CalmWoman
1121
11Resemble AI 45b91687
1093
12Fish Audio 67890dc7
1065
13xAI altair
1062
14MiniMax Turkish_Trustworthyman
1048
15Fish Audio 818b100d
1017
16Async 317bf805
967
17Speechify ayse
921
18Cartesia 630ed21c
906
19Cartesia db6b0ed5
900
20Speechify arda
893
21Cartesia db6b0ed5
877
22Speechify baris
876
23OpenAI onyx
827
24OpenAI nova
825
25Async cca0e076
808
26Speechify aylin
763
27Cartesia 630ed21c
730
28Fish Audio fe96529b
721
95% rangescorebottom of leader’s rangeclearly behind
Google Gemini gemini-3.1-flash-tts-preview Puck leads at 1302 with a range of 1219–1417, clear of every rival.
RankCompanyModelVoiceScore95% rangeVotes
1Google Geminigemini-3.1-flash-tts-previewPuck1301.51219–141758
2Google Geminigemini-3.1-flash-tts-previewKore1194.91094–130266
3Google Geminigemini-2.5-pro-preview-ttsKore1184.71105–128664
4Google Geminigemini-2.5-pro-preview-ttsPuck1183.41103–129360
5xAIgrok-voice-ttsara1157.41083–124861
6MiniMaxspeech-2.8-hdTurkish_CalmWoman1150.61072–125266
7ElevenLabseleven_v3iP95p4xoKVk53GoZ742B1146.61068–123265
8ElevenLabseleven_v3Xb7hH8MSUJpSbSDYk0k21132.51044–124564
9Resemble AIchatterbox-multilingual64ad57701128.41048–122260
10MiniMaxspeech-2.8-turboTurkish_CalmWoman1120.71047–121662
11Resemble AIchatterbox-multilingual45b916871093.41021–117457
12Fish Audios2.1-pro-free67890dc71065.2984–114855
13xAIgrok-voice-ttsaltair1061.5964–114255
14MiniMaxspeech-2.8-hdTurkish_Trustworthyman1048.1980–112856
15Fish Audios2-pro818b100d1017.0936–111960
16Asyncasync_flash_v1.0317bf805967.1884–105161
17Speechifysimba-multilingualayse920.7832–100856
18Cartesiasonic-3.5-2026-05-04630ed21c906.1810–97966
19Cartesiasonic-3.5-2026-05-04db6b0ed5900.0763–99253
20Speechifysimba-multilingualarda893.2795–98860
21Cartesiasonic-3db6b0ed5876.7779–96864
22Speechifysimba-multilingualbaris876.3759–96755
23OpenAIgpt-4o-mini-ttsonyx826.5719–92863
24OpenAIgpt-4o-mini-ttsnova825.4729–90259
25Asyncasync_flash_v1.0cca0e076808.0694–90260
26Speechifysimba-multilingualaylin763.0623–86666
27Cartesiasonic-3630ed21c730.1619–81561
28Fish Audios2.1-pro-freefe96529b721.0578–81769

The Language Gap

Sixteen of the voices were put in front of listeners in both languages — identical company, model and voice ID. That is the cleanest possible test of whether quality travels, because nothing changes except the language being read.

It does not travel.

Voice (identical in both languages)EnglishTurkishDirection
Cartesia db6b0ed5#3 of 44#19 of 28falls in Turkish
Cartesia db6b0ed5#13 of 44#21 of 28falls in Turkish
Async cca0e076#20 of 44#25 of 28falls in Turkish
Cartesia 630ed21c#9 of 44#18 of 28falls in Turkish
xAI altair#5 of 44#13 of 28falls in Turkish
Cartesia 630ed21c#28 of 44#27 of 28falls in Turkish
OpenAI nova#25 of 44#24 of 28falls in Turkish
Async 317bf805#24 of 44#16 of 28falls in Turkish
OpenAI onyx#35 of 44#23 of 28falls in Turkish
Google Gemini Puck#1 of 44#1 of 28falls in Turkish
ElevenLabs iP95p4xoKVk53GoZ742B#15 of 44#7 of 28rises in Turkish
Google Gemini Kore#8 of 44#2 of 28rises in Turkish
xAI ara#19 of 44#5 of 28rises in Turkish
Google Gemini Kore#21 of 44#3 of 28rises in Turkish
ElevenLabs Xb7hH8MSUJpSbSDYk0k2#37 of 44#8 of 28rises in Turkish
Google Gemini Puck#38 of 44#4 of 28rises in Turkish

One important exception, stated plainly: Google Gemini's gemini-3.1-flash-tts-preview Puck is ranked 1st in both languages. A voice can be good in two languages at once. The point is that most are not, and you cannot tell which from an English score.

Pooling all four questions per company gives each one enough comparisons for the cross-language picture to resolve sharply. This is the most useful thing in the report.

Figure 5a
English — 16 companies
All four questions pooled; a tie counts as half a win. 50% is an average voice against this lineup.
Speechify
63.6%
xAI
60.1%
Cartesia
57.7%
Smallest
57.4%
Google Gemini
57.2%
Resemble AI
54.3%
MiniMax
52.6%
Gradium
50.6%
StepFun
49.9%
Hume AI
49.3%
OpenAI
46.0%
ElevenLabs
45.4%
Async
43.5%
Rime
42.5%
Fish Audio
37.3%
Inworld
36.1%
Figure 5b
Turkish — 10 companies
Same construction. Different lineup, so compare positions between the two rather than raw percentages.
Google Gemini
72.7%
xAI
65.3%
ElevenLabs
64.4%
Resemble AI
61.2%
MiniMax
57.5%
Async
40.2%
Fish Audio
39.4%
Speechify
37.3%
Cartesia
35.5%
OpenAI
30.3%

The two highlighted companies move in opposite directions, and neither move is small.

CompanyEnglishTurkishWhat happened
Speechify63.6% — 1st of 1637.3% — 8th of 10top to near-bottom
ElevenLabs45.4% — 12th of 1664.4% — 3rd of 10near-bottom to top
Cartesia57.7% — 3rd of 1635.5% — 9th of 10top to bottom
Google Gemini57.2% — 5th of 1672.7% — 1st of 10strong to dominant
OpenAI46.0% — 11th of 1630.3% — 10th of 10weak in both
xAI60.1% — 2nd of 1665.3% — 2nd of 10the only steady one

Cartesia repeats its pilot result almost exactly, which is the best evidence we have that the effect is real rather than an artefact of one round. Speechify is the new and starker case: first in English, eighth of ten in Turkish.

Six of the sixteen companies fielded nothing in Turkish at all — Smallest, Gradium, StepFun, Hume AI, Rime and Inworld. Turkish speakers choose from 28 voices against 44, and four of the top ten Turkish slots belong to Google.

If your product is not in English, an English benchmark is not evidence. Two of the top three English companies land in the bottom third in Turkish, and the third-best Turkish company is twelfth in English.
#CompanyWin rate95% rangeComparisonsVoices
1Speechify63.6%60.2–66.9%7992
2xAI60.1%56.8–63.5%8212
3Cartesia57.7%55.3–60.2%15574
4Smallest57.4%55.0–59.8%15764
5Google Gemini57.2%54.8–59.6%15904
6Resemble AI54.3%50.8–57.8%7792
7MiniMax52.6%49.9–55.4%12293
8Gradium50.6%47.2–54.0%8182
9StepFun49.9%46.3–53.5%7442
10Hume AI49.3%46.0–52.6%8632
11OpenAI46.0%42.5–49.6%7722
12ElevenLabs45.4%41.9–48.9%7912
13Async43.5%41.0–45.9%15704
14Rime42.5%39.1–45.9%7942
15Fish Audio37.3%34.6–40.1%12133
16Inworld36.1%33.7–38.4%16164
#CompanyWin rate95% rangeComparisonsVoices
1Google Gemini72.7%69.9–75.4%10224
2xAI65.3%61.0–69.5%4852
3ElevenLabs64.4%60.3–68.5%5202
4Resemble AI61.2%56.9–65.6%4852
5MiniMax57.5%53.9–61.0%7443
6Async40.2%35.9–44.4%5092
7Fish Audio39.4%35.9–42.9%7453
8Speechify37.3%34.3–40.4%9724
9Cartesia35.5%32.6–38.5%9994
10OpenAI30.3%26.2–34.4%4772

Voice Versus Model

People compare companies and models, but what you ship is one voice. Within a single model the spread between voices is still large enough to dominate the choice of model.

ModelBest voiceWorst voiceGap
Fish Audio s2.1-pro-free (English)#34 c2623f0c — 949#44 c5f56a6c — 616333 points
Fish Audio s2.1-pro-free (Turkish)#12 67890dc7 — 1065#28 fe96529b — 721344 points
StepFun step-tts-2 (English)#16 lively-girl — 1043#39 magnetic-voiced-male — 891151 points
ElevenLabs eleven_v3 (English)#15 iP95p4xo — 1059#37 Xb7hH8MS — 916143 points

Fish Audio's English pairing is the extreme case: the same model, the same language, 333 points and ten ranks apart, with one of the two finishing last of all 44 voices. Test the voice you are going to ship, not the company that makes it.

Do Listeners Pay Attention

A listening test is only as good as the people doing the listening. Unlike the pilot, this round had the hidden traps running the whole time, so the answer is measured rather than assumed.

Figure 6
Hidden control outcomes
Each control plays one clip against itself. Defensible answers are “about the same” or “audio problem”. A confident pick is a miss. Measured across 242 contributors.
1,448 spotted that the clips matched — 90.7% 149 picked a winner anyway — 9.3%
1,597 hidden controls in total, none of which counted toward any score.

So roughly one time in ten, someone hears the exact same clip twice and still picks a favourite. That is the cost of asking humans to compare things, and it is why we put the traps in. Any listening test without checks like these is carrying about the same noise and does not know it.

Who We Excluded, And Why

Being transparent about this matters more than the result looking clean, so here is exactly what happened before we froze the numbers.

Two contributors tripped the automatic disqualification rule, which needs at least 6 controls seen, at least 3 misses, and a miss rate of 50% or higher. Their votes are out.

A further 51 sat on a softer hold, triggered by two misses. We looked at each one against the 9.3% population miss rate:

GroupPeopleMiss rateChance of that record if listening honestlyDecision
Missed 2 of 2 controls23100%0.9%excluded
Missed 2 of 3767%2.4%excluded
Missed 2 of 4150%4.6%excluded
Missed 2 of 5 to 2 of 122017–40%7–31%readmitted

The 20 whose records are statistically consistent with honest listening were readmitted, returning 1,103 votes to the ranking. The 31 whose records are not — 23 of whom missed every single control they were shown — stay excluded, holding back 555 votes.

That line is a judgement call and we are showing our work rather than hiding it. Note that the automatic rule alone would have cleared all 31, purely because they had not yet been shown 6 controls each. A rule that needs six trials cannot catch someone who fails their first two.

What This Cannot Tell You

We would rather list these than have a reader find them.

Top of the table
The English lead is still a group
The leader's range overlaps 23 other voices. We can tell you which 20 are clearly behind the pack; we cannot order anyone inside it, including first place.
Turkish scale
Turkish is a third the size
851 votes from 44 listeners against 2,185 from 103. Turkish error bars are wider and the field is smaller.
Who voted
The panel is self-selected
224 paid contributors who declared the language as native. Not a representative cross-section of either language.
Baseline
50% means average for this lineup
The English and Turkish lineups differ, so only a company's position transfers between them, never its percentage.
Scope
Short clips only
Clips average about five seconds. Nothing here covers long passages, response speed, cost, or how a voice holds up over ten minutes.
Timing
A snapshot
New models ship constantly. These results are frozen at 25 July 2026 and will not be edited; corrections get published alongside.
The listen check
Playing audio is not listening
The 60% rule tracks how far the audio played, so skipping ahead passes it. The traps, not the rule, catch inattentive listeners.
Coverage
Voices are still missing
We test what we can access when we build a round. A company being absent says nothing about its quality.

Check It Yourself

The results file is public and signed. It is the same file that draws the leaderboard, so there is no private version.

WhatWhere
Browse the leaderboardThe AI Voice Arena — filter by language and question, group by voice or company
Full results filejoinvoicedata.com/api/public/tts-benchmark?download=1
Paid listening workjoinvoicedata.com/tasks

The file carries its own checksum and signature, when it was generated, and a plain description of the method. Each leaderboard reports how many votes it rests on and how many people were behind them.

Working with us

We run these blind English and Turkish tests on unreleased models and tell you where the pronunciation and rhythm problems are before your users find them. If you build AI voices and want to be in the next round, or want a private evaluation, get in touch through patientdesk.ai. If you think we got something wrong, tell us — we publish corrections next to the original rather than quietly editing it.