People listened to pairs of AI voices reading the same sentence and picked the one they preferred, never seeing who made either. This is round two: 12,245 comparisons, 4.6 times the size of our pilot.
What changed since the pilot, how the voting works, what round two found, and what we still cannot tell you.
Summary
Most AI voice companies show you a demo page. Very few tell you how they tested, how confident they are, or how the voice holds up in a language other than English. We needed to pick a voice for a phone agent that talks to real patients, could not find that information anywhere, so we measured it ourselves.
Round 2026.2 covers 44 English voices and 28 Turkish voices from 16 companies. Two voices read the same sentence, a listener picks one, and nobody sees a brand name until voting closes. Every score ships with an error bar, and the full results file is a public download.
What Changed Since The Pilot
Our July pilot was honest about being too small: 2,675 votes across 40 English voices meant 21 of the 22 ranked entries overlapped the leader, and we said plainly that we could not rank English. Round two is 4.6 times bigger, and the picture is different in three specific ways.
| Pilot 2026.1 | Round 2026.2 | |
|---|---|---|
| Votes counted | 2,675 | 12,245 |
| Listeners | 44 | 224 |
| Voices | 66 from 13 companies | 72 from 16 companies |
| Voices with enough votes to rank | 48 of 66 | all 72 |
| Typical English error bar | 340 points | 140 points |
| English voices clearly behind the leader | 1 of 21 | 20 of 43 |
First, English separates now — though not at the top. The error bars roughly halved, and 20 of the 44 voices are now measurably behind the leading pack instead of statistically tangled with it. The other 24 still overlap each other.
Second, the cross-language finding survived and got stronger. It was the one pilot result we said we believed, and more data has not softened it.
Third, and least comfortably, one pilot finding did not survive. We reported that the four listening questions measured genuinely different things, pointing at a near-zero correlation between “sounds human” and overall preference in Turkish. With four times the data that correlation is 0.89. It was noise. We were measuring our own sample size.
| How closely each question tracks overall preference | English pilot | English now | Turkish pilot | Turkish now |
|---|---|---|---|---|
| Sounds human | 0.46 | 0.61 | 0.08 | 0.89 |
| Clear and correct | 0.52 | 0.76 | 0.68 | 0.95 |
| Rhythm and expression | 0.41 | 0.76 | 0.54 | 0.87 |
Splitting the questions is still worth doing — it is how you find out they agree, and it keeps the option open for a future round where they might not. But we no longer claim they measure different things, and anyone who quoted our pilot on that should stop.
How It Works
The point of the setup is to test voices, not brand loyalty or listener patience. Nothing about the method changed between rounds, which is what makes the two comparable.
Both voices read the same sentence
Every comparison uses identical text, from a fixed bank of 24 sentences — 12 English, 12 Turkish — covering plain statements, questions, lists, names, numbers and dates, abbreviations, long winding sentences, and emotional lines.
Nobody sees the brand names
The listener's browser gets an audio file and its length, and nothing else. File names are random IDs, so no company or model name appears in the page, the network traffic, or anywhere a curious listener could dig. Buttons say “First voice” and “Second voice”. The map from clip to company lives in a separate file that is only opened after voting closes.
One question at a time, and you have to listen
A listener gets 25 comparisons in a sitting, all asking the same one of four questions. The server rejects any vote where the listener did not play at least 60% of both clips, so a modified browser cannot fake it. The two players cannot run at once. Replays are unlimited. There is no skip — you pick a voice, say they sound the same, or report an audio problem. Nobody sees the same pair twice, and nobody votes more than once on a given question in a given language.
Hidden checks
Three of the 25 comparisons are traps: the same clip played on both sides, at positions 5, 13 and 21. They look completely normal, both sides getting their own download link for the identical file. The honest answers are “about the same” or “audio problem”; picking a winner counts as a miss. Traps are paid like any other comparison, which is what keeps them hidden, and they never count toward anyone's score.
How the scores work
We use a standard head-to-head method called Bradley–Terry, which turns a pile of pairwise results into a strength number per voice. A tie counts as half a win. Scores are centred so the average voice in each language sits at 1000. Error bars come from resampling listeners rather than individual clicks, because one person voting 25 times is not 25 independent opinions. A voice needs 20 votes to earn a rank; in this round every voice cleared that easily.
What We Collected
Voting ran from 20 to 25 July 2026 and closed at 556 completed sittings.
English Results
This is what the pilot could not give us, with one caveat. The field is 44 voices, the bottom 20 are now measurably behind, and the top 24 are still a crowd.
Google Gemini's gemini-3.1-flash-tts-preview leads at 1138, but its range overlaps every voice down to rank 24. The honest reading is that the English field splits in two: a broad leading pack of 24 voices we cannot order internally, and 20 that are clearly behind it. That is a real gain over the pilot, where only one voice was separable from the leader — but it is separation of the bottom half, not a winner at the top.
| Rank | Company | Model | Voice | Score | 95% range | Votes |
|---|---|---|---|---|---|---|
| 1 | Google Gemini | gemini-3.1-flash-tts-preview | Puck | 1137.8 | 1075–1205 | 93 |
| 2 | Smallest | lightning_v3.1_pro | nolan | 1125.4 | 1055–1209 | 86 |
| 3 | Cartesia | sonic-3.5-2026-05-04 | db6b0ed5 | 1123.7 | 1069–1184 | 100 |
| 4 | Speechify | simba-3.2 | geffen_32 | 1121.2 | 1051–1191 | 107 |
| 5 | xAI | grok-voice-tts | altair | 1107.3 | 1036–1178 | 98 |
| 6 | Speechify | simba-3.2 | dominic_32 | 1099.8 | 1036–1161 | 98 |
| 7 | Resemble AI | chatterbox | 6e37aa15 | 1095.1 | 1028–1175 | 97 |
| 8 | Google Gemini | gemini-3.1-flash-tts-preview | Kore | 1092.4 | 1033–1161 | 116 |
| 9 | Cartesia | sonic-3.5-2026-05-04 | 630ed21c | 1083.2 | 1015–1148 | 89 |
| 10 | Gradium | default | YTpq7expH9539ERJ | 1077.0 | 999–1153 | 112 |
| 11 | MiniMax | speech-2.8-hd | English_Trustworth_Man | 1062.7 | 1005–1127 | 88 |
| 12 | Smallest | lightning_v3.1_pro | chelsea | 1061.4 | 1007–1117 | 109 |
| 13 | Cartesia | sonic-3 | db6b0ed5 | 1060.5 | 997–1138 | 111 |
| 14 | Resemble AI | chatterbox | 819fcc57 | 1060.4 | 1003–1121 | 108 |
| 15 | ElevenLabs | eleven_v3 | iP95p4xoKVk53GoZ742B | 1058.7 | 988–1140 | 91 |
| 16 | StepFun | step-tts-2 | lively-girl | 1042.6 | 971–1108 | 98 |
| 17 | Hume AI | 1 | Female Meditation Guide | 1037.5 | 981–1102 | 104 |
| 18 | Gradium | default | LFZvm12tW_z0xfGo | 1029.9 | 966–1088 | 95 |
| 19 | xAI | grok-voice-tts | ara | 1028.8 | 974–1092 | 108 |
| 20 | Async | async_flash_v1.0 | cca0e076 | 1023.3 | 958–1091 | 106 |
| 21 | Google Gemini | gemini-2.5-pro-preview-tts | Kore | 1022.4 | 973–1079 | 111 |
| 22 | MiniMax | speech-2.8-hd | English_CalmWoman | 1021.2 | 951–1091 | 110 |
| 23 | Async | async_flash_v1.5 | e5a67eaf | 1018.4 | 942–1092 | 83 |
| 24 | Async | async_flash_v1.0 | 317bf805 | 1007.3 | 931–1076 | 90 |
| 25 | OpenAI | gpt-4o-mini-tts | nova | 1005.2 | 936–1072 | 101 |
| 26 | MiniMax | speech-2.8-turbo | English_CalmWoman | 1001.3 | 932–1067 | 114 |
| 27 | Hume AI | 2 | Female Meditation Guide | 992.7 | 931–1061 | 106 |
| 28 | Cartesia | sonic-3 | 630ed21c | 974.5 | 891–1048 | 93 |
| 29 | Rime | coda | masonry | 969.1 | 894–1036 | 77 |
| 30 | Fish Audio | s2-pro | e3cd3841 | 956.8 | 885–1031 | 108 |
| 31 | Async | async_flash_v1.5 | 0aef6559 | 956.2 | 885–1017 | 114 |
| 32 | Smallest | lightning_v3.1 | liam | 955.9 | 882–1031 | 97 |
| 33 | Smallest | lightning_v3.1 | avery | 954.7 | 884–1021 | 97 |
| 34 | Fish Audio | s2.1-pro-free | c2623f0c | 949.3 | 884–1015 | 104 |
| 35 | OpenAI | gpt-4o-mini-tts | onyx | 940.4 | 860–1024 | 84 |
| 36 | Rime | coda | astra | 936.2 | 862–997 | 107 |
| 37 | ElevenLabs | eleven_v3 | Xb7hH8MSUJpSbSDYk0k2 | 915.7 | 842–988 | 101 |
| 38 | Google Gemini | gemini-2.5-pro-preview-tts | Puck | 899.6 | 825–977 | 92 |
| 39 | StepFun | step-tts-2 | magnetic-voiced-male | 891.2 | 814–966 | 92 |
| 40 | Inworld | inworld-tts-2 | Bianca | 883.8 | 815–938 | 109 |
| 41 | Inworld | inworld-tts-2 | Callum | 874.5 | 799–946 | 87 |
| 42 | Inworld | inworld-tts-1.5-max | Bianca | 865.5 | 793–936 | 109 |
| 43 | Inworld | inworld-tts-1.5-max | Callum | 863.1 | 775–943 | 84 |
| 44 | Fish Audio | s2.1-pro-free | c5f56a6c | 616.2 | 399–726 | 86 |
Turkish Results
Turkish has a smaller field and a much clearer top. 19 of 27 challengers are clearly behind the leader.
gemini-3.1-flash-tts-preview Puck leads at 1302 with a range of 1219–1417, clear of every rival.| Rank | Company | Model | Voice | Score | 95% range | Votes |
|---|---|---|---|---|---|---|
| 1 | Google Gemini | gemini-3.1-flash-tts-preview | Puck | 1301.5 | 1219–1417 | 58 |
| 2 | Google Gemini | gemini-3.1-flash-tts-preview | Kore | 1194.9 | 1094–1302 | 66 |
| 3 | Google Gemini | gemini-2.5-pro-preview-tts | Kore | 1184.7 | 1105–1286 | 64 |
| 4 | Google Gemini | gemini-2.5-pro-preview-tts | Puck | 1183.4 | 1103–1293 | 60 |
| 5 | xAI | grok-voice-tts | ara | 1157.4 | 1083–1248 | 61 |
| 6 | MiniMax | speech-2.8-hd | Turkish_CalmWoman | 1150.6 | 1072–1252 | 66 |
| 7 | ElevenLabs | eleven_v3 | iP95p4xoKVk53GoZ742B | 1146.6 | 1068–1232 | 65 |
| 8 | ElevenLabs | eleven_v3 | Xb7hH8MSUJpSbSDYk0k2 | 1132.5 | 1044–1245 | 64 |
| 9 | Resemble AI | chatterbox-multilingual | 64ad5770 | 1128.4 | 1048–1222 | 60 |
| 10 | MiniMax | speech-2.8-turbo | Turkish_CalmWoman | 1120.7 | 1047–1216 | 62 |
| 11 | Resemble AI | chatterbox-multilingual | 45b91687 | 1093.4 | 1021–1174 | 57 |
| 12 | Fish Audio | s2.1-pro-free | 67890dc7 | 1065.2 | 984–1148 | 55 |
| 13 | xAI | grok-voice-tts | altair | 1061.5 | 964–1142 | 55 |
| 14 | MiniMax | speech-2.8-hd | Turkish_Trustworthyman | 1048.1 | 980–1128 | 56 |
| 15 | Fish Audio | s2-pro | 818b100d | 1017.0 | 936–1119 | 60 |
| 16 | Async | async_flash_v1.0 | 317bf805 | 967.1 | 884–1051 | 61 |
| 17 | Speechify | simba-multilingual | ayse | 920.7 | 832–1008 | 56 |
| 18 | Cartesia | sonic-3.5-2026-05-04 | 630ed21c | 906.1 | 810–979 | 66 |
| 19 | Cartesia | sonic-3.5-2026-05-04 | db6b0ed5 | 900.0 | 763–992 | 53 |
| 20 | Speechify | simba-multilingual | arda | 893.2 | 795–988 | 60 |
| 21 | Cartesia | sonic-3 | db6b0ed5 | 876.7 | 779–968 | 64 |
| 22 | Speechify | simba-multilingual | baris | 876.3 | 759–967 | 55 |
| 23 | OpenAI | gpt-4o-mini-tts | onyx | 826.5 | 719–928 | 63 |
| 24 | OpenAI | gpt-4o-mini-tts | nova | 825.4 | 729–902 | 59 |
| 25 | Async | async_flash_v1.0 | cca0e076 | 808.0 | 694–902 | 60 |
| 26 | Speechify | simba-multilingual | aylin | 763.0 | 623–866 | 66 |
| 27 | Cartesia | sonic-3 | 630ed21c | 730.1 | 619–815 | 61 |
| 28 | Fish Audio | s2.1-pro-free | fe96529b | 721.0 | 578–817 | 69 |
The Language Gap
Sixteen of the voices were put in front of listeners in both languages — identical company, model and voice ID. That is the cleanest possible test of whether quality travels, because nothing changes except the language being read.
It does not travel.
| Voice (identical in both languages) | English | Turkish | Direction |
|---|---|---|---|
Cartesia db6b0ed5 | #3 of 44 | #19 of 28 | falls in Turkish |
Cartesia db6b0ed5 | #13 of 44 | #21 of 28 | falls in Turkish |
Async cca0e076 | #20 of 44 | #25 of 28 | falls in Turkish |
Cartesia 630ed21c | #9 of 44 | #18 of 28 | falls in Turkish |
xAI altair | #5 of 44 | #13 of 28 | falls in Turkish |
Cartesia 630ed21c | #28 of 44 | #27 of 28 | falls in Turkish |
OpenAI nova | #25 of 44 | #24 of 28 | falls in Turkish |
Async 317bf805 | #24 of 44 | #16 of 28 | falls in Turkish |
OpenAI onyx | #35 of 44 | #23 of 28 | falls in Turkish |
Google Gemini Puck | #1 of 44 | #1 of 28 | falls in Turkish |
ElevenLabs iP95p4xoKVk53GoZ742B | #15 of 44 | #7 of 28 | rises in Turkish |
Google Gemini Kore | #8 of 44 | #2 of 28 | rises in Turkish |
xAI ara | #19 of 44 | #5 of 28 | rises in Turkish |
Google Gemini Kore | #21 of 44 | #3 of 28 | rises in Turkish |
ElevenLabs Xb7hH8MSUJpSbSDYk0k2 | #37 of 44 | #8 of 28 | rises in Turkish |
Google Gemini Puck | #38 of 44 | #4 of 28 | rises in Turkish |
One important exception, stated plainly: Google Gemini's gemini-3.1-flash-tts-preview Puck is ranked 1st in both languages. A voice can be good in two languages at once. The point is that most are not, and you cannot tell which from an English score.
Pooling all four questions per company gives each one enough comparisons for the cross-language picture to resolve sharply. This is the most useful thing in the report.
The two highlighted companies move in opposite directions, and neither move is small.
| Company | English | Turkish | What happened |
|---|---|---|---|
| Speechify | 63.6% — 1st of 16 | 37.3% — 8th of 10 | top to near-bottom |
| ElevenLabs | 45.4% — 12th of 16 | 64.4% — 3rd of 10 | near-bottom to top |
| Cartesia | 57.7% — 3rd of 16 | 35.5% — 9th of 10 | top to bottom |
| Google Gemini | 57.2% — 5th of 16 | 72.7% — 1st of 10 | strong to dominant |
| OpenAI | 46.0% — 11th of 16 | 30.3% — 10th of 10 | weak in both |
| xAI | 60.1% — 2nd of 16 | 65.3% — 2nd of 10 | the only steady one |
Cartesia repeats its pilot result almost exactly, which is the best evidence we have that the effect is real rather than an artefact of one round. Speechify is the new and starker case: first in English, eighth of ten in Turkish.
Six of the sixteen companies fielded nothing in Turkish at all — Smallest, Gradium, StepFun, Hume AI, Rime and Inworld. Turkish speakers choose from 28 voices against 44, and four of the top ten Turkish slots belong to Google.
| # | Company | Win rate | 95% range | Comparisons | Voices |
|---|---|---|---|---|---|
| 1 | Speechify | 63.6% | 60.2–66.9% | 799 | 2 |
| 2 | xAI | 60.1% | 56.8–63.5% | 821 | 2 |
| 3 | Cartesia | 57.7% | 55.3–60.2% | 1557 | 4 |
| 4 | Smallest | 57.4% | 55.0–59.8% | 1576 | 4 |
| 5 | Google Gemini | 57.2% | 54.8–59.6% | 1590 | 4 |
| 6 | Resemble AI | 54.3% | 50.8–57.8% | 779 | 2 |
| 7 | MiniMax | 52.6% | 49.9–55.4% | 1229 | 3 |
| 8 | Gradium | 50.6% | 47.2–54.0% | 818 | 2 |
| 9 | StepFun | 49.9% | 46.3–53.5% | 744 | 2 |
| 10 | Hume AI | 49.3% | 46.0–52.6% | 863 | 2 |
| 11 | OpenAI | 46.0% | 42.5–49.6% | 772 | 2 |
| 12 | ElevenLabs | 45.4% | 41.9–48.9% | 791 | 2 |
| 13 | Async | 43.5% | 41.0–45.9% | 1570 | 4 |
| 14 | Rime | 42.5% | 39.1–45.9% | 794 | 2 |
| 15 | Fish Audio | 37.3% | 34.6–40.1% | 1213 | 3 |
| 16 | Inworld | 36.1% | 33.7–38.4% | 1616 | 4 |
| # | Company | Win rate | 95% range | Comparisons | Voices |
|---|---|---|---|---|---|
| 1 | Google Gemini | 72.7% | 69.9–75.4% | 1022 | 4 |
| 2 | xAI | 65.3% | 61.0–69.5% | 485 | 2 |
| 3 | ElevenLabs | 64.4% | 60.3–68.5% | 520 | 2 |
| 4 | Resemble AI | 61.2% | 56.9–65.6% | 485 | 2 |
| 5 | MiniMax | 57.5% | 53.9–61.0% | 744 | 3 |
| 6 | Async | 40.2% | 35.9–44.4% | 509 | 2 |
| 7 | Fish Audio | 39.4% | 35.9–42.9% | 745 | 3 |
| 8 | Speechify | 37.3% | 34.3–40.4% | 972 | 4 |
| 9 | Cartesia | 35.5% | 32.6–38.5% | 999 | 4 |
| 10 | OpenAI | 30.3% | 26.2–34.4% | 477 | 2 |
Voice Versus Model
People compare companies and models, but what you ship is one voice. Within a single model the spread between voices is still large enough to dominate the choice of model.
| Model | Best voice | Worst voice | Gap |
|---|---|---|---|
Fish Audio s2.1-pro-free (English) | #34 c2623f0c — 949 | #44 c5f56a6c — 616 | 333 points |
Fish Audio s2.1-pro-free (Turkish) | #12 67890dc7 — 1065 | #28 fe96529b — 721 | 344 points |
StepFun step-tts-2 (English) | #16 lively-girl — 1043 | #39 magnetic-voiced-male — 891 | 151 points |
ElevenLabs eleven_v3 (English) | #15 iP95p4xo — 1059 | #37 Xb7hH8MS — 916 | 143 points |
Fish Audio's English pairing is the extreme case: the same model, the same language, 333 points and ten ranks apart, with one of the two finishing last of all 44 voices. Test the voice you are going to ship, not the company that makes it.
Do Listeners Pay Attention
A listening test is only as good as the people doing the listening. Unlike the pilot, this round had the hidden traps running the whole time, so the answer is measured rather than assumed.
So roughly one time in ten, someone hears the exact same clip twice and still picks a favourite. That is the cost of asking humans to compare things, and it is why we put the traps in. Any listening test without checks like these is carrying about the same noise and does not know it.
Who We Excluded, And Why
Being transparent about this matters more than the result looking clean, so here is exactly what happened before we froze the numbers.
Two contributors tripped the automatic disqualification rule, which needs at least 6 controls seen, at least 3 misses, and a miss rate of 50% or higher. Their votes are out.
A further 51 sat on a softer hold, triggered by two misses. We looked at each one against the 9.3% population miss rate:
| Group | People | Miss rate | Chance of that record if listening honestly | Decision |
|---|---|---|---|---|
| Missed 2 of 2 controls | 23 | 100% | 0.9% | excluded |
| Missed 2 of 3 | 7 | 67% | 2.4% | excluded |
| Missed 2 of 4 | 1 | 50% | 4.6% | excluded |
| Missed 2 of 5 to 2 of 12 | 20 | 17–40% | 7–31% | readmitted |
The 20 whose records are statistically consistent with honest listening were readmitted, returning 1,103 votes to the ranking. The 31 whose records are not — 23 of whom missed every single control they were shown — stay excluded, holding back 555 votes.
That line is a judgement call and we are showing our work rather than hiding it. Note that the automatic rule alone would have cleared all 31, purely because they had not yet been shown 6 controls each. A rule that needs six trials cannot catch someone who fails their first two.
What This Cannot Tell You
We would rather list these than have a reader find them.
Check It Yourself
The results file is public and signed. It is the same file that draws the leaderboard, so there is no private version.
| What | Where |
|---|---|
| Browse the leaderboard | The AI Voice Arena — filter by language and question, group by voice or company |
| Full results file | joinvoicedata.com/api/public/tts-benchmark?download=1 |
| Paid listening work | joinvoicedata.com/tasks |
The file carries its own checksum and signature, when it was generated, and a plain description of the method. Each leaderboard reports how many votes it rests on and how many people were behind them.
Working with us
We run these blind English and Turkish tests on unreleased models and tell you where the pronunciation and rhythm problems are before your users find them. If you build AI voices and want to be in the next round, or want a private evaluation, get in touch through patientdesk.ai. If you think we got something wrong, tell us — we publish corrections next to the original rather than quietly editing it.