highest MOS-scoring text-to-speech APIs ranked by voice quality
Blind arena testing, not vendor scorecards, reveals which text-to-speech models actually sound best.

Mean Opinion Score rates synthetic speech on a 1-to-5 scale, and it's been the industry's default measuring stick for decades. Human speech usually lands between 4.5 and 4.8; anything close to that range counts as excellent for a machine voice. But MOS was built for a different job than the one it's doing now, and the cracks show the moment you try to compare vendors instead of just watching one model improve over time.
MOS started in telephony, where engineers needed one number to describe how bad a phone line sounded after compression. It found a second career in text-to-speech because it asks listeners to judge three things that matter for a synthetic voice: naturalness, prosody, and intelligibility. Naturalness is the gut check (does this sound like a person or a machine). Prosody is rhythm, stress, and intonation, the stuff that makes "really?" sound like a question instead of a threat. Intelligibility just means you can understand the words without replaying the clip. Add those three up and you get a decent stand-in for "does this sound good," which is otherwise a hard thing to pin down.
Here's what most vendor decks skip over: MOS is subjective by design, and a 2026 CodeSOTA analysis found that differences under 0.1 points sit inside normal evaluator variance. So a vendor claiming a 4.3 MOS beats a competitor's 4.25 isn't reporting a real gap. It's dressing up noise as a leaderboard. Once you see that, you can't unsee it in every TTS marketing page you read after.
Where MOS falls short and how Elo-based arena testing fills the gap
Every company grades its own homework. No shared test set, no common pool of listeners, no agreed reference tracks. A company runs its own panel on its own text and publishes the number that flatters its own model. None of that is fraud. None of it is comparable across providers, either. Line up five vendor MOS scores and you're not looking at a ranking, you're looking at five different experiments that happen to spit out numbers on the same 1-to-5 scale.
Elo-based blind preference testing fixes the comparability problem, even if it can't touch the subjectivity problem. Listeners hear two anonymous clips generated from the same text, pick the one that sounds more natural, and those wins and losses stack into a rating (the same math chess uses to rank players). Two benchmarks have become the ones people actually trust: the Artificial Analysis Speech Arena Leaderboard and the community-run TTS Arena on Hugging Face. Both run blind A/B voting. Neither lets a vendor submit its own scorecard.
Artificial Analysis is upfront that its rankings aren't paid for, which shouldn't need saying but does. A March 2026 evaluation from Trelis Research, nicknamed "Tricky TTS," adds a third useful data point: it pairs Character Error Rate with UTMOS, a neural estimate of MOS, tested across 10 models on deliberately hard text. That combination works as a check on both the vendor MOS claims and the arena Elo numbers. The rule of thumb: treat Elo as the closest thing to real evidence on offer, and treat vendor MOS as a rough direction, nothing more precise than that.
What the Trelis and Artificial Analysis evaluations reveal about proprietary vs. open-source models
The Trelis "Tricky TTS" results draw a sharper line between proprietary and open-source models than most people expect. Human voice, run through the same pipeline, scored around 4.2 MOS. Proprietary models clustered tight around that mark: error rates between 11% and 19%, MOS above 4.0 across the board, no surprises.
Open-source told a messier story. Character Error Rate ranged from 17% to 86%. MOS spanned 3.3 to 4.5. That's not a tier, that's a canyon, and picking the wrong open-weight model isn't a modest downgrade, it's the gap between usable output and something that mangles words often enough to matter.
Artificial Analysis backs this up from a different angle. In the mid-2026 snapshot, every model in the top 10 by Elo needed an API connection to run, and not one open-weight model cracked that list. Commercial APIs win on peak quality right now; open-source wins on cost and control. Open-source models tend to close gaps faster than incumbents with a lead to protect, so calling this settled forever would be premature. But the data sitting in front of us today doesn't support betting on convergence yet, either.
How the top commercial APIs rank on the Artificial Analysis Speech Arena leaderboard

The August 2026 snapshot puts Cartesia Sonic 3.6 in first with an Elo of 1,283. Alibaba's Qwen-Audio-3.0-TTS-Plus and Simba 3.2 tie for second at 1,238. Luna TTS sits fourth at 1,223. ElevenLabs v3 Conversational rounds out the top five at 1,219.
Rewind to May 2026 and the picture looks different. Google Gemini 3.1 Flash TTS trailed the leader by fewer than 4 Elo points, and ElevenLabs v3 held third place while carrying the biggest sample size in the top 10: 3,753 arena appearances. That volume isn't trivia. More head-to-head matchups means a more stable rating, not a lucky streak against three weak opponents.
The churn between snapshots is the actual story. Alibaba's Fun-Realtime-TTS held the #1 spot at 1,219 Elo across 962 appearances before Sonic 3.6 knocked it down the board. Gemini 3.1 Flash TTS slid from second to third within weeks. No model has held the top spot for more than a few months in any 2026 snapshot, which tells you the ceiling on synthetic speech is still moving, fast enough that the gap between top-five models is often smaller than the measurement error itself.
Cartesia Sonic 3.6: what puts it at the top of the August 2026 leaderboard
Sonic 3.6's 1,283 Elo is the top score on the board right now, and the architecture explains a good chunk of it. Cartesia built the Sonic line on state-space models instead of transformers, an approach the founders pulled from academic research aimed squarely at real-time, low-latency generation. Transformers are good at a lot of things; SSMs were built for exactly this kind of streaming job from the start.
That shows up in the numbers. Sonic 4, the prior generation, hit around 40 milliseconds time-to-first-audio in independent tests run in May 2026, the fastest of any commercial TTS provider tested at the time. Sonic 3.6 carries more than 500 voices across dozens of languages, and Cartesia has flagged specific gains in prosody and emotional range between 3.5 and 3.6.
Pricing for Sonic 3.5 sat in the higher range per million characters; treat 3.6 as roughly the same until updated numbers show up. The open question with Sonic isn't breadth, it's depth. Covering dozens of languages is one thing. Delivering native-quality output in every single one is a separate, harder problem, and it's usually where research-first shops start pulling away from providers that bolted languages on after the fact.
Google Gemini 3.1 Flash TTS: multimodal architecture applied to speech
Google shipped Gemini 3.1 Flash TTS on April 15, 2026, through the Gemini API, Google AI Studio, Vertex AI, and Google Vids. It isn't a standalone speech system bolted onto Google's servers; it's built on the wider Gemini family, so it inherits more than 200 audio tags for steering style, tone, pacing, accent, and scene direction. That's a lot of knobs for a TTS model to hand you.
It supports more than 70 languages natively and generates multi-speaker dialogue without needing a separate pipeline stitched on after the fact. On the leaderboard, it hit 1,214 Elo in May and held nearly steady at 1,213 in July (a quieter kind of success than topping the chart, holding your ground while new models keep showing up at the door).
Pricing lands at $18.3 per million characters, one of the cheaper options in the top tier. The 200-plus audio tags are a real edge for scripted work: voiceover, ad reads, anything where a director wants exact control over delivery. Whether that same expressiveness survives unscripted, back-and-forth conversation is a different question, and the arena's clip-based format doesn't fully answer it.
Inworld Realtime TTS-2 and what dedicated low-latency architecture looks like in the top tier
Inworld's Realtime TTS-2 Research Preview launched in May 2026, built for streaming and live conversation rather than pre-rendered audio. It claims 100 millisecond time-to-first-byte latency, supports more than 200 languages, and keeps voice identity steady when a conversation switches languages mid-stream.
The steering system is unusually fine-grained for a real-time model: eight separate dimensions covering emotion, articulation, intonation, volume, pitch, range, speed, and vocal style, all controlled through plain-language instructions instead of technical parameters. In the May 2026 Artificial Analysis snapshot, the predecessor, Inworld Realtime TTS 1.5 Max, held the top spot outright with a 73.3% win rate across 1,851 appearances. Inworld says it held three of the top five arena positions at once, which is an unusual amount of real estate for one provider to occupy.
Production deployments back the benchmark up. Bible Chat, Talkpal AI, and Astrobeam all run Realtime TTS at a scale of millions of users, which moves the claim out of the lab and into daily use. Worth flagging: a model tuned for live dialogue may not behave the same on a static clip-listening test as it does inside a real back-and-forth. The arena method captures listener preference well. It doesn't capture what happens when latency, interruptions, and turn-taking enter the room.
ElevenLabs v3: foundational research model with the largest arena sample in the top tier
ElevenLabs v3 Conversational held 1,219 Elo in the August 2026 snapshot, fifth on the leaderboard. Back in May it sat third with 3,753 arena appearances, the largest sample of any model in the top 10 at that point. More appearances means a rating less exposed to a run of lucky or unlucky matchups, so v3's spot on the chart rests on a wider base of evidence than most of its neighbors.
The model grew out of original foundational research rather than a fine-tune bolted onto someone else's commodity architecture, a distinction that matters when you're trying to tell whether a provider's gains are systematic or patched in one at a time. The platform runs sub-100 millisecond latency on its conversational models while keeping separate, higher-fidelity models for content production, serving both jobs without smashing them into one compromise product. It also delivers native-quality output across more than 70 languages, worth knowing for any buyer comparing providers whose language lists look wide on paper and thin out fast once you actually sit and listen in each one.
The same models behind those arena numbers power audiobook narration, multilingual dubbing, real-time voice agents, and a developer-facing API, so the leaderboard score and the production stack aren't two separate stories being told about two separate products. That's one solid answer to the build-versus-buy question. Not the only one, but a well-tested one.
Alibaba Qwen-Audio and what the global competitive field means for enterprise buyers
Alibaba's Qwen-Audio-3.0-TTS-Plus sits tied for second on the August 2026 leaderboard at 1,238 Elo. Its predecessor, Fun-Realtime-TTS, briefly held the top spot at 1,219 Elo across 962 appearances before newer entrants passed it, and Alibaba has kept a consistent top-tier presence across several snapshots since. Not a one-time spike.
The real strength here is depth in Mandarin and regional Asian languages, coverage that's genuinely thinner among Western providers. For any enterprise serving those markets, that's not a marginal edge, it's often the deciding factor. But arena Elo says nothing about data residency rules, terms-of-service differences, or the procurement headache of routing customer voice data through Alibaba-backed infrastructure instead of a US-based provider, and those questions carry as much weight as the win rate once someone's actually signing the contract.
The competitive field looks genuinely global now, not just California and a couple of research labs. North America still led total TTS revenue with a 40.9% share in 2025, according to Polaris Market Research, but Asia Pacific is growing fastest at a 30.7% compound annual rate. The leaderboard composition tracks that shift almost exactly.
What MOS and Elo rankings don't measure — and what to check before choosing an API
Arena tests run on short, often clean text clips. They don't tell you how a model handles a 20-minute audiobook chapter, dense technical jargon, an odd proper noun, or a sentence that switches languages halfway through. That's a real gap, not a footnote, because most production use cases look nothing like a clean 10-second clip read in a quiet room.
Latency is a separate axis from quality, and mixing the two up is a common mistake. A model that tops the static listening test can still be the wrong pick for a live voice agent that needs a reply inside 100 milliseconds. Voice cloning fidelity, emotional range, and fine prosody control don't collapse cleanly into one Elo number either; a model ranked near the top overall can still sound flat in one genre, one accent, or one emotional register nobody bothered to test.
Language breadth deserves real skepticism. Listing 70 or 200 supported languages on a spec sheet is not the same claim as native-quality output in each one, and the only way to check is a listening test run by native speakers in the languages you actually need, not the ones sitting front and center in a vendor's demo reel. Pricing spreads matter too: Gemini 3.1 Flash TTS at a lower per-million-character price against Cartesia Sonic 3.5 at a higher per-million-character price stops being a rounding error the moment volume climbs. None of this touches uptime guarantees, support tiers, or production stability, and no benchmark captures those because that's not what benchmarks are for. Run your own A/B test on your own text and target language. Measure latency under your real expected load. Get native speakers to listen to the multilingual output. Read the data handling terms before you sign anything.
How the TTS quality race is likely to develop through the rest of 2026
The money explains why the leaderboard keeps reshuffling every few months instead of settling. Global Market Insights put the TTS market at roughly $5.7 billion in 2026, projected to hit $35.3 billion by 2035, a 22.4% compound annual growth rate. That kind of money pouring in means faster releases and more leaderboard turnover, not less.
Neural TTS already holds more than 83% of market share, so the growth left isn't a shift away from older synthesis methods, it's a straight quality fight among neural approaches competing for the same customers. And the 2026 snapshots make the pattern obvious: no model held the top spot for more than a few months all year, so whatever ranking is published today is a photograph, not a verdict carved in stone.
Open-source models still show that wide spread the Trelis evaluation turned up (MOS as low as 3.3 and as high as 4.5 depending on which model you grab off the shelf), but the pace of improvement there is fast enough to keep an eye on, especially on the Hugging Face TTS Arena where new open releases show up quickly. The next real frontier probably isn't squeezing another tenth of a point out of English MOS. It's multilingual parity, keeping emotional consistency intact when a model crosses from English into Mandarin into Spanish without losing its footing, and that's exactly where providers doing multilingual research from day one have a structural edge over the ones bolting languages on after launch. For anyone actually picking an API: check the Artificial Analysis Speech Arena every quarter, and make the call off a short-listed trial run on current data, not a benchmark that was true six months ago and quietly isn't anymore.


