text-to-speech pricing compared across premium and budget providers

Pricing varies wildly by billing unit, voice quality tier, and hidden licensing requirements.

Staff Writer · · 8 min read · Updated
Cover illustration for “text-to-speech pricing compared across premium and budget providers”
Text to Speech · August 26, 2026 · 8 min read · 1,791 words

Text-to-speech pricing looks simple until you put two vendors on the same spreadsheet and watch the numbers refuse to line up. Providers bill in different units for different things, and the lowest price per unit usually belongs to the voice quality nobody actually ships. A 10,000-word narration script might run you four dollars or forty, depending on tier and billing unit. Voice quality, latency, cloning rights, and commercial licensing all sit inside that price somewhere, but none show up in the headline rate. The market is set to grow from $4.55 billion in 2024 to $37.55 billion by 2032, and money moving that fast rewrites its own rules every few months, so whatever comparison you bookmarked last year is stale now.

The variables that actually determine what you pay

Model quality tier is the biggest lever on any pricing page. Cloud vendors price standard voices around a few dollars per million characters, then jump to several times more per million or higher once you touch neural voices, which is the tier everyone actually wants. That's a fourfold markup hiding behind what looks like one product line.

Billing unit is the second trap. Cloud hyperscalers charge by the character, while subscription platforms sell credits spent across a handful of capabilities at different burn rates. Agent-focused vendors bill by runtime minute, because a live phone call and an overnight audiobook render are not the same product, even though both come out as speech in the end. Line up a per-character rate against a credit bundle against a per-minute fee and you're not comparing prices anymore. You're doing currency conversion without knowing the exchange rate.

Latency class trips up more buyers than it should. Real-time agent TTS needs infrastructure built to answer in under 100 milliseconds, and that gets priced on its own line, separate from the batch pipes rendering an audiobook while everyone sleeps. Voice cloning and voice library access carry their own logic too: instant cloning, professional cloning, and libraries with hundreds of prebuilt voices are three separate line items inside the same provider, not one feature. Commercial licensing bites people late; some providers gate commercial use behind a paid tier, so the free plan you tested all week turns out to be legally useless for the project you actually wanted it for. The real floor price is whatever the first commercially licensed tier costs, not the free number on the homepage. Volume commitments shift the math again: pay-as-you-go, monthly minimums, and annual subscriptions produce three different effective rates at the same volume. Neural and custom voices already make up more than 83% of the TTS voice-type market, so skipping neural quality to save a few dollars isn't saving money. It just means shipping a voice most listeners clock as dated by the first sentence.

Budget-tier cloud APIs: Google Cloud TTS and Amazon Polly

Google Cloud TTS and Amazon Polly are built the same way: pay-as-you-go, priced per character, with a free monthly allowance generous enough to make small projects effectively free. That structure matters because it's the default everyone else in this market gets measured against.

Google's standard voices run a few dollars per million characters, WaveNet and Neural2 sit at several times that per million, and Studio and Chirp 3 HD reach a higher premium per million. Polly follows the same shape: Standard at a few dollars per million, Neural at several times that per million. The gap between tiers matters more than the $4 rate everyone quotes in comparison posts, since that rate belongs to an engine most production teams abandoned years ago. Once you need neural quality, you're at $16 on both platforms, and the HD tier runs at a premium rate on Google's side. Polly's Generative engine and Google's Chirp 3 HD show both companies pushing upmarket hard, which means "budget" increasingly describes a tier nobody actually ships with.

These platforms make sense where volume is high, cost sensitivity is real, and nobody's listening closely for warmth: IVR systems, accessibility tooling, internal software where the voice is furniture, not the product.

Mid-range pay-as-you-go: developer-focused APIs

OpenAI's tts-1 model prices at a moderate rate per million characters, with the HD version higher, after a substantial price cut in January 2025 put it up against cheaper rivals. The newer gpt-4o-mini-tts model adds steerable voice instructions: you pass a tone or style directive alongside your text instead of picking from a dropdown of presets. That's a genuinely different kind of control, and it's the more interesting move of the two, honestly.

There's no subscription tier, just pay-as-you-go, which turns cost forecasting into a straight multiplication problem, though the catch is that there's no path to a volume discount short of an enterprise deal negotiated directly with OpenAI.

Some APIs in this tier play a different game, built for teams running real-time voice agents. A Voice Agent API that bundles speech-to-text, text-to-speech, and orchestration into one product means putting its rate next to a narration tool's per-character price is like pricing a taxi against a bus ticket. Anyone who benchmarked costs against an older model version should revisit current pay-as-you-go rates, as pricing in this segment has shifted. These tools build for developers wiring an API into a product, not for someone opening a browser tab to knock out a voiceover before lunch.

Enterprise commitment pricing: Microsoft Azure AI Speech

Azure's standard neural pricing puts it in the same general range as Google's Neural2 and Amazon's Polly Neural. Commitment tiers are where Azure's pricing story gets more interesting, and that's before anyone even looks at volume discounts.

Commitment tiers change the math in a way that trips people up, because these aren't usage discounts, they're fixed monthly minimums. You pay the floor whether you use it or not, and unused capacity doesn't roll into next month. At the highest commitment tier, the effective per-million rate can fall well below standard pay-as-you-go options, which makes Azure potentially the most economical choice on this list, but only if your volume forecast holds up. Guess wrong and you've locked in a floor payment for capacity that just evaporates, which is worse than sitting on pay-as-you-go somewhere else.

What justifies the premium is the part that never shows up in a demo: enterprise SLAs, compliance coverage, data residency controls. In regulated industries those aren't extras, they're the entire reason a team is on Azure instead of a cheaper API. Azure also ships a large library of prebuilt neural voices, which matters for global deployments that need regional accent coverage a smaller voice library can't touch.

Subscription platforms: ElevenLabs and the credit-bundle model

Diagram: ElevenLabs Credit Tiers at a Glance. Visualizes: Show the six ElevenLabs subscription tiers as a ranked progression by monthly cost and credit volume: Free ($0 / 10,000 credits), Starter ($6 / 30,000), Creator ($22 / 121,000), Pro ($99 /…

Some subscription platforms price in credits, not characters, and that's the first habit to break if you're coming from a hyperscaler background. Tiers run Free at $0 for 10,000 credits, Starter at $6 for 30,000, Creator at $22 for 121,000, Pro at $99 for 600,000, Scale at $299 for 1,800,000, and Business at $990 for 6,000,000. Annual billing saves roughly two months' cost across the board, dropping effective monthly rates to $5, $18.33, $82.50, $249.17, and $825.

Credits aren't a clean stand-in for characters, and that's by design. Flash TTS burns credits at $0.05 per 1,000 characters, Multilingual output runs $0.10, dubbing costs $0.33 to $2.20 per minute, and transcription bills at $0.22 per hour. Conversational AI agent pricing got restructured in late 2025 to bundle agent minutes directly into each paid plan (75 minutes on Starter, running up to 12,375 on Business), with overage billed at $0.08 per minute plus separate LLM token costs passed through from whatever model powers the agent.

Commercial licensing starts at the Starter tier; the free plan is non-commercial only, and that's the real floor for anyone building something meant to ship or sell. What the credit system buys beyond raw audio is voice cloning, a sizable voice library, dubbing, multilingual output, and a content interface a bare character-billed API doesn't offer at any price. The tradeoff is operational: tracking which model burns credits at what rate, watching usage across capability types, and budgeting for LLM pass-through on agent minutes is not a one-time calculation. It's a running spreadsheet you keep open.

For a plain benchmark: a 30-minute voiceover produced through a subscription platform at this tier runs roughly a few dollars, against tens to hundreds of dollars or more for a human voice actor doing the same job.

Low-latency agent specialists and where they fit

Several specialist providers build their products around ultra-low latency for real-time voice agents, where response time isn't a feature bullet but the whole point of the product.

Some platforms sit across both categories instead of picking one, with a low-latency model handling the agent use case while standard models cover narration and content work, all inside one platform rather than two separate products. The agent pricing model, bundled minutes plus overage, reflects a cost structure built for orchestration and connection time, not raw text length. That's the same logic behind every bundled agent product in this market: price the session, because the character count was never what made it expensive.

Anyone evaluating TTS for a conversational agent should price the whole stack, speech-to-text, orchestration, text-to-speech, together, instead of isolating the TTS line and comparing it to a narration rate built for a completely different job.

Matching pricing structure to use case: a practical decision framework

Table: TTS Pricing Models by Use Case. Compares Representative Vendors, Billing Unit, Best For, Key Risk, and 1 more by Hyperscaler Pay-as-You-Go, Subscription Credit Bundle, Agent Specialist and Enterprise Commitment.

Cheapest isn't the question; whether the billing model matches your actual usage pattern is the question.

High-volume batch work with predictable character counts belongs on hyperscaler pay-as-you-go pricing (Polly or Google Cloud), or on an Azure commitment tier once volume justifies the fixed floor. Developer prototyping or low-volume API testing without subscription overhead fits a pay-as-you-go API without subscription tiers best. Content work, narration, podcasts, audiobooks, dubbing, especially anything needing voice cloning or multilingual output, points toward a subscription platform with credit bundles that cover a range from solo creator tiers up through production team plans. Real-time conversational agents get judged on session latency and bundle economics, not per-character rates; Cartesia Sonic-3 and similar low-latency specialist tools are the ones that matter here. Enterprise compliance needs (HIPAA, data residency, contractual SLAs) point toward Azure AI Speech, with the caveat that commitment pricing only pays off if the volume forecast holds up.

Commercial licensing deserves one more mention, since it's the easiest line to miss on a pricing page: confirm it's included at the tier you're actually buying, not somewhere higher up the list where you assumed it would be. With a large share of enterprise applications expected to run AI agents by 2026, text-to-speech is turning into an infrastructure line item with real cost forecasting attached, not a one-off purchase of a creative tool. Budget it that way, and run the numbers before you sign anything.

Sources

  1. cloud.google.com
  2. texttolab.com
Filed underText to Speech

More in Text to Speech