Modern text-to-speech (TTS) sounds good in a demo. You paste a paragraph into a vendor's playground, pick a voice, and it reads it back almost like a person would. Then you wire it into a real product and things get harder. Responses take too long to start. The voice reads "Dr." as "drive". The monthly bill grows faster than your user count.
None of this means TTS isn't ready. It means production TTS is an engineering problem, not a voice-picking problem. This post covers the three trade-offs that matter most — latency, cost, and voice quality — and how to make sensible choices on each.
Two very different jobs
Before comparing anything, be clear about which job you're doing. TTS shows up in two shapes, and they have almost opposite requirements.
Batch narration. Audiobooks, article read-alouds, course videos, pre-recorded phone menus. You have the full text up front, nobody is waiting, and quality is everything. You can afford to synthesize slowly, listen to the result, and re-render a paragraph that came out wrong.
Real-time speech. Voice assistants, phone agents, live accessibility features. Someone is waiting for the system to talk. Every moment of silence is felt, and you often don't even have the full text yet — it's still streaming out of an LLM.
Most mistakes come from evaluating for one job and deploying for the other. A voice that sounds wonderful on a pre-rendered paragraph may only be available through a slow, non-streaming endpoint. A fast streaming voice may sound flat over a twenty-minute chapter.
Latency: measure time to first audio
For real-time use, the number that matters is time to first audio: the gap between "we have something to say" and "the user hears sound". Total synthesis time matters much less. Once audio starts playing, the rest only has to arrive faster than it plays.
TTS is usually just one stage in a longer chain. In a typical voice agent, the delay a user feels is the sum of:
- Detecting that the user has finished speaking
- Transcribing their speech
- The LLM producing its first useful tokens
- Collecting enough text to send to TTS
- TTS producing its first audio chunk
- Network transfer and client-side buffering before playback
Instrument each stage separately. Teams often blame the TTS vendor when the real delay is a too-cautious end-of-speech detector, or waiting for the LLM's entire answer before synthesizing anything.
Stream in, stream out
TTS APIs generally fall into three groups:
- Full text in, full file out. Simple, and fine for batch. Too slow for conversation.
- Full text in, streamed audio out. Audio chunks arrive as they're generated, so playback starts early.
- Streamed text in, streamed audio out. You push text as it arrives, usually over a WebSocket, and the service decides when to speak.
If you're feeding TTS from an LLM, you need at least the second pattern plus your own chunking, or the third. With chunking, the trade-off is direct. Send tiny fragments and you get speed but choppy, robotic phrasing — the model can't plan intonation for a sentence it hasn't seen. Wait for full paragraphs and the phrasing is natural but the user sits in silence.
Splitting at sentence boundaries, with a minimum length, is a sensible default:
import re
# Sentence end: . ! or ? optionally followed by a closing quote or bracket, then whitespace
SENTENCE_END = re.compile(r'[.!?]["\')\]]?\s')
def chunk_for_tts(token_stream, min_chars=40):
"""Group streamed LLM tokens into sentence-sized chunks for TTS."""
buffer = ""
for token in token_stream:
buffer += token
# Search from min_chars so short openers like "Sure." merge with the next sentence
match = SENTENCE_END.search(buffer, min_chars)
while match:
cut = match.end()
yield buffer[:cut].strip()
buffer = buffer[cut:]
match = SENTENCE_END.search(buffer, min_chars)
if buffer.strip():
yield buffer.strip()
This is deliberately naive — "Dr. Lee" will split in the wrong place — but it shows the shape. Tune min_chars by listening, not by guessing.
Chunking has a side effect: each chunk is synthesized on its own, so intonation can reset at every boundary. Some APIs let you pass the surrounding text as context to smooth this out. If yours does, use it.
The unglamorous latency wins
- Pre-render fixed phrases. Greetings, "one moment while I check that", error messages. Synthesize them once at deploy time and play them from storage instantly.
- Keep connections warm. Opening a new connection or WebSocket for every utterance adds avoidable setup time.
- Ask for the right format. More on this below.
Voice quality: test on your text, not theirs
Vendor demos use text chosen to sound good. Your production text contains order numbers, abbreviations, prices, and product names nobody has heard of. Judge quality on that.
Build a torture script
Write 30–50 sentences drawn from what your product actually says, plus the hard cases:
- Numbers in every form: "$1,250.00", "3/4", "2026-10-05", "1.5x", phone numbers
- Abbreviations that depend on context: "Dr. Lee lives on Lee Dr.", "St. Louis", "5 mg"
- Acronyms people say differently: "SQL", "API", "GIF"
- Your brand names, product names, and the names of your customers' cities
- Long sentences, questions, and lists — where prosody usually breaks first
Have several people listen blind and rank candidates. Listen on the devices your users will use: phone speakers, earbuds, a car. And listen for a long stretch. A voice that charms in one sentence can grate after ten minutes.
Fix pronunciation before the text reaches TTS
Most pronunciation bugs are really text normalization bugs. You have three tools, roughly in order of reliability:
- Normalize the text yourself. Expand "$1,250.00" to "one thousand two hundred fifty dollars" in code, where you control the rules and can unit-test them.
- Use pronunciation dictionaries or SSML where the vendor supports them, for names and jargon that keep coming up.
- Tell the LLM it's writing for speech. Ask for no markdown, no bullet symbols, no URLs, short sentences, and numbers written the way they should be spoken. This fixes a surprising amount, and it shortens responses too.
Watch for generative glitches
Newer neural TTS models sound more natural than older systems, but they can fail in new ways. They sometimes skip a word, repeat a phrase, or produce a strange noise, and some don't produce identical audio for identical input. For batch work, a cheap guard is to run the output back through speech-to-text and compare it with the input text. A large mismatch flags a clip for re-rendering or a human listen.
Cost: understand what you're actually paying for
Hosted TTS is usually priced per character of input text, or per second of audio produced. Premium or custom voices often cost more than standard ones. Self-hosting an open-weight model trades that per-unit bill for compute you pay for whether it's busy or not.
| Hosted API | Self-hosted open model | |
|---|---|---|
| Pricing | Per character or per audio second | GPU/CPU time, busy or idle |
| Cheapest when | Volume is low, spiky, or uncertain | Volume is high and steady |
| Voice selection | Large catalog, custom voices on offer | Limited to what the model supports |
| Hidden costs | Premium voice tiers, re-renders | Ops, scaling, model upgrades |
Before you compare rates, look at what's driving your character count:
- LLM verbosity. Every character the model writes is a character you pay to speak. A system prompt that asks for concise spoken answers cuts TTS cost and latency at the same time.
- Repeated text. Many apps say the same things again and again. Cache synthesized audio keyed on a hash of the text, voice, model, and settings. When nondeterminism doesn't matter, a cache hit is free and instant.
- Wasted synthesis. If the user interrupts after one sentence, stop synthesizing the rest. Barge-in handling saves money as well as frustration.
- Previews and retries. UI previews, retried requests, and test environments all bill like production.
Self-hosting starts to make sense with high, steady volume; strict data-residency rules; a need to run offline or on-device; or a GPU fleet you already operate for other work. Otherwise, a hosted API is usually cheaper once you count engineering and on-call time — especially early, when your volume is still a guess.
The details that bite later
Audio format. Ask for the format your output channel needs instead of transcoding afterward. Phone systems typically expect 8kHz μ-law audio. Browsers handle MP3 and Opus well. Raw PCM avoids decode time but takes more bandwidth. Generating high-sample-rate audio for a phone line just wastes bytes and adds a resampling step.
Failure handling. Your TTS vendor will have a bad day. Decide in advance what happens then: a timeout, then a fallback to a second provider or a simpler voice, then falling back to showing text if you have a screen. A voice product that goes silent with no explanation feels broken in a way a slow web page doesn't.
Voice cloning and consent. Custom voices are a real product feature, but only clone a voice with the speaker's explicit, documented consent, and read your vendor's terms on it. Rules about disclosing synthetic voices — especially in automated phone calls — are tightening in several places, so check what applies to you before launch.
Swappability. Keep TTS behind a small internal interface: text and voice settings in, an audio stream out.
A decision path that works
- Name the job. Batch or real-time? This rules out half the options immediately.
- Write the torture script from your real text and the hard cases above.
- Shortlist two or three voices and run blind listening tests on target devices.
- Measure time to first audio in your actual pipeline, not in a vendor playground.
- Model the cost from realistic character counts, after trimming LLM verbosity and adding caching.
- Build in fallbacks and an abstraction layer before launch, not after the first outage.
The short version
Production TTS is mostly about everything around the voice: streaming text in sensible chunks, normalizing text before it's spoken, caching what repeats, and planning for failure. Measure time to first audio rather than total synthesis time. Test quality on your own text and your users' devices. And keep the provider swappable, because this field won't sit still.
Adding voice to your product and trying to make it feel fast and natural? Get in touch — we design and build audio and AI pipelines end to end.