Gradium TTS: no longer choose between latency and accuracy
Gradium releases a new TTS model available today that achieves 81% pass rate on a hard-case evaluation of 500 sentences across five languages (English, French, Spanish, Portuguese, German) and reduces time to first audio (TTFA P50) to 216 ms. The model correctly reads phone numbers, email addresses, IBANs and other structured entities without pre-processing, addressing common failure points for voice agents in production.
Voice AI models are still considered the dumber ones, and rightfully so. But it's great to watch the rocket-fast development where voice models will respond instantly and the quality of information provided will rival frontier AI models.
What is TTFA and why does it matter for text-to-speech model evaluation?
TTFA (time to first audio) is the latency between receiving text and playing the first audio output. It is critical for voice agents because high latency degrades user experience in phone calls and creates uncomfortable pauses in conversation flow.
Which types of text entities are most challenging for voice agents?
Structured entities such as phone numbers, email addresses, IBAN codes, confirmation numbers, and reference codes are most problematic because they require precise pronunciation of every digit or character. Missing a single digit makes the entire output unusable.
How does Gradium's testing approach differ from some competitors?
Gradium does not rewrite text with a hidden LLM before synthesis—meaning demo outputs match what the API produces. Some competitors use such preprocessing in demos to improve results, which do not carry over to production.
- Bland releases benchmark-topping new voice model — bland.ai 74 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 73 % match
- Static vs. dynamic languages for AI coding: complex problems shrink efficiency differences — danluu.com 73 % match
- Gradium
- Gradium TTS
- Voice Design
- ElevenLabs
- Cartesia Sonic
- Fish Audio
- Inworld
- NVIDIA8
- Hugging Face12