TECH FLOW Svět Androida
← Back to the stream
gradium.ai · picked by Petr Mišák · 26d ago

Gradium TTS: no longer choose between latency and accuracy

Source preview: Gradium TTS: no longer choose between latency and accuracy
AI summary

Gradium releases a new TTS model available today that achieves 81% pass rate on a hard-case evaluation of 500 sentences across five languages (English, French, Spanish, Portuguese, German) and reduces time to first audio (TTFA P50) to 216 ms. The model correctly reads phone numbers, email addresses, IBANs and other structured entities without pre-processing, addressing common failure points for voice agents in production.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

4 people have already opened the source

Tip author’s note

Voice AI models are still considered the dumber ones, and rightfully so. But it's great to watch the rocket-fast development where voice models will respond instantly and the quality of information provided will rival frontier AI models.

AI questions & answers
What is TTFA and why does it matter for text-to-speech model evaluation?

TTFA (time to first audio) is the latency between receiving text and playing the first audio output. It is critical for voice agents because high latency degrades user experience in phone calls and creates uncomfortable pauses in conversation flow.

Which types of text entities are most challenging for voice agents?

Structured entities such as phone numbers, email addresses, IBAN codes, confirmation numbers, and reference codes are most problematic because they require precise pronunciation of every digit or character. Missing a single digit makes the entire output unusable.

How does Gradium's testing approach differ from some competitors?

Gradium does not rewrite text with a hidden LLM before synthesis—meaning demo outputs match what the API produces. Some competitors use such preprocessing in demos to improve results, which do not carry over to production.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions