Why we built the fastest robust TTS model

Our latest Text-to-Speech model takes around 50ms for the first audio chunk to arrive. These are the lowest latency figures out there among frontier TTS models.

Generally latency and accuracy are traded values. ElevenLabs, for example, ships its newest model in two variants: Eleven v4, "tuned for produced content where quality matters most", and Eleven v4 Turbo, a separate variant focused on lower latency. Choosing a fast model does not have to come with cutting corners for model robustness. Gradium's latest Text-to-Speech model beats Eleven v4 Turbo on latency while improving on naturalness and expressivity in all languages, and retaining our strengths of handling hard cases like phone numbers, alpha-numeric contextual text.