Speech generation

The fastest and most natural text to speech model

Ranked #1 for naturalness, sub-90ms latency, and natively multilingual across 44 languages.

Naturalness

#1 ranked

Speech Arena style leaderboard claim from the source pages. Hear it via sales — no public player here.

Latency

Sub-90ms

Fast enough for live conversation, not a rendered file you wait on.

Languages

44 native

One model family. No separate stack per locale.

American EnglishBritish EnglishSpanishFrench GermanPortugueseHindiJapanese KoreanMandarin+ 34 more