Reviews & reception

What benchmarks and the community say — and how it stacks up against others. A non-official digest of public opinion; tap a card for detail and source.

11 result(s)

A genuine long-form breakthrough

Reviewers widely call 90-min, 4-speaker dialogue a real step change.

Long-formSlatorPraise

Open, free, self-hostable

Praised for MIT licensing and avoiding subscription/privacy concerns.

Open sourceCostPraise

Speaker consistency & naturalness

HN community credits its voice consistency and conversational naturalness.

Hacker NewsNaturalnessPraise

67% rate expressiveness higher

An AllAboutAI survey: 67% of technical users rate it above Chatterbox-TTS.

AllAboutAIExpressiveness

Intonation still slips

An HN user: intonation is off on nearly every phrase, with robotic modulation.

Hacker NewsIntonationCriticism

Impressive, but not by today's bar

Some find it impressive vs years ago, yet uninspiring by current standards.

Hacker NewsCriticism

Language & feature limits

Reviews agree: EN/ZH-first, no overlapping speech, music or sound effects.

LimitsLanguagesCriticism

The 2026 models arrived quietly

Streaming ASR shipped with no launch post and, so far, no independent benchmarks.

StreamingBenchmarksCriticism

MOS 4.5 — tops the benchmarks

Per MS's report, the 7B model's MOS beat ElevenLabs v3 and Gemini TTS.

BenchmarksMOSvs others

vs ElevenLabs

Wins on long-form and cost; ElevenLabs still leads on polish and languages.

ElevenLabsvs others

vs F5-TTS / XTTS

Better long-dialogue consistency; F5 still praised for single-utterance quality.

F5-TTSXTTSvs others