VibeVoice-TTS · 1.5B
Long-form multi-speaker synthesis: up to 90 minutes and 4 distinct speakers in a single generation.
Six open checkpoints, one design language. Tap a card for the full detail.
5 result(s)
Long-form multi-speaker synthesis: up to 90 minutes and 4 distinct speakers in a single generation.
Long-form recognition: 60 minutes in one pass, with who / when / what — speaker, timestamp and content.
Lightweight streaming: ~300 ms first-audio latency, streaming text input, robust to ~10 minutes.
Transcribes who said what while the audio is still arriving, emitting a line roughly every 2.9 seconds.
The ASR model squeezed from 4.62 GB to 1.58 GB — faster than real time on a few CPU threads, no GPU.