Model comparison

All five VibeVoice models side by side — pick the right one for the job.

Realtime-0.5B
0.5B
real-time
TTS-1.5B
1.5B
flagship
ASR-7B
7B
recognition
ASR-Streaming
1.5B / 7B
live transcript
ASR-BitNet
1.58 GB
edge CPU
TaskTTS (stream)Text-to-SpeechSpeech-to-TextSTT (stream)Speech-to-Text
Max length~10 min90 min60 mincontinuous60 min
Speakers1up to 4
Context8K64K64Krolling64K
Streaming text inputcheck
Low latency (~300 ms)check~2.9 s chunk
Multi-speaker consistencycheck
Timestamps + diarizationcheckcheckcheck
Runs without a GPUcheck
LanguagesMulti-voiceEN / ZH first50+ languages10 languagesEN / ZH +
Best forAgent voiceLong podcastsTranscriptionLive captionsOn-device