The model family

Six open checkpoints, one design language. Tap a card for the full detail.

5 result(s)

VibeVoice-TTS · 1.5B

Long-form multi-speaker synthesis: up to 90 minutes and 4 distinct speakers in a single generation.

1.5B90 min4 speakersEN / ZH

VibeVoice-ASR · 7B

Long-form recognition: 60 minutes in one pass, with who / when / what — speaker, timestamp and content.

7B60 min50+ languages64K

VibeVoice-Realtime · 0.5B

Lightweight streaming: ~300 ms first-audio latency, streaming text input, robust to ~10 minutes.

0.5B~300 msstreamingmultilingual

VibeVoice-ASR-Streaming · 1.5B / 7B

Transcribes who said what while the audio is still arriving, emitting a line roughly every 2.9 seconds.

1.5B / 7Bstreaming10 languagesSep 2026

VibeVoice-ASR-BitNet · CPU

The ASR model squeezed from 4.62 GB to 1.58 GB — faster than real time on a few CPU threads, no GPU.

1.58 GBCPU onlyRTF < 1Jul 2026