Open-source long-form voice AI

VibeVoice synthesises up to 90 minutes of natural, multi-speaker audio in a single pass — powered by a 7.5 Hz continuous speech tokenizer and a next-token diffusion design. Six open checkpoints now cover long-form TTS, real-time speech, long-form and streaming recognition, and CPU-only edge inference.

90min in one pass
4speakers at once
7.5Hz tokenizer
6open checkpoints

The model family

Six open checkpoints, one design language. Tap a card for the full detail.

Model comparison

All five VibeVoice models side by side — pick the right one for the job.

How it works

A 7.5 Hz tokenizer plus next-token diffusion: how VibeVoice fits 90 minutes of audio inside a 64K context — and how that design later stretched to streaming and to CPUs.

Specs & numbers

The headline figures behind VibeVoice, plus a full per-model spec table covering all five models.

TTS comparison

VibeVoice next to popular and open-source TTS models. A fair, non-exhaustive sketch — strengths differ by use case.

ASR comparison

VibeVoice-ASR next to FunASR and Whisper — open-source speech-to-text, side by side. Updated for the streaming and CPU builds.

Realtime voice

Low-latency streaming voice — VibeVoice-Realtime next to OpenAI's latest realtime model and CosyVoice 2.

Reviews & reception

What benchmarks and the community say — and how it stacks up against others. A non-official digest of public opinion; tap a card for detail and source.

Where it fits

Long-form, multi-speaker synthesis unlocks workflows that single-utterance TTS struggles with — and the 2026 recognition builds add live and on-device ones.

Development timeline

From the first TTS release to streaming, speaker-attributed recognition — and a build that runs on a CPU.

Limitations & responsible AI

What VibeVoice can't do, what it needs, and how to use it responsibly. Search to jump to a question.

Community builds & safety

People have wrapped VibeVoice into ComfyUI nodes, desktop apps, API servers and native ports. Here is what exists — and what to check before you run any of it. Signals captured September 2026; tap a card for the detail and the repo link.

Resources & links

Official source, weights, a live demo and further reading. Filter by type or search.