Where it fits

Long-form, multi-speaker synthesis unlocks workflows that single-utterance TTS struggles with — and the 2026 recognition builds add live and on-device ones.

90 min

Blog → podcast

Turn an article into a natural two-host conversation, generated in one pass with consistent voices.

Educational dialogue

Teacher–student scripts, interviews and role-play voiced for courseware.

AI agent voice

Low-latency voice interface for agents (Realtime).

Corporate e-learning

Scale training narration without a studio.

Long-form narration

Narrate technical docs and long articles in a single coherent take.

Multi-speaker narration

Keep up to 4 voices consistent inside one model.

Live captions & meeting notes

Streaming ASR labels who is speaking as the meeting happens — no waiting for the recording to end.

On-device & offline

The 1.58 GB CPU build keeps sensitive audio on the machine.