Blog → podcast
Turn an article into a natural two-host conversation, generated in one pass with consistent voices.
Long-form, multi-speaker synthesis unlocks workflows that single-utterance TTS struggles with — and the 2026 recognition builds add live and on-device ones.
Turn an article into a natural two-host conversation, generated in one pass with consistent voices.
Teacher–student scripts, interviews and role-play voiced for courseware.
Low-latency voice interface for agents (Realtime).
Scale training narration without a studio.
Narrate technical docs and long articles in a single coherent take.
Keep up to 4 voices consistent inside one model.
Streaming ASR labels who is speaking as the meeting happens — no waiting for the recording to end.
The 1.58 GB CPU build keeps sensitive audio on the machine.