TL;DR
Microsoft's cutting-edge open-source voice AI framework, including long-form TTS, real-time TTS, and multilingual ASR models.
Key features
VibeVoice-TTS: Long-form multi-speaker text-to-speech supporting up to 90 minutes and 4 speakers.
VibeVoice-Realtime-0.5B: Real-time TTS with streaming input, supporting multiple languages and various voice styles.
VibeVoice-ASR: Single processing of up to 60 minutes, structured transcription with speaker/time/content, supporting over 50 languages, compatible with Transformers and vLLM.
Fine-tuning code and Colab demos provided.
When to use it
When generating long-form audiobooks, podcasts, or multi-speaker dialogues.
For real-time speech synthesis applications (chatbots, voice assistants).
Suitable for research and services requiring multilingual long-form speech recognition and speaker diarization transcription.