VibeVoice TTS
Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.
Kpọchie ngwe gị n'ime SSML táàbụ̀ maka nlekọta ziri ezi:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Táàbụ̀ nke móòdù ahụ a họọrọ na-aghọta - pịa ka ịkpụga otu n'ime ngwe gị ebe ọ na-eme:
Móòdù a na-agụ ngwe nkịtị, yabụ na a na-ewepụta inline táàbụ̀. Maka táàbụ̀-n'okpuru n'émóòdù, gbanwee ka móòdù na-egosi ihe dịka Orpheus mọọbụ Bark.
Ndesịta okwu emeredịkachọrọ:
_N'ihe banyere VibeVoice
VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.
Ọkachasị maka: Podcasts, dialogues, long-form narration, multi-speaker content
Nlegharịa niile VibeVoice ụdaN'ime nlele
- Ńkwádò
- Microsoft
- Ikikere
- MIT
- Tier
- standard
- Nhazi
- fast
- Nhazi ụda
- Ọ bụghị
- Asụsụ ndị ahụ
- English, Chinese
- Ụhara Max
- 50000