VibeVoice TTS
Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.
Fi àkọlé rẹ pamọ́ sí àwọn àmì-ìwé SSML fún ìdáràn:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Àwọn Àmì-ìwé tí àwọn ìṣàmúlò-ètò tí a yàn gbọ́ - tẹ̀ láti fi ọkan sínú àkọ́lé rẹ̀ nínú àwọn ààyè-iṣẹ́ tí o bá jẹ́:
Àwọn àwọn àkọlé àwọn ààyè-iṣẹ́ àwọn àwọn àmì-ìwé àwọn àmì-ìwé àwọn àwọn àmì-ìwé àwọn à
Àwọn àwọn ìṣàfarawé àwọn àwọn ìṣàfarawé àwọn (ọrọ = ìṣàfàlì):
Ààyè-iṣẹ́ VibeVoice
VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.
Tí o dara jù fún: Podcasts, dialogues, long-form narration, multi-speaker content
Wá Gbogbo àwòrán VibeVoice Àwọn àwòránNínú àwọn ìṣàfarawé
- Àwọn Àkọlé
- Microsoft
- Àwọn Ààyè-iṣẹ́
- MIT
- Àwọn àwọn ààyè-iṣẹ́
- standard
- Ìjánu-ìṣàmúlò-ètò
- fast
- Ìṣàfarawé àwọn àmì-ìwé
- Àwọn àwọn àgbéwọlé
- Àwọn
- English, Chinese
- Àwọn àyọkà ìpele
- 50000