VibeVoice

VibeVoice TTS

Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.

Ṣẹ̀dà fun àwọn àmì-àṣírí 5,000

Fi àkọlé rẹ pamọ́ sí àwọn àmì-ìwé SSML fún ìdáràn:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Àwọn Àmì-ìwé tí àwọn ìṣàmúlò-ètò tí a yàn gbọ́ - tẹ̀ láti fi ọkan sínú àkọ́lé rẹ̀ nínú àwọn ààyè-iṣẹ́ tí o bá jẹ́:

Àwọn àwọn àkọlé àwọn ààyè-iṣẹ́ àwọn àwọn àmì-ìwé àwọn àmì-ìwé àwọn àwọn àmì-ìwé àwọn à

Àwọn àwọn ìṣàfarawé àwọn àwọn ìṣàfarawé àwọn (ọrọ = ìṣàfàlì):

-12 +12
0.5x 2.0x
Free pẹlu Piper, VITS, MeloTTS
Àwọn àwòrán tí o ti ṣẹ̀dà tí o bá han níbẹ̀. Yan àwọn àwòrán, tẹ̀lẹ̀ àkọlé, ki o si tẹ̀ Ṣẹ̀dà.
Àwọn àwọn àwòrán tí a ṣẹ̀dà
0:00
Ṣàfikún Àwọn Àmì-ìwé Ṣàfikún.srt Líǹkì náà kù nínú 24h
Ìjádé ọ̀fẹ́: ìlòjútó ara ẹni. Lisensi Iṣowo ori lati $5/mo
O fẹ́ TTS.ai? Fì sọ̀kalẹ̀ fún àwọn ọrẹ̀ rẹ̀!

Ààyè-iṣẹ́ VibeVoice

VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.

Tí o dara jù fún: Podcasts, dialogues, long-form narration, multi-speaker content

Wá Gbogbo àwòrán VibeVoice Àwọn àwòrán

Nínú àwọn ìṣàfarawé

Àwọn Àkọlé
Microsoft
Àwọn Ààyè-iṣẹ́
MIT
Àwọn àwọn ààyè-iṣẹ́
standard
Ìjánu-ìṣàmúlò-ètò
fast
Ìṣàfarawé àwọn àmì-ìwé
Àwọn àwọn àgbéwọlé
Àwọn
English, Chinese
Àwọn àyọkà ìpele
50000

VibeVoice Àwọn àwòrán

Speaker 1

English
Àwọn ìpéwọ̀n Neutral

Speaker 1 (Chinese)

Chinese
Àwọn ìpéwọ̀n Neutral

Speaker 2

English
Àwọn ìpéwọ̀n Neutral

Speaker 2 (Chinese)

Chinese
Àwọn ìpéwọ̀n Neutral

Speaker 3

English
Àwọn ìpéwọ̀n Neutral

Speaker 4

English
Àwọn ìpéwọ̀n Neutral

VibeVoice Àwọn Àtòjọ-ẹ̀yàn

VibeVoice supports up to 4 distinct speakers and up to 90 minutes of continuous output, with speaker tags for multi-turn dialogue — built for podcasts and long-form narration. It accepts up to 50,000 characters per request.

Yes. Alongside the 1.5B long-form model, a Realtime 0.5B variant achieves roughly 300ms latency for interactive use.

VibeVoice is MIT-licensed. It supports English and Chinese and does not currently support voice cloning.
← Gbogbo àwọn ìrànwọ́