VibeVoice

VibeVoice TTS

Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.

Bhalisa Uluhlu lwezinto zobumnini Zolwaleko...

Ulawulo oluchanekileyo:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Ii-tags imodeli ekhethiweyo iqonda - nqakraza ukushiya enye kumbhalo wakho apho isenza khona:

Le modeli ifunda umbhalo oqhelekileyo, ngoko ke i-inline tags ilahleka. Uphawu olusekelwe kwi-emotions, tshintshela kwimodeli ebonisa umbono njenge-Orpheus okanye i-Bark.

Chaza ubeko lwephepha

-12 +12
0.5x 2.0x
Ikhululekile nge Piper, VITS, MeloTTS
Isandi sakho esivelisweyo siza kuvela apha. Khetha imodeli, ngenisa umbhalo, kwaye unqakraze Yenza.
Isandi Sizaliswe Ngempumelelo
0:00
Layisha ezantsi Layisha ezantsi Ikhonkco liphelelwe lixesha kwiyure ezi-24
Inqanaba elikhululekileyo: ukusetyenziswa komuntu siqu. Ilayisensi yezorhwebo ukusuka kwi- $5/inyanga
Uthando TTS.ai? Nceda utshele abalandeli bakho!

I-About VibeVoice

VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.

Elungileyo: Podcasts, dialogues, long-form narration, multi-speaker content

Khangela konke VibeVoice iilizwi

Kwingxelo

Umbhekisi phambili
Microsoft
Ilayisensi
MIT
I-Tier
standard
Isantya
fast
Ukuphinda usebenzise ilizwi
Akukho nanye
Iilwimi
English, Chinese
Ubukhulu bamagama
50000

VibeVoice iilizwi

Speaker 1

English
Emiselweyo Neutral

Speaker 1 (Chinese)

Chinese
Emiselweyo Neutral

Speaker 2

English
Emiselweyo Neutral

Speaker 2 (Chinese)

Chinese
Emiselweyo Neutral

Speaker 3

English
Emiselweyo Neutral

Speaker 4

English
Emiselweyo Neutral

VibeVoice TTS - Imibuzo ebuzwa rhoqo

VibeVoice supports up to 4 distinct speakers and up to 90 minutes of continuous output, with speaker tags for multi-turn dialogue — built for podcasts and long-form narration. It accepts up to 50,000 characters per request.

Yes. Alongside the 1.5B long-form model, a Realtime 0.5B variant achieves roughly 300ms latency for interactive use.

VibeVoice is MIT-licensed. It supports English and Chinese and does not currently support voice cloning.
← Zonke iingoma