VibeVoice

VibeVoice TTS

Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.

Ku soo biir 5,000 xaraf xaddid

Wrap qoraalka ku SSML tags si loo hubiyo xakamaynta saxda ah:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags qaabka la doortay fahmo — riix si aad u hoos mid ka mid ah qoraalka aad halkaas oo uu ka dhacaa:

Model this akhriyo qoraalka caadiga ah, sidaas inline tags waa la iska indho tiri. For tag-ku salaysan dareenka, u dhaqaaqo si ay u muujiyaan qaabka sida Orpheus ama Bark.

Define custom pronunciations (word = dhawaaqa):

-12 +12
0.5x 2.0x
Bilaash ah oo leh Piper, VITS, MeloTTS
Your audio soo saaro halkan ka muuqan doonaa. Dooro qaab, ku qor qoraalka, oo guji soo saaro.
Dhaqdhaqaaqa ayaa la soo saaray
0:00
Soo dejisa Soo dejisan.srt Xidhiidhku wuxuu dhamaaday 24h
Free tier: isticmaalka shakhsiga ah. Liisan ganacsi laga bilaabo $ 5 / mo
Jecel TTS.ai? Ka warran saaxiibadaa!

About VibeVoice

VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.

Ugu Fiican: Podcasts, dialogues, long-form narration, multi-speaker content

Taabo oo kala soo bax VibeVoice cod

Eeg

Soo-saarayaasha
Microsoft
Liisan
MIT
Qiyaamaha
standard
Xawaaraha
fast
Duubista Codka
Ha
Afaf
English, Chinese
Noocyada ugu badan
50000

VibeVoice cod

Speaker 1

English
Caadi Neutral

Speaker 1 (Chinese)

Chinese
Caadi Neutral

Speaker 2

English
Caadi Neutral

Speaker 2 (Chinese)

Chinese
Caadi Neutral

Speaker 3

English
Caadi Neutral

Speaker 4

English
Caadi Neutral

VibeVoice Su'aalaha La Weydiiyo

VibeVoice supports up to 4 distinct speakers and up to 90 minutes of continuous output, with speaker tags for multi-turn dialogue — built for podcasts and long-form narration. It accepts up to 50,000 characters per request.

Yes. Alongside the 1.5B long-form model, a Realtime 0.5B variant achieves roughly 300ms latency for interactive use.

VibeVoice is MIT-licensed. It supports English and Chinese and does not currently support voice cloning.
← Codadka oo dhan