VibeVoice

VibeVoice TTS

Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.

Misoratra anarana fetra 5000 marika

Ampidiro anatin'ny tag SSML ny lahabolana mba hahazoana fifehezana mazava tsara:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag fantatry ny modely voafaritra — tsindrio mba hametrahana iray ao anatin'ny lahatsoratrao izay misy azy:

Mamakiana lahabolana tsotra io modely io, ka tsy raharahaina ny tag anatin'ny andalana. Raha mila fihetseham-po mifototra amin'ny tag ianao, dia miova ho modely maneho fihetseham-po toy ny Orpheus na Bark.

Mamaritra ny fanononana safidy (teny = fanononana):

-12 +12
0.5x 2.0x
Malalaka miaraka amin'ny Piper, VITS, MeloTTS
Hiseho eto ny feo namoronanao. Misafidiana modely iray, soraty ny lahabolana, dia tsindrio ny Mamorona.
Namorona feo tsara
0:00
Handefa feo Hidina.srt Tapitra ao anatin'ny 24 ora ity rohy ity
Ny faritr'ora dia GMT+1. : Tranonkala ofisialy Lisansa ara-barotra manomboka amin'ny $5/volana
Tianao ve ny TTS.ai? Lazao amin'ny namanao!

Mombamomba VibeVoice

VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.

Tsara indrindra ho an'ny: Podcasts, dialogues, long-form narration, multi-speaker content

Jereo izy rehetra VibeVoice feo

Amin'ny fijery fohy

Mpamorona
Microsoft
Lisansa
MIT
Taona
standard
Hafainganan'ny fanovana
fast
Fandraisana feo
Tsy misy
Teny
English, Chinese
Marika betsaka indrindra
50000

VibeVoice feo

Speaker 1

English
Standard Neutral

Speaker 1 (Chinese)

Chinese
Standard Neutral

Speaker 2

English
Standard Neutral

Speaker 2 (Chinese)

Chinese
Standard Neutral

Speaker 3

English
Standard Neutral

Speaker 4

English
Standard Neutral

VibeVoice TTS - Fanontaniana matetika

VibeVoice supports up to 4 distinct speakers and up to 90 minutes of continuous output, with speaker tags for multi-turn dialogue — built for podcasts and long-form narration. It accepts up to 50,000 characters per request.

Yes. Alongside the 1.5B long-form model, a Realtime 0.5B variant achieves roughly 300ms latency for interactive use.

VibeVoice is MIT-licensed. It supports English and Chinese and does not currently support voice cloning.
← Ny feo rehetra