VibeVoice TT TT TTT TTT T TT TT T T TTT TTT TTT
Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.
在 SSML 标记中折行文本以精确控制 :
<speak><prosody rate="slow">Slow speech</prosody></speak>
标记选中模式的理解度 - 单击将一个输入到文本中, 发生时 :
这个模型读的是简单的文字, 所以内嵌标签会被忽略。 对于基于标签的情感, 请切换到像 Orpheus 或 Bark 这样的表达模式 。
定义自定义发音( Word = 发音) :
关于 VibeVoice
VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.
最佳: Podcasts, dialogues, long-form narration, multi-speaker content
全部浏览 VibeVoice 声音一眼看一眼,
- 开发者
- Microsoft
- 许可证
- MIT
- 级别
- standard
- 速度
- fast
- 语音克隆
- 否 无
- 语言
- English, Chinese
- 最大字符
- 50000