IndexTTS-2 TT TT TTT TTT T TT TT T T TTT TTT TTT
A zero-shot TTS model with fine-grained emotion control via emotion vectors, no emotion-specific training data required.
在 SSML 标记中折行文本以精确控制 :
<speak><prosody rate="slow">Slow speech</prosody></speak>
标记选中模式的理解度 - 单击将一个输入到文本中, 发生时 :
这个模型读的是简单的文字, 所以内嵌标签会被忽略。 对于基于标签的情感, 请切换到像 Orpheus 或 Bark 这样的表达模式 。
定义自定义发音( Word = 发音) :
关于 IndexTTS-2
IndexTTS-2, from the Index Team, is an expressive text-to-speech system that pairs zero-shot voice synthesis with precise emotional control. Rather than relying on emotion-labeled training data, it uses emotion vectors to dial in tones like happy, sad, angry, or fearful independently of the voice itself. Built on a Qwen2 backbone with BigVGAN as the vocoder, it supports English and Chinese and can clone a voice from roughly five seconds of reference audio. It suits audiobooks, virtual assistants, and any content where the same voice needs to shift emotional register. Its weights use the Bilibili Model License, which permits commercial use below large usage and revenue thresholds.
最佳: Emotionally expressive content, audiobooks, virtual assistants
全部浏览 IndexTTS-2 声音一眼看一眼,
- 开发者
- Index Team
- 许可证
- Bilibili Model License
- 级别
- standard
- 速度
- medium
- 语音克隆
- 是
- 语言
- English, Chinese
- 最大字符
- 1000