IndexTTS-2 TTS
A zero-shot TTS model with fine-grained emotion control via emotion vectors, no emotion-specific training data required.
Zabalte svůj text do značek SSML pro přesné ovládání:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Značky vybraného modelu rozumí? klikněte na tlačítko pro kapku jednoho do textu, kde se to stane:
Tento model čte prostý text, takže inline značky jsou ignorovány. Pro tag-based emotion, přepněte na expresivní model, jako je Orpheus nebo Bark.
Definovat vlastní výslovnosti (slovo = výslovnost):
O aplikaci IndexTTS-2
IndexTTS-2, from the Index Team, is an expressive text-to-speech system that pairs zero-shot voice synthesis with precise emotional control. Rather than relying on emotion-labeled training data, it uses emotion vectors to dial in tones like happy, sad, angry, or fearful independently of the voice itself. Built on a Qwen2 backbone with BigVGAN as the vocoder, it supports English and Chinese and can clone a voice from roughly five seconds of reference audio. It suits audiobooks, virtual assistants, and any content where the same voice needs to shift emotional register. Its weights use the Bilibili Model License, which permits commercial use below large usage and revenue thresholds.
Nejlepší pro: Emotionally expressive content, audiobooks, virtual assistants
Procházet vše IndexTTS-2 hlasyNa první pohled
- Vývojář
- Index Team
- Licence
- Bilibili Model License
- Úroveň
- standard
- Rychlost
- medium
- Klonování hlasu
- Ano.
- Jazyky
- English, Chinese
- Max znaků
- 1000