IndexTTS-2

IndexTTS-2 TTS

A zero-shot TTS model with fine-grained emotion control via emotion vectors, no emotion-specific training data required.

Prijavite se za ograničenje od 5.000 znakova

Omotajte tekst u SSML oznake za preciznu kontrolu:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Oznake koje odabrani model razumije — kliknite da biste ih ubacili u tekst gdje se pojavljuju:

Ovaj model čita običan tekst, tako da se inline oznake ignoriraju. Za emocije zasnovane na oznakama, prebacite se na ekspresivni model poput Orpheusa ili Bark-a.

Definirajte vlastite izgovore (riječ = izgovor):

-12 +12
0.5x 2.0x
Besplatno sa Piper, VITS, MeloTTS
Ovdje će se pojaviti vaš generirani audio. Izaberite model, unesite tekst i kliknite na Generiraj.
Audio uspješno generisan
0:00
Preuzmi audio Preuzmi.srt Link istječe za 24h
Free tier: personal use. Komercijalna licenca od $5/mjesečno
Volite TTS.ai?

O meni IndexTTS-2

IndexTTS-2, from the Index Team, is an expressive text-to-speech system that pairs zero-shot voice synthesis with precise emotional control. Rather than relying on emotion-labeled training data, it uses emotion vectors to dial in tones like happy, sad, angry, or fearful independently of the voice itself. Built on a Qwen2 backbone with BigVGAN as the vocoder, it supports English and Chinese and can clone a voice from roughly five seconds of reference audio. It suits audiobooks, virtual assistants, and any content where the same voice needs to shift emotional register. Its weights use the Bilibili Model License, which permits commercial use below large usage and revenue thresholds.

Najbolje za: Emotionally expressive content, audiobooks, virtual assistants

Pregledaj sve IndexTTS-2 glasovi

Na prvi pogled

Programer
Index Team
Licenca
Bilibili Model License
Životinje
standard
Brzina
medium
Kloniranje glasa
Da.
Jezici
English, Chinese
Maksimalan broj znakova
1000

IndexTTS-2 glasovi

Chinese Default

Chinese
Standardni Neutral

Default

English
Standardni Neutral

IndexTTS-2 FAQ

It uses emotion vectors that let you specify tones such as happy, sad, angry, or fearful without needing emotion-specific training data, and the emotional expression is controlled independently from the voice identity.

Yes. It performs zero-shot voice cloning from a short reference, typically around five seconds of audio, in English or Chinese.

Its weights are released under the Bilibili Model License, which allows commercial use for products below defined user and revenue thresholds. Larger deployments should review the license terms.
← Svi glasovi