IndexTTS-2 TTSName
A zero-shot TTS model with fine-grained emotion control via emotion vectors, no emotion-specific training data required.
ტექსტის გადატანა SSML ჭდეებში ზუსტი კონტროლისთვის:
<speak><prosody rate="slow">Slow speech</prosody></speak>
მონიშნული მოდელისთვის გასაგები ჭდეები - დააწკაპუნეთ, რომ ერთი ჭდე თქვენს ტექსტში ჩააგდოთ, სადაც ის მოხდება:
ეს მოდელი კითხულობს ჩვეულებრივ ტექსტს, ამიტომ შიგნით თავსებული ჭდეები იგნორირდება. ჭდეებზე დაფუძნებული ემოციებისთვის გადადით გამოხატვის მოდელში, როგორიცაა Orpheus ან Bark.
ინდივიდუალური გამოთქმების განსაზღვრა (სიტყვი = გამოთქმა):
ინფორმაცია IndexTTS-2
IndexTTS-2, from the Index Team, is an expressive text-to-speech system that pairs zero-shot voice synthesis with precise emotional control. Rather than relying on emotion-labeled training data, it uses emotion vectors to dial in tones like happy, sad, angry, or fearful independently of the voice itself. Built on a Qwen2 backbone with BigVGAN as the vocoder, it supports English and Chinese and can clone a voice from roughly five seconds of reference audio. It suits audiobooks, virtual assistants, and any content where the same voice needs to shift emotional register. Its weights use the Bilibili Model License, which permits commercial use below large usage and revenue thresholds.
საუკეთესოა: Emotionally expressive content, audiobooks, virtual assistants
ყველას დათვალიერება IndexTTS-2 ხმაჟ მალკჲ ოჲდლვენსგაŒვ
- პროგრამისტი
- Index Team
- ლიცენზია
- Bilibili Model License
- იანვარი
- standard
- სიჩქარე
- medium
- ხმა
- ეა
- ენაName
- English, Chinese
- სიმბოლოების მაქსიმალური რაოდენობა
- 1000