VITS टीटीएस
The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.
अचूक नियंत्रण करीता SSML टॅग अंतर्गत पाठ्य वेल्ड करा:
<speak><prosody rate="slow">Slow speech</prosody></speak>
निवडलेले नमूना समजून घेणारे टॅग - पाठ्य अंतर्गत एक टाका जेथे ते घडते:
हे मॉडेल सादा पाठ्य वाचते, त्यामुळे इनलाईन टॅग दुर्लक्ष केले जातात. टॅग आधारीत भावना करीता, Orpheus किंवा Bark सारखे अभिव्यक्ती मॉडेल करीता बदलवा.
इच्छिक उच्चारण निश्चित करा (शब्द = उच्चारण):
विषयी VITS
VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.
सर्वोत्तम: General-purpose text-to-speech with natural prosody
सर्व ब्राऊज करा VITS आवाजएक नजर
- डेव्हलपर
- Jaehyeon Kim et al.
- परवाना
- MIT
- टर
- free
- वेग
- fast
- आवाज क्लोन
- नाही
- भाषाName
- English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
- कमाल अक्षरे
- 2000