VITS TTS
The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.
Whāriki i tōna kupu i roto i ngā tohu SSML mō te whakahaere tika:
<speak><prosody rate="slow">Slow speech</prosody></speak>
E mōhio ana ngā tohu ki te tauira i kōwhiria - ka kōwhiria kia whakawātea tētahi ki roto i tōna kupu i reira ka puta ai:
Ka pānui tēnei tauira i te kupu noa, nā reira ka whakakāhoretia ngā tohu ā-waitara. Mō te āhua o te tohu-taihi, ka huri ki tētahi tauira whakamārama pēnei i a Orpheus, Bark rānei.
Ka tautuhia ngā tohutohu ā-ringa (wāhi = tohutohu):
Mo VITS
VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.
Pai mo: General-purpose text-to-speech with natural prosody
Ka tirohia katoa VITS ngā oroI te tirohanga
- Ka whakawhanakehia
- Jaehyeon Kim et al.
- Ka taea te whakawātea
- MIT
- Karaka
- free
- Āhuatanga
- fast
- Whakakōrero reo
- Kāore
- reo
- English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
- Kāri nui rawa
- 2000