VITS TTS
The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về VITS
VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.
Tốt nhất cho: General-purpose text-to-speech with natural prosody
& Xem tất cả VITS giọng nóiMột cái nhìn
- Nhà phát triển
- Jaehyeon Kim et al.
- Giấy phép
- MIT
- Thú
- free
- Tốc độ
- fast
- Ký âm
- Không
- Ngôn ngữ
- English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
- Tối đa các ký tự
- 2000