VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Đăng ký giới hạn 5000 ký tự

Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:

Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.

Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):

-12 +12
0.5x 2.0x
Miễn phí với Piper, VITS, MeloTTS
Âm thanh đã tạo sẽ xuất hiện ở đây. Chọn một mô hình, nhập văn bản, và nhấn vào Tạo.
Âm thanh đã được tạo thành công
0:00
Tải về âm thanh Tải về Liên kết hết hạn trong 24h
Free tier: sử dụng cá nhân. Giấy phép thương mại từ $5/mo
Cảm ơn bạn đã tin tưởng TTS.ai!

Về VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Tốt nhất cho: General-purpose text-to-speech with natural prosody

& Xem tất cả VITS giọng nói

Một cái nhìn

Nhà phát triển
Jaehyeon Kim et al.
Giấy phép
MIT
Thú
free
Tốc độ
fast
Ký âm
Không
Ngôn ngữ
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Tối đa các ký tự
2000

VITS giọng nói

CSS10 (Dutch)

Dutch
Tự do Neutral

CSS10 (Finnish)

Finnish
Tự do Neutral

CSS10 (French)

French
Tự do Neutral

CSS10 (German)

German
Tự do Neutral

CSS10 (Hungarian)

Hungarian
Tự do Neutral

CSS10 (Spanish)

Spanish
Tự do Neutral

Common Voice (Bulgarian)

Bulgarian
Tự do Neutral

Common Voice (Portuguese)

Portuguese
Tự do Neutral

Default

English
Tự do Neutral

MAI (Polish)

Polish
Tự do Female

MAI (Ukrainian)

Ukrainian
Tự do Neutral

VITS TTS - FAQ

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Tất cả giọng nói