VITS

VITS ТТС

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Бүртгүүлэх 5000 тэмдэгтээс хэтрэхгүй

Тодорхой хяналтын тулд SSML тэмдгээр текстээ буулгах:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Бүх

Энэ загвар нь энгийн утга уншдаг, тиймээс доторх тэмдгийг үл тоомсорлодог. Тэг- суурилсан сэтгэл хөдлөл, Orpheus эсвэл Bark- шиг илэрхийлэх загвар руу шилжинэ.

Өөрийн дуудлагыг тодорхойлох (үг = дуудлага):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS-тэй чөлөөт
Таны үүсгэсэн дууны файл энд гарч ирнэ. Модель сонгож, текстийг оруулж, Бүтээгдэх товчийг дарна уу.
Аудио амжилттай бүтээгдсэн
0:00
Дуу татаж авах .srt татаж авах Холбоо 24 цагийн дараа дуусна
Хязгааргүй: хувийн хэрэглээ. Бизнесийн лиценз $5/сараас
TTS.ai-г хайрладаг уу? Найзуудаа хэлж өгөөрэй!

Тодорхойлолт VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Хамгийн сайн: General-purpose text-to-speech with natural prosody

Бүхнийг харах VITS дуунууд

Нүүр хуудас

Хөгжүүлэгч
Jaehyeon Kim et al.
Лиценз
MIT
Үхрийн
free
Хурд
fast
Дууны дугуй
Үгүй
Хэл
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Хамгийн их үсэг
2000

VITS дуунууд

CSS10 (Dutch)

Dutch
Хязгааргүй Neutral

CSS10 (Finnish)

Finnish
Хязгааргүй Neutral

CSS10 (French)

French
Хязгааргүй Neutral

CSS10 (German)

German
Хязгааргүй Neutral

CSS10 (Hungarian)

Hungarian
Хязгааргүй Neutral

CSS10 (Spanish)

Spanish
Хязгааргүй Neutral

Common Voice (Bulgarian)

Bulgarian
Хязгааргүй Neutral

Common Voice (Portuguese)

Portuguese
Хязгааргүй Neutral

Default

English
Хязгааргүй Neutral

MAI (Polish)

Polish
Хязгааргүй Female

MAI (Ukrainian)

Ukrainian
Хязгааргүй Neutral

VITS ТТС - Тодорхойгүй асуултууд

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Бүх дуунууд