VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Sa palibot sa Aïn Ouaïd. Limitahan sa 5,000 ka karakter

Ang yuta palibot sa Ssm kay medyo kabukiran.

<speak><prosody rate="slow">Slow speech</prosody></speak>

Ang mga tag sa gipili nga modelo makasabut - i-klik aron ihulog ang usa sa imong teksto diin kini mahitabo:

Ang modelong kini mobasa sa yano nga teksto, busa ang mga inline tags gi-ignore. Alang sa mga tag-based nga emosyon, i-usab sa usa ka ekspresyonal nga modelo sama sa Orpheus o Bark.

Ang yuta palibot sa Cerro La Pronunciación kay lain-lain.

-12 +12
0.5x 2.0x
Sa rehiyon palibot sa Piper, mga lawis talagsaon komon.
Ang imong na-generate nga audio mopakita dinhi. Pilia ang usa ka modelo, i-type ang teksto, ug i-klik ang Genere.
Ang audio maayong natukod
0:00
I-download ang Audio Sa palibot sa Srt. Hapit nalukop sa kaumahan ang palibot sa 24H.
Sa palibot sa ‘En ‘Alam. Ang yuta palibot sa $5 Mine kay lain-lain.
Love TTS.ai? Tell your friends!

Sa palibot sa Aïn el Aïd. VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Sa palibot sa Best.: General-purpose text-to-speech with natural prosody

Lawak ang lahat VITS Tingog

Sa palibot sa Glance.

Pag-uswag
Jaehyeon Kim et al.
Lisensiya
MIT
Tigre
free
Katulin
fast
Sa palibot sa Klondike.
Wala
Linguistics
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Maksimum nga mga karakter
2000

VITS Tingog

CSS10 (Dutch)

Dutch
Libre Neutral

CSS10 (Finnish)

Finnish
Libre Neutral

CSS10 (French)

French
Libre Neutral

CSS10 (German)

German
Libre Neutral

CSS10 (Hungarian)

Hungarian
Libre Neutral

CSS10 (Spanish)

Spanish
Libre Neutral

Common Voice (Bulgarian)

Bulgarian
Libre Neutral

Common Voice (Portuguese)

Portuguese
Libre Neutral

Default

English
Libre Neutral

MAI (Polish)

Polish
Libre Female

MAI (Ukrainian)

Ukrainian
Libre Neutral

VITS Sa palibot sa FAQ.

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Sa palibot sa Votos.