VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Tia sahihi kwa kiwango cha tabia 5,000

Pakua maandishi yako katika tovuti ya SSML kwa ajili ya udhibiti sahihi:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag anaelewa mfano unaochaguliwa na unajibu ujumbe huu:

Mfano huu unasomeka maandishi rahisi, kwa hiyo alama za vidole hupuuzwa. Kwa hisia za ndani za watu, geukia kigezo kinachoonesha hisia kama Orfeus au Bark.

Matamshi ya desturi (neno = matamshi):

-12 +12
0.5x 2.0x
Nikiwa huru na Piper, VITS, MelloTTTS
Unaweza kuchagua mfano, maandishi, na kidofo kinachoitwa Genete.
Edio Iliyorekebishwa kwa Mafanikio
0:00
Paketi ya Audio Paketisha.srt Kiungo kinakufa mnamo 24
Safu huru: matumizi ya kibinafsi. Hati ya biashara kutoka dola 5/mo
Fanya hii sauti yako mwenyewe Chokoa sauti kwa sekunde 30
Waeleze rafiki zako kuhusu mapenzi ya TTS.ai?

Habari VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Bora kwa: General-purpose text-to-speech with natural prosody

Ng'ombe wote VITS sauti

Kutupia jicho

Mbuni
Jaehyeon Kim et al.
Lenzi
MIT
Tier
free
Mwendo
fast
Kufanyizwa kwa Sauti
Hapana
Lugha
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Wahusika wa Max
2000

VITS sauti

CSS10 (Dutch)

Dutch
Huru Neutral

CSS10 (Finnish)

Finnish
Huru Neutral

CSS10 (French)

French
Huru Neutral

CSS10 (German)

German
Huru Neutral

CSS10 (Hungarian)

Hungarian
Huru Neutral

CSS10 (Spanish)

Spanish
Huru Neutral

Common Voice (Bulgarian)

Bulgarian
Huru Neutral

Common Voice (Portuguese)

Portuguese
Huru Neutral

Default

English
Huru Neutral

MAI (Polish)

Polish
Huru Female

MAI (Ukrainian)

Ukrainian
Huru Neutral

VITS TTS ngumuSTEGAQ

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Sauti zote