VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Prihlásiť sa na odber Limit 5 000 znakov

Zabaliť text do SSML značiek pre presnú kontrolu:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Značky, ktorým vybraný model rozumie — kliknutím ich umiestnite do textu tam, kde sa vyskytujú:

Tento model číta obyčajný text, takže vnorené značky sa ignorujú.Pre emócie založené na značkách prejdite na expresívny model ako Orpheus alebo Bark.

Definovať vlastné výslovnosti (slovo = výslovnosť):

-12 +12
0.5x 2.0x
Zadarmo s Piper, VITS, MeloTTS
Vyberte si model, zadajte text a kliknite na tlačidlo Generovať.Generate.
Audio generované úspešne
0:00
Stiahnuť audio na stiahnutie Stiahnuť.srt súbor Platnosť odkazu vyprší za 24h
Bezplatná úroveň: osobné použitie. Komerčná licencia od 5 USD/mesiac
Láska TTS.ai? Povedzte svojim priateľom!

O nás VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Najlepšie pre: General-purpose text-to-speech with natural prosody

Prehľadávať všetky VITS hlasy

Na prvý pohľad

Vývojár
Jaehyeon Kim et al.
Licencia
MIT
Zvieratá
free
Rýchlosť
fast
Klonovanie hlasu
Nie
Jazyky
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Max. počet znakov
2000

VITS hlasy

CSS10 (Dutch)

Dutch
Voľný Neutral

CSS10 (Finnish)

Finnish
Voľný Neutral

CSS10 (French)

French
Voľný Neutral

CSS10 (German)

German
Voľný Neutral

CSS10 (Hungarian)

Hungarian
Voľný Neutral

CSS10 (Spanish)

Spanish
Voľný Neutral

Common Voice (Bulgarian)

Bulgarian
Voľný Neutral

Common Voice (Portuguese)

Portuguese
Voľný Neutral

Default

English
Voľný Neutral

MAI (Polish)

Polish
Voľný Female

MAI (Ukrainian)

Ukrainian
Voľný Neutral

VITS TTS — Často kladené otázky

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Všetky hlasy