VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Signa per 5000 caràcters límit

Ajusta el text a les etiquetes SSML per al control precís:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetes del model seleccionat entenen el clic show clic per a deixar- ne un al text a on succeeix:

Aquest model llegeix text pla, així que les etiquetes inserides s' ignoren. Per a emocions basades en etiquetes, canvieu a un model expressiu com Orfeus o Bark.

Defineix pronúncies personalitzades (word = pronunciació):

-12 +12
0.5x 2.0x
Lliure amb Pipista, VITS, MeloTTS
Aquí apareixerà el vostre àudio generat. Escolliu un model, introduïu text i cliqueu Genera.
L' àudio s' ha generat correctament
0:00
Descarrega àudio Descarrega.srt L' enllaç expirarà el 24h
Correlitzador lliure: ús personal. Llicència de venda de 5/mo
Fes que això sigui la teva pròpia veu Clona una veu en 30 segons
Els teus amics!

Quant a VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Millor per: General-purpose text-to-speech with natural prosody

Navega- ho tot VITS veus

En una mirada

Desenvolupador
Jaehyeon Kim et al.
Llicència
MIT
TierCity name (optional, probably does not need a translation)
free
Velocitat
fast
clonació de veu
No
Idiomes
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Nombre màxim de caràcters
2000

VITS veus

CSS10 (Dutch)

Dutch
Lliure Neutral

CSS10 (Finnish)

Finnish
Lliure Neutral

CSS10 (French)

French
Lliure Neutral

CSS10 (German)

German
Lliure Neutral

CSS10 (Hungarian)

Hungarian
Lliure Neutral

CSS10 (Spanish)

Spanish
Lliure Neutral

Common Voice (Bulgarian)

Bulgarian
Lliure Neutral

Common Voice (Portuguese)

Portuguese
Lliure Neutral

Default

English
Lliure Neutral

MAI (Polish)

Polish
Lliure Female

MAI (Ukrainian)

Ukrainian
Lliure Neutral

VITS PMF TTS

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Totes les veus