VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Inscrever-se para o limite de 5000 caracteres

Envolva o seu texto em tags SSML para controle preciso:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas o modelo selecionado entende — clique para soltar um para o seu texto onde acontece:

Este modelo lê texto simples, por isso as etiquetas inline são ignoradas. Para emoção baseada em tags, mude para um modelo expressivo como Orpheus ou Bark.

Definir pronúncias personalizadas (palavra = pronúncia):

-12 +12
0.5x 2.0x
Grátis com Piper, VITS, MeloTTS
Seu áudio gerado aparecerá aqui. Escolha um modelo, introduza texto e clique em Gerar.
O áudio gerado com sucesso
0:00
Baixe áudio Baixar.srt A ligação expira em 24h
Gratuito nível: uso pessoal. Licença comercial a partir de $5/mo
Gosta do TTS.ai? Conte aos seus amigos!

Sobre VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Melhor para: General-purpose text-to-speech with natural prosody

Procurar todos VITS vozes

De uma olhada

Desenvolvedor
Jaehyeon Kim et al.
Licença
MIT
Tier
free
Velocidade
fast
Clonagem de voz
Não
Línguas
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Número máximo de caracteres
2000

VITS vozes

CSS10 (Dutch)

Dutch
Grátis Neutral

CSS10 (Finnish)

Finnish
Grátis Neutral

CSS10 (French)

French
Grátis Neutral

CSS10 (German)

German
Grátis Neutral

CSS10 (Hungarian)

Hungarian
Grátis Neutral

CSS10 (Spanish)

Spanish
Grátis Neutral

Common Voice (Bulgarian)

Bulgarian
Grátis Neutral

Common Voice (Portuguese)

Portuguese
Grátis Neutral

Default

English
Grátis Neutral

MAI (Polish)

Polish
Grátis Female

MAI (Ukrainian)

Ukrainian
Grátis Neutral

VITS TTS — FAQ

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Todas as vozes