VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

Aanmelden voor 5.000 tekenlimiet

Wrap uw tekst in SSML-tags voor nauwkeurige controle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags het geselecteerde model begrijpt

Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.

Definieer aangepaste uitspraaken (woord = uitspraak):

-12 +12
0.5x 2.0x
Gratis met Piper, VITS, MeloTTS
Uw gegenereerde audio zal hier verschijnen. Kies een model, voer tekst in en klik op Genereren.
Audio Generated Succesvol
0:00
Audio downloaden Download.srt Link verloopt in 24 uur
Gratis niveau: persoonlijk gebruik. Commerciële licentie van $5/mo
Hou van TTS.ai? Vertel het je vrienden!

Info VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

Beste voor: General-purpose text-to-speech with natural prosody

Alles doorbladeren VITS stemmen

In een oogopslag

Ontwikkelaar
Jaehyeon Kim et al.
Licentie
MIT
Niveau
free
Snelheid
fast
Klonen van stemmen
Nee
Talen
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
Max. tekens
2000

VITS stemmen

CSS10 (Dutch)

Dutch
Vrij Neutral

CSS10 (Finnish)

Finnish
Vrij Neutral

CSS10 (French)

French
Vrij Neutral

CSS10 (German)

German
Vrij Neutral

CSS10 (Hungarian)

Hungarian
Vrij Neutral

CSS10 (Spanish)

Spanish
Vrij Neutral

Common Voice (Bulgarian)

Bulgarian
Vrij Neutral

Common Voice (Portuguese)

Portuguese
Vrij Neutral

Default

English
Vrij Neutral

MAI (Polish)

Polish
Vrij Female

MAI (Ukrainian)

Ukrainian
Vrij Neutral

VITS Veelgestelde vragen

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← Alle stemmen