VITS

VITS TTS

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

ثبت نام برای حد ۵۰۰۰ کاراکتر

برای کنترل دقیق ، متن خود را در برچسبهای SSML بپیچید:

<speak><prosody rate="slow">Slow speech</prosody></speak>

برچسبهایی که مدل برگزیده می‌فهمد — برای انداختن یکی در متن خود ، جایی که اتفاق می‌افتد ، کلیک کنید:

این مدل متن ساده را می‌خواند ، بنابراین برچسب‌های خطی نادیده گرفته می‌شوند. برای احساسات مبتنی بر برچسب ، به یک مدل بیانی مانند Orpheus یا Bark تغییر دهید.

تعریف تلفظ سفارشی) کلمه = تلفظ (:

-12 +12
0.5x 2.0x
آزاد با Piper, VITS, MeloTTS
صدای تولید شده شما در اینجا ظاهر خواهد شد. یک مدل را انتخاب کنید ، متن را وارد کنید ، و تولید را فشار دهید.
صدا با موفقیت تولید شد
0:00
بارگیری صدا دانلود پیوند در ۲۴ ساعت پایان می‌یابد
1- استفاده شخصی: استفاده شخصی. مجوز تجاری از $5/mo
دوست داريد TTS.ai؟ به دوستانتون بگو!

در مورد VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

بهترین برای: General-purpose text-to-speech with natural prosody

مرور همۀ VITS صداها

يه نگاهي بنداز

توسعه‌دهنده
Jaehyeon Kim et al.
مجوز
MIT
حیوان
free
سرعت
fast
شبیه‌سازی صدا
نه
زبانها
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
بیشینه نویسه‌ها
2000

VITS صداها

CSS10 (Dutch)

Dutch
آزاد Neutral

CSS10 (Finnish)

Finnish
آزاد Neutral

CSS10 (French)

French
آزاد Neutral

CSS10 (German)

German
آزاد Neutral

CSS10 (Hungarian)

Hungarian
آزاد Neutral

CSS10 (Spanish)

Spanish
آزاد Neutral

Common Voice (Bulgarian)

Bulgarian
آزاد Neutral

Common Voice (Portuguese)

Portuguese
آزاد Neutral

Default

English
آزاد Neutral

MAI (Polish)

Polish
آزاد Female

MAI (Ukrainian)

Ukrainian
آزاد Neutral

VITS FAQ - پرسش و پاسخ

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← همه صداها