VITS

VITS 音声翻訳

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

登録 5000文字の制限を設けました

SSML タグでテキストを囲み、正確な制御を行う:

<speak><prosody rate="slow">Slow speech</prosody></speak>

選択したモデルが理解するタグ - クリックしてテキストにドラッグします:

このモデルは単純テキストを読み込み、インラインタグは無視されます。タグベースの感情を表現するには、Orpheus や Bark のような表現モデルに切り替えてください。

カスタム発音を定義 (単語=発音):

-12 +12
0.5x 2.0x
ピパー、VITS、MeloTTS をフリーで使用
生成したオーディオがここに表示されます。モデルを選択し、テキストを入力して、生成をクリックします。
オーディオを作成しましたName
0:00
音声をダウンロード ダウンロード リンクは24時間で失効します
無料階級:個人用。 商用ライセンス $5/月から
これを自分の声にしよう 30秒で声をクローン
TTS.aiが気に入りましたか?友達に教えてあげましょう!

情報 VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

適合する: General-purpose text-to-speech with natural prosody

すべてブラウズ VITS 声

概要

開発者
Jaehyeon Kim et al.
ライセンス
MIT
動物
free
スピード
fast
声のクローン
いや
言語
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
最大文字数
2000

VITS 声

CSS10 (Dutch)

Dutch
自由 Neutral

CSS10 (Finnish)

Finnish
自由 Neutral

CSS10 (French)

French
自由 Neutral

CSS10 (German)

German
自由 Neutral

CSS10 (Hungarian)

Hungarian
自由 Neutral

CSS10 (Spanish)

Spanish
自由 Neutral

Common Voice (Bulgarian)

Bulgarian
自由 Neutral

Common Voice (Portuguese)

Portuguese
自由 Neutral

Default

English
自由 Neutral

MAI (Polish)

Polish
自由 Female

MAI (Ukrainian)

Ukrainian
自由 Neutral

VITS よくある質問

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← すべての声