IndexTTS-2

IndexTTS-2 音声翻訳

A zero-shot TTS model with fine-grained emotion control via emotion vectors, no emotion-specific training data required.

登録 5000文字の制限を設けました

SSML タグでテキストを囲み、正確な制御を行う:

<speak><prosody rate="slow">Slow speech</prosody></speak>

選択したモデルが理解するタグ - クリックしてテキストにドラッグします:

このモデルは単純テキストを読み込み、インラインタグは無視されます。タグベースの感情を表現するには、Orpheus や Bark のような表現モデルに切り替えてください。

カスタム発音を定義 (単語=発音):

-12 +12
0.5x 2.0x
ピパー、VITS、MeloTTS をフリーで使用
生成したオーディオがここに表示されます。モデルを選択し、テキストを入力して、生成をクリックします。
オーディオを作成しましたName
0:00
音声をダウンロード ダウンロード リンクは24時間で失効します
無料階級:個人用。 商用ライセンス $5/月から
これを自分の声にしよう 30秒で声をクローン
TTS.aiが気に入りましたか?友達に教えてあげましょう!

情報 IndexTTS-2

IndexTTS-2, from the Index Team, is an expressive text-to-speech system that pairs zero-shot voice synthesis with precise emotional control. Rather than relying on emotion-labeled training data, it uses emotion vectors to dial in tones like happy, sad, angry, or fearful independently of the voice itself. Built on a Qwen2 backbone with BigVGAN as the vocoder, it supports English and Chinese and can clone a voice from roughly five seconds of reference audio. It suits audiobooks, virtual assistants, and any content where the same voice needs to shift emotional register. Its weights use the Bilibili Model License, which permits commercial use below large usage and revenue thresholds.

適合する: Emotionally expressive content, audiobooks, virtual assistants

すべてブラウズ IndexTTS-2 声

概要

開発者
Index Team
ライセンス
Bilibili Model License
動物
standard
スピード
medium
声のクローン
はい
言語
English, Chinese
最大文字数
1000

IndexTTS-2 声

Chinese Default

Chinese
標準 Neutral

Default

English
標準 Neutral

IndexTTS-2 よくある質問

It uses emotion vectors that let you specify tones such as happy, sad, angry, or fearful without needing emotion-specific training data, and the emotional expression is controlled independently from the voice identity.

Yes. It performs zero-shot voice cloning from a short reference, typically around five seconds of audio, in English or Chinese.

Its weights are released under the Bilibili Model License, which allows commercial use for products below defined user and revenue thresholds. Larger deployments should review the license terms.
← すべての声