MOSS-TTSD

MOSS-TTSD 音声翻訳

A 7B dialogue model that continues conversations from an audio prompt — up to five speakers and 60 minutes of coherent audio.

登録 5000文字の制限を設けました

SSML タグでテキストを囲み、正確な制御を行う:

<speak><prosody rate="slow">Slow speech</prosody></speak>

選択したモデルが理解するタグ - クリックしてテキストにドラッグします:

このモデルは単純テキストを読み込み、インラインタグは無視されます。タグベースの感情を表現するには、Orpheus や Bark のような表現モデルに切り替えてください。

カスタム発音を定義 (単語=発音):

-12 +12
0.5x 2.0x
ピパー、VITS、MeloTTS をフリーで使用
生成したオーディオがここに表示されます。モデルを選択し、テキストを入力して、生成をクリックします。
オーディオを作成しましたName
0:00
音声をダウンロード ダウンロード リンクは24時間で失効します
無料階級:個人用。 商用ライセンス $5/月から
これを自分の声にしよう 30秒で声をクローン
TTS.aiが気に入りましたか?友達に教えてあげましょう!

情報 MOSS-TTSD

MOSS-TTSD v1.0 from OpenMOSS is a 7-billion-parameter dialogue text-to-speech model that continues a conversation from a short audio prompt rather than reading isolated lines. It handles up to five simultaneous speakers via [S1]/[S2]-style tags, zero-shot voice cloning from 3-to-10-second references, and stretches of coherent multi-turn dialogue up to 60 minutes long. It is distinct from the OpenMOSS MOSS-TTS model — the TTSD variant is specialized for podcast, audiobook, and dubbing workflows where long, consistent conversational audio is the goal. Released under Apache 2.0, it needs around 12GB of VRAM given its size.

適合する: Podcasts, audiobooks, dubbed dialogue, conversational content with multiple voices

すべてブラウズ MOSS-TTSD 声

概要

開発者
OpenMOSS
ライセンス
Apache 2.0
動物
standard
スピード
medium
声のクローン
はい
言語
English, Chinese
最大文字数
5000

MOSS-TTSD 声

Default (Chinese)

Chinese
標準 Neutral

Default Speaker

English
標準 Neutral

MOSS-TTSD よくある質問

Up to five simultaneous speakers, addressed via speaker tags like [S1] and [S2], with the ability to clone each voice from a short reference clip.

It can produce up to 60 minutes of coherent multi-turn dialogue, which is what makes it suited to full podcast episodes and audiobook chapters rather than short clips.

MOSS-TTSD is a dialogue-specialized variant that continues conversations from an audio prompt and targets podcast, audiobook, and dubbing workflows, whereas the base MOSS-TTS is a general single-voice synthesis model.
← すべての声