MOSS-TTSD 音声翻訳
A 7B dialogue model that continues conversations from an audio prompt — up to five speakers and 60 minutes of coherent audio.
SSML タグでテキストを囲み、正確な制御を行う:
<speak><prosody rate="slow">Slow speech</prosody></speak>
選択したモデルが理解するタグ - クリックしてテキストにドラッグします:
このモデルは単純テキストを読み込み、インラインタグは無視されます。タグベースの感情を表現するには、Orpheus や Bark のような表現モデルに切り替えてください。
カスタム発音を定義 (単語=発音):
情報 MOSS-TTSD
MOSS-TTSD v1.0 from OpenMOSS is a 7-billion-parameter dialogue text-to-speech model that continues a conversation from a short audio prompt rather than reading isolated lines. It handles up to five simultaneous speakers via [S1]/[S2]-style tags, zero-shot voice cloning from 3-to-10-second references, and stretches of coherent multi-turn dialogue up to 60 minutes long. It is distinct from the OpenMOSS MOSS-TTS model — the TTSD variant is specialized for podcast, audiobook, and dubbing workflows where long, consistent conversational audio is the goal. Released under Apache 2.0, it needs around 12GB of VRAM given its size.
適合する: Podcasts, audiobooks, dubbed dialogue, conversational content with multiple voices
すべてブラウズ MOSS-TTSD 声概要
- 開発者
- OpenMOSS
- ライセンス
- Apache 2.0
- 動物
- standard
- スピード
- medium
- 声のクローン
- はい
- 言語
- English, Chinese
- 最大文字数
- 5000