MOSS-TTSD Mga TNT
A 7B dialogue model that continues conversations from an audio prompt — up to five speakers and 60 minutes of coherent audio.
I-wrap ang iyong teksto sa SSML tags para sa tumpak na kontrol:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags ang napili modelo nauunawaan — i-click upang ihulog ang isa sa iyong teksto kung saan ito ay nangyayari:
Ang modelong ito ay nagbabasa ng karaniwang teksto, kaya inline tags ay hindi pinapansin. Para sa tag-based na damdamin, lumipat sa isang makahulugang modelo tulad ng Orpheus o Bark.
Tukuyin ang mga pasadyang mga panlapi (word = panlapi):
Tungkol sa MOSS-TTSD
MOSS-TTSD v1.0 from OpenMOSS is a 7-billion-parameter dialogue text-to-speech model that continues a conversation from a short audio prompt rather than reading isolated lines. It handles up to five simultaneous speakers via [S1]/[S2]-style tags, zero-shot voice cloning from 3-to-10-second references, and stretches of coherent multi-turn dialogue up to 60 minutes long. It is distinct from the OpenMOSS MOSS-TTS model — the TTSD variant is specialized for podcast, audiobook, and dubbing workflows where long, consistent conversational audio is the goal. Released under Apache 2.0, it needs around 12GB of VRAM given its size.
Pinakamahusay para sa: Podcasts, audiobooks, dubbed dialogue, conversational content with multiple voices
Mag-browse ng lahat MOSS-TTSD Mga bosesSa isang sulyap
- Developer
- OpenMOSS
- Lisensya
- Apache 2.0
- Mga hayop
- standard
- Bilis
- medium
- Pag-clone ng boses
- Oo
- Wika
- English, Chinese
- Max character
- 5000