MOSS-TTSD

MOSS-TTSD የድምፅ ፋይል

A 7B dialogue model that continues conversations from an audio prompt — up to five speakers and 60 minutes of coherent audio.

ምዝገባ ፊደል(ሎች)

ርዕሱን በSSML መለያዎች ውስጥ ለጥሩ ቁጥጥር ይዞሩት:

<speak><prosody rate="slow">Slow speech</prosody></speak>

የተመረጠው ሞዴል የሚያውቃቸው መለያዎች - በጽሑፍዎ ውስጥ የሚከሰትበትን ቦታ ለመውሰድ ጠቅ ያድርጉ፦

ይህ ሞዴል ቀላል ጽሑፍን ያነባል፣ ስለዚህም በመስመር ውስጥ ያሉ ምልክቶች ይዘገያሉ፡፡ ለታክስ-ተኮር ስሜት እንደ ኦርፊየስ ወይም ባርክ ያሉ ግልጽ ሞዴሎችን ይለውጡ።

የራሱን ተናጋሪ ግለጽ (ቃል = ተናጋሪ):

-12 +12
0.5x 2.0x
ነጻ ከፒፐር, VITS, MeloTTS ጋር
የእርስዎ የተፈጠረ ድምፅ እዚህ ይታይ. ሞዴል ይምረጡ፣ ጽሑፍ ያስገቡ፣ እና ይፈጥሩ ላይ ጠቅ ያድርጉ
ድምፅ በፍጥነት ተፈጠረ
0:00
ድምፅ ያውርዱ ያውርዱ መገናኛ በ24 ሰዓት ውስጥ ይቋረጣል
ነጻ ደረጃ: የግል ጥቅም የኮሜርሲ ውል ከ $5/mo
TTS.aiን ወዳጅነት?

ስለ MOSS-TTSD

MOSS-TTSD v1.0 from OpenMOSS is a 7-billion-parameter dialogue text-to-speech model that continues a conversation from a short audio prompt rather than reading isolated lines. It handles up to five simultaneous speakers via [S1]/[S2]-style tags, zero-shot voice cloning from 3-to-10-second references, and stretches of coherent multi-turn dialogue up to 60 minutes long. It is distinct from the OpenMOSS MOSS-TTS model — the TTSD variant is specialized for podcast, audiobook, and dubbing workflows where long, consistent conversational audio is the goal. Released under Apache 2.0, it needs around 12GB of VRAM given its size.

ምርጥ ለ: Podcasts, audiobooks, dubbed dialogue, conversational content with multiple voices

ሁሉንም አጥፉ MOSS-TTSD ድምጾች

በጥቂቱ

የድር አዘጋጅ
OpenMOSS
ፈቃድ
Apache 2.0
ዐምድ
standard
ፍጥነት
medium
የድምፅ ቅጂ
አዎ
ቋንቋዎች
English, Chinese
ፊደላት
5000

MOSS-TTSD ድምጾች

Default (Chinese)

Chinese
መደበኛ Neutral

Default Speaker

English
መደበኛ Neutral

MOSS-TTSD የትርጉም መሳሪያ

Up to five simultaneous speakers, addressed via speaker tags like [S1] and [S2], with the ability to clone each voice from a short reference clip.

It can produce up to 60 minutes of coherent multi-turn dialogue, which is what makes it suited to full podcast episodes and audiobook chapters rather than short clips.

MOSS-TTSD is a dialogue-specialized variant that continues conversations from an audio prompt and targets podcast, audiobook, and dubbing workflows, whereas the base MOSS-TTS is a general single-voice synthesis model.
← ሁሉንም ድምጾች