Ming-Omni TTS

Ming-Omni TTS អត្ថបទ​ទៅ​សំឡេង

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

ចុះឈ្មោះ កំណត់​សម្រាប់​តួអក្សរ ៥, ០០០

រុំ​អត្ថបទ​របស់​អ្នក​ក្នុង​ស្លាក SSML សម្រាប់​ការ​ត្រួតពិនិត្យ​ជាក់លាក់ ៖

<speak><prosody rate="slow">Slow speech</prosody></speak>

ស្លាក​ដែល​ម៉ូដែល​ដែល​បាន​ជ្រើស​យល់​ — ចុច​ដើម្បី​ទម្លាក់​មួយ​ទៅ​ក្នុង​អត្ថបទ​របស់​អ្នក​នៅ​កន្លែង​ដែល​វា​កើតឡើង ៖

ម៉ូដែល​នេះ​អាន​អត្ថបទ​ធម្មតា ដូច្នេះ​ស្លាក​ក្នុង​បន្ទាត់​ត្រូវ​បាន​មិន​អើពើ ។ សម្រាប់​អារម្មណ៍​ដែល​មាន​មូលដ្ឋាន​លើ​ស្លាក ប្ដូរ​ទៅ​ម៉ូដែល​បង្ហាញ​ដូច​ជា Orpheus ឬ Bark ។

កំណត់​ការ​បញ្ចេញ​សំឡេង​ផ្ទាល់ខ្លួន (ពាក្យ = ការ​បញ្ចេញ​សំឡេង) ៖

-12 +12
0.5x 2.0x
ឥតគិតថ្លៃ​ជាមួយ Piper, VITS, MeloTTS
អូឌីយ៉ូ​ដែល​បាន​បង្កើត​របស់​អ្នក​នឹង​លេចឡើង​នៅ​ទីនេះ ។ ជ្រើស​ម៉ូដែល បញ្ចូល​អត្ថបទ ហើយ​ចុច បង្កើត ។
បាន​បង្កើត​អូឌីយ៉ូ​ដោយ​ជោគជ័យ
0:00
ទាញយក​អូឌីយ៉ូ ទាញយក.srt តំណផុតកំណត់ក្នុង 24h
កម្រិត​ឥត​គិត​ថ្លៃ ៖ ការ​ប្រើ​ផ្ទាល់​ខ្លួន ។ អាជ្ញាប័ណ្ណពាណិជ្ជកម្មពី $5/ខែ
ធ្វើ​ឲ្យ​នេះ​ជា​សំឡេង​របស់​អ្នក ក្លូន​សំឡេង​ក្នុង ៣០ វិនាទី
ស្រឡាញ់ TTS.ai? ប្រាប់មិត្តភក្តិរបស់អ្នក!

អំពី Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

ល្អបំផុត​សម្រាប់: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

រកមើល​ទាំងអស់ Ming-Omni TTS សំឡេង

ទិដ្ឋភាព​ទូទៅ

អ្នក​អភិវឌ្ឍន៍
inclusionAI
អាជ្ញាបណ្ណ
Apache 2.0
ផ្កាយ
free
ល្បឿន
medium
ការ​ក្លូន​សំឡេង
បាទ/ ចាស
ភាសា
English, Chinese
តួអក្សរ​អតិបរមា
1000

Ming-Omni TTS សំឡេង

Default

English
ឥត​គិត​ថ្លៃ Neutral

Default (Chinese)

Chinese
ឥត​គិត​ថ្លៃ Neutral

Ming-Omni TTS TTS - សំណួរ​ដែល​សួរ​ញឹកញាប់

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← សំឡេង​ទាំងអស់