Ming-Omni TTS

Ming-Omni TTS ടിടിഎസ്

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

ഒപ്പ് വയ്ക്ക്. 5,000 ക്യാരക്ടര്‍ പരിധിയ്ക്കു്

കൃത്യമായ നിയന്ത്രണത്തിനായി SSML തൊങ്ങലില്‍ വാചകം പൊതിയുക:

<speak><prosody rate="slow">Slow speech</prosody></speak>

തെരഞ്ഞെടുത്ത മാതൃക മനസ്സിലാക്കുന്നത് ടാഗ് (കുടികള്‍) :

ഈ മോഡ് സാധാരണ പദാവലി വായിക്കുന്നു, അതുകൊണ്ട് ഇന്‍ലൈന്‍ തൊങ്ങല്‍ അവഗണിപ്പിക്കുന്നു. ടാഗ് അടിസ്ഥാനപരമായ വികാരങ്ങള്‍ക്കു് ഓര്‍ഫിയസ് അല്ലെങ്കില്‍ ബാര്‍ക് പോലുള്ള ഒരു ചിത്രീകരണ മോഡില്‍ മാറുക.

ഇഷ്ടപ്പെട്ട ഉച്ചാരണം നിര്‍വ്വചിക്കുക (വാക്ക് = ഉച്ചാരണം):

-12 +12
0.5x 2.0x
പൈപ്പര്‍, വി.ടി.സ്, മെലോട്ടിക്സ്
നിങ്ങള്‍ ഉണ്ടാക്കിയ ഓഡിയോ ഇവിടെ പ്രത്യക്ഷപ്പെടും. ഒരു മാതൃക തെരഞ്ഞെടുക്കുക, പദാവലി നല്‍കുക, നിര്‍മ്മിക്കുക എന്നിവ നിര്‍മ്മിക്കുക.
വിജയകരമായി ഉണ്ടാക്കിയ ശബ്ദങ്ങള്‍
0:00
ഓഡിയോ ഡൌണ്‍ലോഡ് ചെയ്യുക ഡൌണ്‍ലോട് ചെയ്യുക ലിങ്ക് 24hല്‍ അവസാനിച്ചിരിക്കുന്നു
സ്വതന്ത്രമായ അക്ഷരങ്ങളുടെ താഴേയ്ക്ക് പ്രവര്‍ത്തിപ്പിയ്ക്കുന്നു ഓരോ മാസവും 200K അക്ഷരങ്ങള്‍ ലഭ്യമാക്കുക — 550/mo അല്ലെങ്കില്‍ ഒരു സമയം 100K പാക്ക് 5000- നു്
ഇത് നിന്റെ ശബ്ദം തന്നെ ആക്കൂ. ശബ്ദം 30 സെക്കന്‍റില്‍ ക്ലിയര്‍ ചെയ്യുന്നു
ടിടിഎസ് സ്‌നേഹിക്കുന്നു, കൂട്ടുകാരോട് പറയൂ!

സംബന്ധിച്ച് Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

അതിനു വേണ്ടിയുള്ള ഏറ്റവും നല്ല സ്ഥലം.: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

എല്ലാം പരതുക Ming-Omni TTS ശബ്ദങ്ങള്‍

ഒരു നോക്കുമ്പോള്‍

രചയിതാവു്
inclusionAI
അനുമതി
Apache 2.0
ടിയെര്‍
free
വേഗത
medium
ശബ്ദമിശ്രണോപാധി
അതെ
ഭാഷകള്‍
English, Chinese
ഏറ്റവും കൂടിയ ക്യാരക്ടറുകള്‍
1000

Ming-Omni TTS ശബ്ദങ്ങള്‍

Default

English
ഫ്രീ Neutral

Default (Chinese)

Chinese
ഫ്രീ Neutral

Ming-Omni TTS ടിടിഎസ്‌ —⁠ എഫ്‌എക്‌സി

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← എല്ലാ ശബ്ദങ്ങളും