Ming-Omni TTS TTS
A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.
Wrap ou tèks nan SSML tags pou presizyon kontwòl:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags ke modèl la chwazi konprann — klike pou mete yon nan tèks ou kote li rive:
Modèl sa a li tèks senp, se poutèt sa atik ki nan liy yo pa pran an kont. Pou efè ki baze sou atik, chanje pou yon modèl ekspresyon tankou Orpheus oswa Bark.
Define prononciations Custom (mot = prononciation):
Atik Ming-Omni TTS
Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.
Pi bon pou: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content
Navigue tout Ming-Omni TTS VoyYon ti gade
- Pwogramè
- inclusionAI
- Lisans
- Apache 2.0
- Nivo
- free
- Vitès
- medium
- Klonaj vwa
- Wi
- Lang
- English, Chinese
- Karakteris maksimòm
- 1000