Ming-Omni TTS

Ming-Omni TTS TTS

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

Enskri Limit pou 5,000 karaktè

Wrap ou tèks nan SSML tags pou presizyon kontwòl:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags ke modèl la chwazi konprann — klike pou mete yon nan tèks ou kote li rive:

Modèl sa a li tèks senp, se poutèt sa atik ki nan liy yo pa pran an kont. Pou efè ki baze sou atik, chanje pou yon modèl ekspresyon tankou Orpheus oswa Bark.

Define prononciations Custom (mot = prononciation):

-12 +12
0.5x 2.0x
Gratis ak Piper, VITS, MeloTTS
Son ou kreye a ap parèt isit la. Chwazi yon modèl, antre tèks la, epi klike Kreye.
Audio Generated Successfully
0:00
Telechaje son Telechaje.srt Link expires in 24h
Free tier: itilize pèsonèl. Lisans Komèsyal soti nan $5/mo
Love TTS.ai? Di zanmi ou yo!

Atik Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

Pi bon pou: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

Navigue tout Ming-Omni TTS Voy

Yon ti gade

Pwogramè
inclusionAI
Lisans
Apache 2.0
Nivo
free
Vitès
medium
Klonaj vwa
Wi
Lang
English, Chinese
Karakteris maksimòm
1000

Ming-Omni TTS Voy

Default

English
Gratis Neutral

Default (Chinese)

Chinese
Gratis Neutral

Ming-Omni TTS TTS — FAQ

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← Tout vwa