Ming-Omni TTS TTS
A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.
Întoarceți textul în etichetele SSML pentru un control precis:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etichetele modelului selectat înțeleg — click pentru a lăsa unul în textul tău unde se întâmplă:
Acest model citește textul simplu, astfel încât etichetele inline sunt ignorate. Pentru emoții bazate pe tag, schimbați la un model expresiv cum ar fi Orpheus sau Bark.
Definiți pronunțiare personalizată (cuvânt = pronunție):
Despre Ming-Omni TTS
Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.
Cel mai bun pentru: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content
Navigați toate Ming-Omni TTS vociLa o privire
- Dezvoltator
- inclusionAI
- Licență
- Apache 2.0
- Nivel
- free
- Viteză
- medium
- Clonarea vocală
- Da.
- Limbi
- English, Chinese
- Caractere maxime
- 1000