Ming-Omni TTS TTS
A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về Ming-Omni TTS
Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.
Tốt nhất cho: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content
& Xem tất cả Ming-Omni TTS giọng nóiMột cái nhìn
- Nhà phát triển
- inclusionAI
- Giấy phép
- Apache 2.0
- Thú
- free
- Tốc độ
- medium
- Ký âm
- Có
- Ngôn ngữ
- English, Chinese
- Tối đa các ký tự
- 1000