Ming-Omni TTS

Ming-Omni TTS TTS

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

Đăng ký giới hạn 5000 ký tự

Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:

Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.

Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):

-12 +12
0.5x 2.0x
Miễn phí với Piper, VITS, MeloTTS
Âm thanh đã tạo sẽ xuất hiện ở đây. Chọn một mô hình, nhập văn bản, và nhấn vào Tạo.
Âm thanh đã được tạo thành công
0:00
Tải về âm thanh Tải về Liên kết hết hạn trong 24h
Free tier: sử dụng cá nhân. Giấy phép thương mại từ $5/mo
Cảm ơn bạn đã tin tưởng TTS.ai!

Về Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

Tốt nhất cho: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

& Xem tất cả Ming-Omni TTS giọng nói

Một cái nhìn

Nhà phát triển
inclusionAI
Giấy phép
Apache 2.0
Thú
free
Tốc độ
medium
Ký âm
Ngôn ngữ
English, Chinese
Tối đa các ký tự
1000

Ming-Omni TTS giọng nói

Default

English
Tự do Neutral

Default (Chinese)

Chinese
Tự do Neutral

Ming-Omni TTS TTS - FAQ

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← Tất cả giọng nói