Ming-Omni TTS

Ming-Omni TTS ສຽງ​ເປັນ​ຂໍ້ຄວາມ

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

ລົງທະບຽນ ຈໍາກັດ​ຕົວ​ອັກສອນ​ໃຫ້​ໄດ້ 5,000

ວາງ​ຂໍ້ຄວາມ​ຂອງທ່ານ​ໄວ້​ໃນ​ແທັກ SSML ເພື່ອ​ຄວບຄຸມ​ຢ່າງ​ລະອຽດ:

<speak><prosody rate="slow">Slow speech</prosody></speak>

ແທັກ​ທີ່​ຕົວແບບ​ທີ່​ໄດ້​ເລືອກ​ເຂົ້າໃຈ — ກົດ​ເພື່ອ​ປ່ອຍ​ພວກ​ມັນ​ລົງ​ໃນ​ຂໍ້ຄວາມ​ຂອງທ່ານ​ບ່ອນ​ທີ່​ມັນ​ເກີດຂຶ້ນ:

ແບບ​ນີ້​ອ່ານ​ຂໍ້ຄວາມ​ປົກກະຕິ, ສະນັ້ນ​ແທັກ​ໃນ​ແຖບ​ຈະ​ບໍ່​ຖືກ​ລະບຸ​ໄວ້. ສຳ​ລັບ​ຄວາມ​ຮູ້ສຶກ​ທີ່​ອີງ​ໃສ່​ແທັກ, ປ່ຽນ​ໄປ​ຫາ​ແບບ​ທີ່​ສະແດງ​ອອກ​ຄື Orpheus ຫຼື Bark.

ຕັ້ງຄ່າ​ການ​ອອກສຽງ​ແບບ​ສ່ວນ​ຕົວ (ຄໍາ = ການອອກສຽງ):

-12 +12
0.5x 2.0x
ຟຣີ​ກັບ Piper, VITS, MeloTTS
ສຽງ​ທີ່​ໄດ້​ສ້າງ​ຂຶ້ນ​ຂອງ​ທ່ານ​ຈະ​ປາກົດ​ຢູ່​ທີ່​ນີ້. ເລືອກ​ແບບ, ເຂົ້າ​ເຖິງ​ຂໍ້ຄວາມ ແລະ ຄລິກ​ໃສ່ ສ້າງ.
ສ້າງ​ສຽງ​ໄດ້​ຢ່າງ​ສຳເລັດ​ຜົນ
0:00
ດາວໂຫລດ​ສຽງ ດາວໂຫລດ.srt ການ​ເຊື່ອມຕໍ່​ຈະ​ໝົດ​ອາຍຸ​ໃນ 24 ຊົ່ວໂມງ
ລະດັບຟຣີ: ການໃຊ້ສ່ວນຕົວ. ໃບອະນຸຍາດການຄ້າຈາກ $5/ເດືອນ
ກຳລັງໃຊ້​ງານ​ດົນ​ເກີນ​ໄປ​ໃນ​ຕົວອັກສອນ​ທີ່​ມີ ໄດ້ຮັບຕົວອັກສອນ 200K ທຸກໆເດືອນ - $5/ເດືອນ ຫຼື 100K ຄັ້ງດຽວສໍາລັບ $5
ສ້າງສຽງ​ຂອງ​ທ່ານ​ເອງ ສ້າງ​ສຽງ​ແບບ​ຄລາສສິກ​ໃນ​ 30 ວິນາທີ
ຮັກ TTS.ai? ເວົ້າກັບເພື່ອນຂອງທ່ານ!

ກ່ຽວ​ກັບ Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

ດີທີ່ສຸດ ສຳ ລັບ: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

ຄົ້ນຫາ​ທັງໝົດ Ming-Omni TTS ສຽງ

ເບິ່ງ​ຢ່າງ​ໄວວາ

ຜູ້​ພັດທະນາ
inclusionAI
ໃບອະນຸຍາດ
Apache 2.0
ສັດ
free
ໄວ
medium
ການປິດ​ສຽງ
​ແມ່ນ
ພາສາ
English, Chinese
ຕົວອັກສອນ​ສູງສຸດ
1000

Ming-Omni TTS ສຽງ

Default

English
ບໍ່ມີ Neutral

Default (Chinese)

Chinese
ບໍ່ມີ Neutral

Ming-Omni TTS TTS - ຄໍາຖາມ​ທີ່​ຖາມ​ເລື້ອຍໆ

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← ສຽງ​ທັງ​ໝົດ