Ming-Omni TTS

Ming-Omni TTS TTS

A compact 0.5B omni-modal speech model with near-CD-quality 44.1kHz output and zero-shot voice cloning.

Aanmelden voor 5.000 tekenlimiet

Wrap uw tekst in SSML-tags voor nauwkeurige controle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags het geselecteerde model begrijpt

Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.

Definieer aangepaste uitspraaken (woord = uitspraak):

-12 +12
0.5x 2.0x
Gratis met Piper, VITS, MeloTTS
Uw gegenereerde audio zal hier verschijnen. Kies een model, voer tekst in en klik op Genereren.
Audio Generated Succesvol
0:00
Audio downloaden Download.srt Link verloopt in 24 uur
Gratis niveau: persoonlijk gebruik. Commerciële licentie van $5/mo
Hou van TTS.ai? Vertel het je vrienden!

Info Ming-Omni TTS

Ming-omni-tts-0.5B by inclusionAI is a compact omni-modal speech model built on the BailingMM dense backbone with a patch-by-patch flow-matching audio decoder. Despite its small 500M-parameter size, it outputs 44.1kHz audio approaching CD quality and supports zero-shot voice cloning from a reference of three seconds or more. It includes built-in emotion, dialect, and even background-music control driven by JSON instructions, and is notably stable — reporting a 0.83% word error rate on Chinese benchmarks. With Apache 2.0 licensing and modest 3GB VRAM needs, it fits high-fidelity bilingual narration, emotion-controlled voice acting, and Chinese audiobook production.

Beste voor: High-fidelity bilingual narration, emotion-controlled voice acting, Chinese audiobook content

Alles doorbladeren Ming-Omni TTS stemmen

In een oogopslag

Ontwikkelaar
inclusionAI
Licentie
Apache 2.0
Niveau
free
Snelheid
medium
Klonen van stemmen
Ja.
Talen
English, Chinese
Max. tekens
1000

Ming-Omni TTS stemmen

Default

English
Vrij Neutral

Default (Chinese)

Chinese
Vrij Neutral

Ming-Omni TTS Veelgestelde vragen

It outputs 44.1kHz audio, close to CD quality — high for a model of only 0.5B parameters — thanks to its patch-by-patch flow-matching audio decoder.

Beyond voice cloning, it supports emotion, dialect, and background-music control via JSON instructions, and it is very stable, reporting a 0.83% word error rate on Chinese benchmarks.

English and Chinese, with zero-shot voice cloning from a reference clip of three seconds or longer.
← Alle stemmen