VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Ṣẹ̀dà fun àwọn àmì-àṣírí 5,000

Fi àkọlé rẹ pamọ́ sí àwọn àmì-ìwé SSML fún ìdáràn:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Àwọn Àmì-ìwé tí àwọn ìṣàmúlò-ètò tí a yàn gbọ́ - tẹ̀ láti fi ọkan sínú àkọ́lé rẹ̀ nínú àwọn ààyè-iṣẹ́ tí o bá jẹ́:

Àwọn àwọn àkọlé àwọn ààyè-iṣẹ́ àwọn àwọn àmì-ìwé àwọn àmì-ìwé àwọn àwọn àmì-ìwé àwọn à

Àwọn àwọn ìṣàfarawé àwọn àwọn ìṣàfarawé àwọn (ọrọ = ìṣàfàlì):

-12 +12
0.5x 2.0x
Free pẹlu Piper, VITS, MeloTTS
Àwọn àwòrán tí o ti ṣẹ̀dà tí o bá han níbẹ̀. Yan àwọn àwòrán, tẹ̀lẹ̀ àkọlé, ki o si tẹ̀ Ṣẹ̀dà.
Àwọn àwọn àwòrán tí a ṣẹ̀dà
0:00
Ṣàfikún Àwọn Àmì-ìwé Ṣàfikún.srt Líǹkì náà kù nínú 24h
Ìjádé ọ̀fẹ́: ìlòjútó ara ẹni. Lisensi Iṣowo ori lati $5/mo
O fẹ́ TTS.ai? Fì sọ̀kalẹ̀ fún àwọn ọrẹ̀ rẹ̀!

Ààyè-iṣẹ́ VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Tí o dara jù fún: High-fidelity audio, audiobooks, long-form content with voice consistency

Wá Gbogbo àwòrán VoxCPM Àwọn àwòrán

Nínú àwọn ìṣàfarawé

Àwọn Àkọlé
OpenBMB
Àwọn Ààyè-iṣẹ́
Apache 2.0
Àwọn àwọn ààyè-iṣẹ́
standard
Ìjánu-ìṣàmúlò-ètò
fast
Ìṣàfarawé àwọn àmì-ìwé
Yà
Àwọn
English, Chinese
Àwọn àyọkà ìpele
2000

VoxCPM Àwọn àwòrán

Default

English
Àwọn ìpéwọ̀n Neutral

Default Chinese

Chinese
Àwọn ìpéwọ̀n Neutral

VoxCPM Àwọn Àtòjọ-ẹ̀yàn

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Gbogbo àwọn ìrànwọ́