VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Whakawhanake mō te tepe o ngā tohu 5,000

Whāriki i tōna kupu i roto i ngā tohu SSML mō te whakahaere tika:

<speak><prosody rate="slow">Slow speech</prosody></speak>

E mōhio ana ngā tohu ki te tauira i kōwhiria - ka kōwhiria kia whakawātea tētahi ki roto i tōna kupu i reira ka puta ai:

Ka pānui tēnei tauira i te kupu noa, nā reira ka whakakāhoretia ngā tohu ā-waitara. Mō te āhua o te tohu-taihi, ka huri ki tētahi tauira whakamārama pēnei i a Orpheus, Bark rānei.

Ka tautuhia ngā tohutohu ā-ringa (wāhi = tohutohu):

-12 +12
0.5x 2.0x
Waihoki me Piper, VITS, MeloTTS
Ka puta tēnei te oro i waihangatia e koe. Ka kōwhiria tētahi tauira, ka tāurua te kupu, a, ka kōwhiria te Whakatū.
Kua angitu te whakaputanga oro
0:00
Waihoki i te oro Whakahua.srt Ka ngaro te pānga i roto i te 24h
Tauwhāinga wātea: te whakamahinga whaiaro. Whakawhiwhinga hokohoko mai i te $5/mo
E manakohia ana e TTS.ai? Whakapāpāho ki ōna hoa!

Mo VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Pai mo: High-fidelity audio, audiobooks, long-form content with voice consistency

Ka tirohia katoa VoxCPM ngā oro

I te tirohanga

Ka whakawhanakehia
OpenBMB
Ka taea te whakawātea
Apache 2.0
Karaka
standard
Āhuatanga
fast
Whakakōrero reo
He
reo
English, Chinese
Kāri nui rawa
2000

VoxCPM ngā oro

Default

English
Paerewa Neutral

Default Chinese

Chinese
Paerewa Neutral

VoxCPM TTS - FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Ko nga oro katoa