VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Irreġistra issa għal 5,000 karattru limitu

Wrap test tiegħek fil-tags SSML għall-kontroll preċiż:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags li l-mudell magħżul jifhem — ikklikkja biex tqiegħed waħda fit-test tiegħek fejn jiġri:

Dan il-mudell jaqra test sempliċi, għalhekk it-tags inline huma injorati.Għal emozzjoni bbażata fuq it-tag, aqleb għal mudell espressiv bħal Orpheus jew Bark.

Iddefinixxi pronunzji tad-dwana (kelma = pronunzja):

-12 +12
0.5x 2.0x
Ħieles ma Piper, VITS, MeloTTS
L-awdjo iġġenerat tiegħek se jidher hawnhekk. Agħżel mudell, daħħal it-test, u kklikkja Iġġenera.
Awdjo Iġġenerat b'suċċess
0:00
Niżżel l-awdjo Niżżel.srt Il-link tiskadi f'24 siegħa
Livell Ħieles: użu personali. Liċenzja Kummerċjali minn $5/mo
Agħmel dan il-vuċi tiegħek stess Klona vuċi f'30 sekonda
Imħabba TTS.ai? Għid lill-ħbieb tiegħek!

Dwar VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

L-aħjar għal: High-fidelity audio, audiobooks, long-form content with voice consistency

Ibbrawżja kollox VoxCPM vuċijiet

Daqqa t'għajn

Żviluppatur
OpenBMB
Liċenzja
Apache 2.0
Annimali
standard
Veloċità
fast
Klonazzjoni tal-vuċi
Iva
Lingwi
English, Chinese
Karattri massimi
2000

VoxCPM vuċijiet

Default

English
Standard Neutral

Default Chinese

Chinese
Standard Neutral

VoxCPM TTS — Mistoqsijiet Frekwenti

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Il-vuċijiet kollha