VoxCPM TTS
A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.
Wrap wanu malemba mu SSML tags kwa kuwongolera moyenera:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags chosankhidwa chitsanzo amamvetsa - dinani kuti aphe mmodzi m'mawu anu pamene chimachitika:
Izi ndi njira yolemba malemba oyera, kotero ma tag ophatikizidwa amasiya kuganiziridwa. Kuti mupange ma tag ogwirizana ndi maganizo, gwiritsani ntchito njira yolemba malemba monga Orpheus kapena Bark.
Define custom pronunciations (word = pronunciation):
Za VoxCPM
VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.
Best kwa: High-fidelity audio, audiobooks, long-form content with voice consistency
Pezani zonse VoxCPM maganizoPa mphindi
- Wopanga
- OpenBMB
- License
- Apache 2.0
- Mtundu
- standard
- Kuyenda
- fast
- Kusintha kwa mawu
- Yes
- Zilankhulo
- English, Chinese
- Max characters
- 2000