VoxCPM TTS
A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.
Ajusta el text a les etiquetes SSML per al control precís:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etiquetes del model seleccionat entenen el clic show clic per a deixar- ne un al text a on succeeix:
Aquest model llegeix text pla, així que les etiquetes inserides s' ignoren. Per a emocions basades en etiquetes, canvieu a un model expressiu com Orfeus o Bark.
Defineix pronúncies personalitzades (word = pronunciació):
Quant a VoxCPM
VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.
Millor per: High-fidelity audio, audiobooks, long-form content with voice consistency
Navega- ho tot VoxCPM veusEn una mirada
- Desenvolupador
- OpenBMB
- Llicència
- Apache 2.0
- TierCity name (optional, probably does not need a translation)
- standard
- Velocitat
- fast
- clonació de veu
- Sí
- Idiomes
- English, Chinese
- Nombre màxim de caràcters
- 2000