VoxCPM TTS
A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.
Ampidiro anatin'ny tag SSML ny lahabolana mba hahazoana fifehezana mazava tsara:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tag fantatry ny modely voafaritra — tsindrio mba hametrahana iray ao anatin'ny lahatsoratrao izay misy azy:
Mamakiana lahabolana tsotra io modely io, ka tsy raharahaina ny tag anatin'ny andalana. Raha mila fihetseham-po mifototra amin'ny tag ianao, dia miova ho modely maneho fihetseham-po toy ny Orpheus na Bark.
Mamaritra ny fanononana safidy (teny = fanononana):
Mombamomba VoxCPM
VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.
Tsara indrindra ho an'ny: High-fidelity audio, audiobooks, long-form content with voice consistency
Jereo izy rehetra VoxCPM feoAmin'ny fijery fohy
- Mpamorona
- OpenBMB
- Lisansa
- Apache 2.0
- Taona
- standard
- Hafainganan'ny fanovana
- fast
- Fandraisana feo
- Eny
- Teny
- English, Chinese
- Marika betsaka indrindra
- 2000