VoxCPM TTS
A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.
Ojehaijey ñe'ẽnguéra etiquetas SSML-pe peteĩ control hekopete g̃uarã:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etiquetas ohechakuaáva modelo ojeporavóva - tesãirã peteĩ peteĩva ñe'ẽnguérape, oĩhápe:
Ko modelo ohai texto ndahasyivéva, upévare umi etiqueta oĩva línea ryepýpe ndojehechakuaái. Umi emoción oñemopyendáva etiqueta-pe g̃uarã, oñemoambue peteĩ modelo expresivo-pe taha'e Orfeo térã Bark.
Oñemohenda ñe'ẽnguéra ojehechapyréva (tembiapo = ñe'ẽnguéra):
Mba'épa VoxCPM
VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.
Oñeha'ãvéva: High-fidelity audio, audiobooks, long-form content with voice consistency
Ojehecha opavave VoxCPM ñe'ẽPeteĩ jehecha
- Desarrollador
- OpenBMB
- Licencia
- Apache 2.0
- Ta'ãnga
- standard
- Velocidad
- fast
- Clonación ñe'ẽnguéra rehe
- Ha'e
- Ñe'ẽ
- English, Chinese
- Caracteres máx.
- 2000