VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Ojejapo 5000 caracter rehegua límite

Ojehaijey ñe'ẽnguéra etiquetas SSML-pe peteĩ control hekopete g̃uarã:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas ohechakuaáva modelo ojeporavóva - tesãirã peteĩ peteĩva ñe'ẽnguérape, oĩhápe:

Ko modelo ohai texto ndahasyivéva, upévare umi etiqueta oĩva línea ryepýpe ndojehechakuaái. Umi emoción oñemopyendáva etiqueta-pe g̃uarã, oñemoambue peteĩ modelo expresivo-pe taha'e Orfeo térã Bark.

Oñemohenda ñe'ẽnguéra ojehechapyréva (tembiapo = ñe'ẽnguéra):

-12 +12
0.5x 2.0x
Libre Piper, VITS, MeloTTS ndive
Audio-kuéra oguenohẽva ojekuaauka ko'ápe. Oñeporavo peteĩ modelo, omoĩnge ñe'ẽ ha ohesa'ỹijo Generar.
Audio oñemoheñói porã
0:00
Oñeguenohẽ marandu myambue guive Oñeguenohẽ.srt Ko enlace hi'are 24 h rire
Nivel libre: jeiporu personal. Licencia comercial $5/ha'e rupi
Ehayhuetéva TTS.ai? He'i umi iñangirũpe!

Mba'épa VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Oñeha'ãvéva: High-fidelity audio, audiobooks, long-form content with voice consistency

Ojehecha opavave VoxCPM ñe'ẽ

Peteĩ jehecha

Desarrollador
OpenBMB
Licencia
Apache 2.0
Ta'ãnga
standard
Velocidad
fast
Clonación ñe'ẽnguéra rehe
Ha'e
Ñe'ẽ
English, Chinese
Caracteres máx.
2000

VoxCPM ñe'ẽ

Default

English
Estándar Neutral

Default Chinese

Chinese
Estándar Neutral

VoxCPM Pregunta ojehechavéva

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Opaite ñe'ẽ