VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Inscríbete límite de 5. 000 caracteres

Incluír o texto en etiquetas SSML para un control preciso:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas que o modelo escollido entende - prema para deixar unha no texto onde ocorre:

Este modelo le texto simple, polo que se ignoran as etiquetas inline. Para emocións baseadas en etiquetas, cambie a un modelo expresivo como Orpheus ou Bark.

Definir pronunciacións personalizadas (palabra = pronunciación):

-12 +12
0.5x 2.0x
Libre con Piper, VITS, MeloTTS
O son xerado aparecerá aquí. Escolla un modelo, introduza o texto e prema Xerar.
O son xerou correctamente
0:00
Obter o son Obter.srt A ligazón caduca en 24 horas
Nivel libre: uso persoal. Licenza comercial desde $5/mes
Encántalle TTS.ai? Cóntallo aos teus amigos!

Acerca de VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Mellor para: High-fidelity audio, audiobooks, long-form content with voice consistency

Examinar todo VoxCPM voces

De un vistazo

Desenvolvente
OpenBMB
Licenza
Apache 2.0
Tier
standard
Velocidade
fast
Clonaxe de voz
Si
Linguas
English, Chinese
Caracteres máximos
2000

VoxCPM voces

Default

English
Estándar Neutral

Default Chinese

Chinese
Estándar Neutral

VoxCPM TTS - FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Todas as voces