VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Înscrie-te pentru limitele de 5000 de caractere

Întoarceți textul în etichetele SSML pentru un control precis:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etichetele modelului selectat înțeleg — click pentru a lăsa unul în textul tău unde se întâmplă:

Acest model citește textul simplu, astfel încât etichetele inline sunt ignorate. Pentru emoții bazate pe tag, schimbați la un model expresiv cum ar fi Orpheus sau Bark.

Definiți pronunțiare personalizată (cuvânt = pronunție):

-12 +12
0.5x 2.0x
Gratuit cu Piper, VITS, MeloTTS
Audio generat va apărea aici. Alegeți un model, introduceți text și faceți clic pe Generați.
Audio generat cu succes
0:00
Descarcă audio Descărcare.srt Legătura expiră în 24 ore
Gratuit: utilizare personală. Licență comercială de la 5$/mo
Spune-i prietenilor tăi!

Despre VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Cel mai bun pentru: High-fidelity audio, audiobooks, long-form content with voice consistency

Navigați toate VoxCPM voci

La o privire

Dezvoltator
OpenBMB
Licență
Apache 2.0
Nivel
standard
Viteză
fast
Clonarea vocală
Da.
Limbi
English, Chinese
Caractere maxime
2000

VoxCPM voci

Default

English
Standard Neutral

Default Chinese

Chinese
Standard Neutral

VoxCPM TTS – FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Toate vocile