VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Anmelden Limit fir 5. 000 Zeichen

Wrap your text in SSML tags for precise control:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags déi d'gewielt Modell verstinn - klickt fir eng an Ärem Text ze setzen wou se geschitt:

Dëse Modell liest einfache Text, sou datt Inline-Tags ignoréiert ginn. Fir Tag-baséiert Emotiounen, wielt e expressiven Modell wéi Orpheus oder Bark.

Eegen Aussproochen definéieren (Wuert = Aussprooch):

-12 +12
0.5x 2.0x
Free mat Piper, VITS, MeloTTS
Äert generéiert Audio wäert hei erscheinen. Wielt e Modell, gitt Text an a klickt op Generéieren.
Audio gouf erfollegräich generéiert
0:00
Audio erofgelueden Lëscht vu lëtzebuergesche Schrëftsteller Link expires in 24h
Den Haaptuert ass Personnes. Kommerziell Lizenz vun $5/mo
Maacht dat Är eege Stëmm Klonen eng Stëmm an 30 Sekonnen
Liewe TTS.ai? Erzielt Är Frënn!

Iwwer VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Bescht fir: High-fidelity audio, audiobooks, long-form content with voice consistency

All sichen VoxCPM Stimmen

Op ee Bléck

Entwéckler
OpenBMB
Lizenz
Apache 2.0
Tier
standard
Geschwindegkeet
fast
Sprooche-Klonen
Ja
Sproochen
English, Chinese
Maximal Zeichen
2000

VoxCPM Stimmen

Default

English
Standard Neutral

Default Chinese

Chinese
Standard Neutral

VoxCPM Lëscht vun de FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← All Stimmen