VoxCPM

VoxCPM TTS-värden

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Registrera dig för 5 000 teckengräns

Radera din text i SSML-taggar för exakt kontroll:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Taggar den valda modellen förstår — klicka för att släppa en i din text där det händer:

Den här modellen läser vanlig text, så inline-taggar ignoreras. För taggbaserade känslor, byt till en uttrycksfull modell som Orpheus eller Bark.

Definiera egna uttal (ord = uttal):

-12 +12
0.5x 2.0x
Gratis med Piper, VITS, Melotts
Ditt genererade ljud visas här. Välj en modell, skriv in text och klicka på Generera.
Ljud genereras framgångsrikt
0:00
Ladda ner ljud Ladda ner.srt Länken går ut i 24 timmar
Fri nivå: personlig användning. Kommersiell licens från $5/mo
Berätta för dina vänner!

Om jag inte kan VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Bäst för: High-fidelity audio, audiobooks, long-form content with voice consistency

Bläddra alla VoxCPM röster

Med en blick

Utvecklare
OpenBMB
Licens
Apache 2.0
Nivå
standard
Varvtal
fast
Röstkloning
Ja, det är jag.
Språk
English, Chinese
Max tecken
2000

VoxCPM röster

Default

English
Standardvärde Neutral

Default Chinese

Chinese
Standardvärde Neutral

VoxCPM TTS – FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Alla röster