VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Registrer deg for 5000 tegn- grense

Bryt teksten i SSML- tagger for nøyaktig kontroll:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Merker den valgte modellen forstår – klikk for å slippe en i teksten der den skjer:

Denne modellen leser ren tekst, så merker i teksten blir ignorert. Bytt til en ekspressiv modell som Orfeus eller Bark for å bruke tagger.

Definer selvvalgte uttaler (ord = uttale):

-12 +12
0.5x 2.0x
Fri for piper, VITS, MeloTTS
Her vises din genererte lyd. Velg en modell, skriv inn tekst og trykk Generer.
Lydgenerert vellykket
0:00
Last ned lyd Last ned.srt Lenke utløper om 24 timer
Fritt nivå: personlig bruk. Handelslisens fra $5/mo
Elsker TTS.ai? Fortell vennene dine!

Om VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Best for: High-fidelity audio, audiobooks, long-form content with voice consistency

Bla gjennom alle VoxCPM stemmer

Med et blikk

Utvikler
OpenBMB
Lisens
Apache 2.0
Nivå
standard
Hastighet
fast
Stemmekloning
Ja
Språk
English, Chinese
Største antall tegn
2000

VoxCPM stemmer

Default

English
Standard Neutral

Default Chinese

Chinese
Standard Neutral

VoxCPM TTS — OSS

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Alle stemmer