VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

אַרײַנשרײַבן פֿאַר 5,000 שריפֿטצײכן

איבער־פֿאַרקער דעם טעקסט אין SSML הענטלעך פֿאַר אַ פּשוטער קאָנטראָל:

<speak><prosody rate="slow">Slow speech</prosody></speak>

הענטלעך װאָס דער אויסגעקליבן מאָדעל פֿאַרשטײט — קליק צו אַרײַנשרײַבן אײן אין װײַז־טעקסט װוּ עס פּאַסט זיך:

דאָס מאָדעל לייענט נאָרמאַלן טעקסט, אַזוי אַרײַנגעלייגטע הענטלעך ווערן איגנאָרירט. פֿאַר הענטלעך־באזירטע װײַב־איגנאָרירונג, װײַז צו אַ װײַז־מאָדעל װי אורפיאָ אָדער װאַרק

װײַז פֿאָרױסװײַז

-12 +12
0.5x 2.0x
פֿרײַ מיט Piper, VITS, MeloTTS
די אױדיו־טעקע װעט דאָ װײַזן זיך. קלײַב אַ מאָדעל אױס, אַרײַנשרײַב דעם טעקסט און קליק אױף אױסשרײַבן
אוודיאָ אױסגעגרײט
0:00
דאַונלאָוד אַרײַנשטעלן.srt פֿאַרבינדונג ענדיקט זיך אין 24 שעה
פרייע מדרגה: פּערזענלעכער ניצן קאָממערשעל ליסענסע פֿון $5/חודש
ליבע TTS.ai? זאָגן דיין פריינט

אױף VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

בעסטער פֿאַר: High-fidelity audio, audiobooks, long-form content with voice consistency

בלעטער VoxCPM שריפֿטצײכן

אין אַ שריט

אַנטוויקלער
OpenBMB
לינקס
Apache 2.0
װײַז
standard
שאַטירונג
fast
שפּראַך
יאָ
שפּראַכן
English, Chinese
גרעסטע שריפֿטצײכן
2000

VoxCPM שריפֿטצײכן

Default

English
סטאַנדאַרד Neutral

Default Chinese

Chinese
סטאַנדאַרד Neutral

VoxCPM FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← אַלע שפּראַכן