VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

رجسٹر کریں 5000 حروف کی حد

SSML ٹیگ میں اپنے متن کو دقيق کنٹرول کے لیے لپیٹیں:

<speak><prosody rate="slow">Slow speech</prosody></speak>

تگ منتخب ماڈل سمجھتا هے - آپ کے متن ميں ایک ڈالنے کے ليے کلک کريں جہاں هيں :

هي ماڈل عام متن پڑھتا هے ، تو ان لائن ٹگز کو نظر انداز کريں تا گ ڑي گي ايميشن کے ليے ، اورف يا بارک کے ليے بيان کر نے والے ماڈل ميں تبديل کريں

خود ساختہ تلفظ (لفظ = تلفظ):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS کے ساتھ مفت
آپ کا بنا يا او ڊيو یہاں نظر آ ئي گا ماڈل منتخب کريں ، متن داخل کريں اور بنا ئيں کلک کريں
آڈیو کامیابی سے پیدا کی گئی
0:00
آڈیو ڈاؤن لوڈ کریں .srt ڈاؤن لوڈ کریں رابطہ 24 گھنٹوں میں ختم ہو جاتا ہے
مفت سطح: ذاتی استعمال. $5/مئی سے تجارتی لائسنس
TTS.ai سے محبت؟ اپنے دوستوں کو بتائیں!

متعلقہ VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

بہترین: High-fidelity audio, audiobooks, long-form content with voice consistency

سب براؤز کریں VoxCPM آوازیں

ایک نظر میں

ڈیولپر
OpenBMB
لائسنس
Apache 2.0
تير
standard
رفتار
fast
آواز کا کلوننگ
جی ہاں
زبانیں
English, Chinese
زیادہ سے زیادہ حروف
2000

VoxCPM آوازیں

Default

English
معیار Neutral

Default Chinese

Chinese
معیار Neutral

VoxCPM TTS - FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← تمام آوازیں