VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

ثبت نام برای حد ۵۰۰۰ کاراکتر

برای کنترل دقیق ، متن خود را در برچسبهای SSML بپیچید:

<speak><prosody rate="slow">Slow speech</prosody></speak>

برچسبهایی که مدل برگزیده می‌فهمد — برای انداختن یکی در متن خود ، جایی که اتفاق می‌افتد ، کلیک کنید:

این مدل متن ساده را می‌خواند ، بنابراین برچسب‌های خطی نادیده گرفته می‌شوند. برای احساسات مبتنی بر برچسب ، به یک مدل بیانی مانند Orpheus یا Bark تغییر دهید.

تعریف تلفظ سفارشی) کلمه = تلفظ (:

-12 +12
0.5x 2.0x
آزاد با Piper, VITS, MeloTTS
صدای تولید شده شما در اینجا ظاهر خواهد شد. یک مدل را انتخاب کنید ، متن را وارد کنید ، و تولید را فشار دهید.
صدا با موفقیت تولید شد
0:00
بارگیری صدا دانلود پیوند در ۲۴ ساعت پایان می‌یابد
1- استفاده شخصی: استفاده شخصی. مجوز تجاری از $5/mo
دوست داريد TTS.ai؟ به دوستانتون بگو!

در مورد VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

بهترین برای: High-fidelity audio, audiobooks, long-form content with voice consistency

مرور همۀ VoxCPM صداها

يه نگاهي بنداز

توسعه‌دهنده
OpenBMB
مجوز
Apache 2.0
حیوان
standard
سرعت
fast
شبیه‌سازی صدا
آره
زبانها
English, Chinese
بیشینه نویسه‌ها
2000

VoxCPM صداها

Default

English
پیش‌فرض Neutral

Default Chinese

Chinese
پیش‌فرض Neutral

VoxCPM FAQ - پرسش و پاسخ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← همه صداها