VoxCPM

VoxCPM ТТС

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Регистрация 5000 символга кадәр

Матныгызны төгәл контроль өчен SSML теглары белән әйләндерегез:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Сайланган модель аңлаган теглар — аларны җөмләгә төшерү өчен өстәп куегыз:

Бу модель гади текстны укый, шуңа күрә эчтәге теглар игътибарга алынмый. Тэгларга нигезләнгән эмоцияләр өчен, Orpheus яки Bark кебек иҗади модельгә күчегез.

Үзенчәлекле әйтелешне билгеләгез (сүз = әйтелеш):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS белән бушлай
Сезнең барлыкка китерелгән аудио монда күренәчәк. Модельне сайлагыз, мәтнне кертегез, һәм "Ярату" төймәсен басыгыз.
Аудио уңышлы төзелде
0:00
Аудио йөкләү .srt файлын төшерү Сүзнең вакыты 24 сәгатьтән соң бетә
РФ су реестры мәгълүматлары: Персона. Коммерцияле лицензия $5/аена
TTS.ai-ны яратасызмы? Дусларыгызга сөйләгез!

Бәйләнешләр VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Иң яхшысы: High-fidelity audio, audiobooks, long-form content with voice consistency

Барлыгын карау VoxCPM тавышлар

Бер карашка

Программист
OpenBMB
Лицензия
Apache 2.0
Гыйнвар
standard
Югары тизлек
fast
Сүзләрне клонлау
Әйе
Телләр
English, Chinese
Макс. символлар саны
2000

VoxCPM тавышлар

Default

English
Стандарт Neutral

Default Chinese

Chinese
Стандарт Neutral

VoxCPM РФ су реестры мәгълүматлары: Фурга.

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Барлык тавышлар