VoxCPM

VoxCPM ТТС

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Каттоо 5000 символго чейин

Текстти SSML тегдерине өткөрүп берүү:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Белгилер, тандалган модель түшүнөт — текстке бирин коюу үчүн, аны чыкылдатыңыз:

Бул модель жөнөкөй текстти окуйт, ошондуктан тексттеги тегдер эске алынбайт. Тегдер менен эмоцияларды жаратууда Orpheus же Bark сыяктуу экспрессивдүү моделдерге өтүү керек.

Өзгөчө сүйлөмдөрдү аныктоо (сөз = сүйлөм):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS менен акысыз
Сиздин түзүлгөн аудио файлыңыз бул жерде пайда болот. Модель тандап, текстти киргизип, Жаңылоо баскычын басыңыз.
Аудио ийгиликтүү түзүлгөн
0:00
Аудиону жүктөп алуу .srt жүктөп алуу Ссылканын мөөнөтү 24 сааттан кийин аяктайт
Free level: жеке колдонуу. Коммерциялык лицензия $5/айдан
Бул үндү өзүңүздүн үнүңүзгө айландырыңыз 30 секундда үндү клондоо
TTS.ai сизге жактыбы? Досторуңузга айтып коюңуз!

Маалымат VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Эң жакшысы: High-fidelity audio, audiobooks, long-form content with voice consistency

Баарын кароо VoxCPM үн

Бир көз салуу

Жазуучу
OpenBMB
Лицензия
Apache 2.0
Тигр
standard
Жылдамдыгы
fast
Сөздү клондоо
Ооба
Тилдер
English, Chinese
Макс. символдор
2000

VoxCPM үн

Default

English
Стандарттык Neutral

Default Chinese

Chinese
Стандарттык Neutral

VoxCPM ТТС — Жогорудагы суроолор

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Бардык үн