VoxCPM

VoxCPM ТТС

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Запиши се. за 5.000 ограничувања на знаци

Опакувајте го вашиот текст во SSML ознаки за прецизна контрола:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Означувања на избраниот модел кои го разбира — кликнете за да го пуштите еден во вашиот текст каде што се случува:

Овој модел чита обичен текст, па значи, во редовните ознаки се игнорираат. За емоции базирани на таг, префрли се на изразителен модел како што е Орфеус или Барк.

Дефинирај ги сопствените изговори (слов = изговор):

-12 +12
0.5x 2.0x
Слободен со Пајпер, ВИТС, Мелотс
Тука ќе се појави вашиот генериран звук. Изберете модел, внесете текст и кликнете на Генерирај.
Аудио генериран успешно
0:00
Симни аудио Симнување.srt Врската истекува за 24 часа
Слободна ставка: лична употреба. Комерцијална дозвола од 5 долари/мо
Кажи им на пријателите!

За VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Најдобро за: High-fidelity audio, audiobooks, long-form content with voice consistency

Прелистувај ги сите VoxCPM гласови

На еден поглед

Развивач
OpenBMB
Лиценца
Apache 2.0
Ниво
standard
Брзина
fast
Гласовно клонирање
Да.
Јазици
English, Chinese
Макс. знаци
2000

VoxCPM гласови

Default

English
Стандардно Neutral

Default Chinese

Chinese
Стандардно Neutral

VoxCPM TTS — Прашања

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Сите гласови