VoxCPM

VoxCPM ТТС

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Падпісацца Абмежаванне на 5000 знакаў

Захоўваць тэкст у тэгах SSML для дакладнага кантролю:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Тэгі, якія разумее выбраная мадэль - націсніце, каб перанесці іх у тэкст:

Гэтая мадэль чытае звычайны тэкст, таму ўбудаваныя тэгі ігнаруюцца. Для эмоцый, заснаваных на тэгах, пераключыцеся на мадэлі выразнасці, такія як Orpheus або Bark.

Вызначыць уласнае вымаўленне (слова = вымаўленне):

-12 +12
0.5x 2.0x
Свабодны з Piper, VITS, MeloTTS
Створаны вамі гук з' явіцца тут. Выберыце мадэль, увядзіце тэкст і націсніце Стварыць.
Аўдыё паспяхова створанаName
0:00
Сцягнуць гук Сцягнуць.srt Тэрмін дзеяння спасылкі скончыцца праз 24 гадзіны
Любіце TTS.ai? Раскажыце сваім сябрам!

Пра VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Лепшы для: High-fidelity audio, audiobooks, long-form content with voice consistency

Прагляд усіх VoxCPM галасы

Кароткае апісанне

Распрацоўшчык
OpenBMB
Ліцэнзія
Apache 2.0
Стварыць
standard
Хуткасць
fast
Клонаванне голасу
Так
Мовы
English, Chinese
Найбольшая колькасць знакаў
2000

VoxCPM галасы

Default

English
Па змаўчанні Neutral

Default Chinese

Chinese
Па змаўчанні Neutral

VoxCPM Частыя пытанні

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Усе галасы