VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Ndaftar for 5,000 characters limit

Nglapisi teks ing tag SSML kanggo kontrol sing tepat:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag kang dipahami model kang dipilih — klik kanggo ngethok siji menyang teks sampeyan ing ngendi iku kedadeyan:

Model ieu maca teks biasa, jadi tag inline diabaikan. Pikeun emosi dumasar tag, ganti ka model ekspresif kayaning Orpheus atawa Bark.

Nyathet pangucapan standar (kata = pangucapan):

-12 +12
0.5x 2.0x
Bebas karo Piper, VITS, MeloTTS
Audio anu dihasilkeun bakal muncul di dieu. Pilih model, ketok teks, sarta ketok Janji.
Audio berhasil diciptakan
0:00
Muat turun audio Muat turun.srt Link expires in 24h
Kacamatan iki kalebu: Kacamatan Semarang. Lisénsi komersial saka $ 5 / mo
Love TTS.ai? Nyathet kanca-kancamu!

About VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Paling apik kanggo: High-fidelity audio, audiobooks, long-form content with voice consistency

Nglayar kabeh VoxCPM suara

Ing cetha

Pangembang
OpenBMB
Lisensi
Apache 2.0
Tingkat
standard
Kecepatan
fast
Kloning suara
Iya
Basa
English, Chinese
Karakter paling akeh
2000

VoxCPM suara

Default

English
Standar Neutral

Default Chinese

Chinese
Standar Neutral

VoxCPM TTS — FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Sekabeh swara