VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Izena eman 5.000 karaktereko muga

Itzulbiratu zure testua SSML etiketetan kontrol zehatzagoa lortzeko:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Hautatutako modeloak ulertzen dituen etiketak — egin klik testuan jartzeko:

Eredu honek testu arrunta irakurtzen du, beraz, lerro-barneko etiketei ez zaie jaramonik egiten. Etiketetan oinarritutako emozioetarako, aldatu Orpheus edo Bark bezalako adierazpen-modelo batera.

Definitu ahoskera pertsonalizatuak (hitza = ahoskera):

-12 +12
0.5x 2.0x
Librea Piper, VITS, MeloTTS-ekin
Zure sortutako audioa hemen agertuko da. Aukeratu modelo bat, idatzi testua eta egin klik Sortu botoian.
Audioa behar bezala sortu da
0:00
Deskargatu audioa Deskargatu.srt Esteka 24 ordutan iraungiko da
Librea: erabiltzaile pribatuentzat. Lizentzia komertziala $5/mo-tik
Maite TTS.ai? Esan zure lagunei!

Honi buruz VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Honako hauentzako onena: High-fidelity audio, audiobooks, long-form content with voice consistency

Arakatu dena VoxCPM ahotsak

Begirada batean

Garatzailea
OpenBMB
Lizentzia
Apache 2.0
Tier
standard
Abiadura
fast
Ahots klonaketa
Bai
Hizkuntzak
English, Chinese
Gehienezko karaktereak
2000

VoxCPM ahotsak

Default

English
Lehenetsia Neutral

Default Chinese

Chinese
Lehenetsia Neutral

VoxCPM TTS — Galdera ohikoenak

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Ahots guztiak