VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Prijavite se za ograničenje od 5.000 znakova

Omotajte tekst u SSML oznake za preciznu kontrolu:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Oznake koje odabrani model razumije — kliknite da biste ih ubacili u tekst gdje se pojavljuju:

Ovaj model čita običan tekst, tako da se inline oznake ignoriraju. Za emocije zasnovane na oznakama, prebacite se na ekspresivni model poput Orpheusa ili Bark-a.

Definirajte vlastite izgovore (riječ = izgovor):

-12 +12
0.5x 2.0x
Besplatno sa Piper, VITS, MeloTTS
Ovdje će se pojaviti vaš generirani audio. Izaberite model, unesite tekst i kliknite na Generiraj.
Audio uspješno generisan
0:00
Preuzmi audio Preuzmi.srt Link istječe za 24h
Free tier: personal use. Komercijalna licenca od $5/mjesečno
Volite TTS.ai?

O meni VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Najbolje za: High-fidelity audio, audiobooks, long-form content with voice consistency

Pregledaj sve VoxCPM glasovi

Na prvi pogled

Programer
OpenBMB
Licenca
Apache 2.0
Životinje
standard
Brzina
fast
Kloniranje glasa
Da.
Jezici
English, Chinese
Maksimalan broj znakova
2000

VoxCPM glasovi

Default

English
Standardni Neutral

Default Chinese

Chinese
Standardni Neutral

VoxCPM FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Svi glasovi