VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Zaregistrovat se pro 5000 znaků limit

Zabalte svůj text do značek SSML pro přesné ovládání:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Značky vybraného modelu rozumí? klikněte na tlačítko pro kapku jednoho do textu, kde se to stane:

Tento model čte prostý text, takže inline značky jsou ignorovány. Pro tag-based emotion, přepněte na expresivní model, jako je Orpheus nebo Bark.

Definovat vlastní výslovnosti (slovo = výslovnost):

-12 +12
0.5x 2.0x
Zdarma s Piper, VITS, MeloTTS
Zde se objeví váš vygenerovaný zvuk. Vyberte model, zadejte text a klikněte na Generovat.
Audio generované úspěšně
0:00
Stáhnout zvuk Stáhnout.srt Odkaz vyprší v 24 hodin
Volný stupeň: osobní použití. Obchodní licence od 5 dolarů/mo
Miluju TTS.ai? Řekni to svým přátelům!

O aplikaci VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Nejlepší pro: High-fidelity audio, audiobooks, long-form content with voice consistency

Procházet vše VoxCPM hlasy

Na první pohled

Vývojář
OpenBMB
Licence
Apache 2.0
Úroveň
standard
Rychlost
fast
Klonování hlasu
Ano.
Jazyky
English, Chinese
Max znaků
2000

VoxCPM hlasy

Default

English
Standardní Neutral

Default Chinese

Chinese
Standardní Neutral

VoxCPM FAQ TTS

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Všechny hlasy