VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Melden Sie sich an für 5.000 Zeichen-Grenze

Verpacken Sie Ihren Text in SSML-Tags für eine präzise Kontrolle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags, die das ausgewählte Modell versteht — klicken Sie, um einen in Ihren Text zu legen, wo es passiert:

Dieses Modell liest Text, so dass Inline-Tags ignoriert werden. Für tag-basierte Emotion, wechseln Sie zu einem ausdrucksstarken Modell wie Orpheus oder Bark.

Benutzerdefinierte Aussprachen definieren (Wort = Aussprache):

-12 +12
0.5x 2.0x
Frei mit Piper, VITS, MeloTTS
Hier erscheint Ihr generiertes Audio. Wählen Sie ein Modell, geben Sie Text ein und klicken Sie auf Generieren.
Audio-Erzeugung erfolgreich
0:00
Audio herunterladen Download.srt Link läuft in 24h aus
Freier Dienstgrad: persönlicher Gebrauch. Kommerzielle Lizenz ab $5/mo
Gefällt dir TTS.ai? Erzähl es deinen Freunden!

Über VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Das Beste für: High-fidelity audio, audiobooks, long-form content with voice consistency

Alle durchsuchen VoxCPM Stimmen

Auf einen Blick

Entwickler
OpenBMB
Lizenz
Apache 2.0
Tierart
standard
Geschwindigkeit
fast
Klonen der Stimme
Nein
Sprachen
English, Chinese
Maximale Zeichen
2000

VoxCPM Stimmen

Default

English
Standard Neutral

Default Chinese

Chinese
Standard Neutral

VoxCPM TTS — FAQ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Alle Stimmen