VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Aanmelden voor 5.000 tekenlimiet

Wrap uw tekst in SSML-tags voor nauwkeurige controle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags het geselecteerde model begrijpt

Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.

Definieer aangepaste uitspraaken (woord = uitspraak):

-12 +12
0.5x 2.0x
Gratis met Piper, VITS, MeloTTS
Uw gegenereerde audio zal hier verschijnen. Kies een model, voer tekst in en klik op Genereren.
Audio Generated Succesvol
0:00
Audio downloaden Download.srt Link verloopt in 24 uur
Gratis niveau: persoonlijk gebruik. Commerciële licentie van $5/mo
Hou van TTS.ai? Vertel het je vrienden!

Info VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Beste voor: High-fidelity audio, audiobooks, long-form content with voice consistency

Alles doorbladeren VoxCPM stemmen

In een oogopslag

Ontwikkelaar
OpenBMB
Licentie
Apache 2.0
Niveau
standard
Snelheid
fast
Klonen van stemmen
Ja.
Talen
English, Chinese
Max. tekens
2000

VoxCPM stemmen

Default

English
Standaard Neutral

Default Chinese

Chinese
Standaard Neutral

VoxCPM Veelgestelde vragen

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Alle stemmen