VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Kugadzwa for 5,000 character limit

Wrap yako tenzi mu SSML tags kuti zvive nyore kudzora:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags iyo yakasarudzwa model inonzwisiswa — tinya kuti utore imwe muchinyorwa chako apo inoitika:

Iyi modhi inoverenga chete mazita ezvinyorwa, saka zvinyorwa zvinonyorwa mumitsara hazvina kutariswa. Kuti uwane pfungwa dziri mumitsara, chinja kune imwe modhi inoratidza pfungwa seOrpheus kana Bark.

Define custom pronunciations (word = pronunciation):

-12 +12
0.5x 2.0x
Free with Piper, VITS, MeloTTS
Yako yakagadzirwa audio ichaonekwa pano. Choose a model, enter text, and click Generate.
Audio Yakagadzirwa Nekubudirira
0:00
Download Audio Dhawunirodha.srt Link inotanga kushanda mu 24h
Free tier: personal usage. Commercial license from $5/mo
Love TTS.ai? Tiudza shamwari dzako!

Chii VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Yakanaka kune: High-fidelity audio, audiobooks, long-form content with voice consistency

Tarisa zvese VoxCPM mazwi

Mufananidzo

Developer
OpenBMB
License
Apache 2.0
Tier
standard
Speed
fast
Kutaura
Yes
Zvinhu
English, Chinese
Max characters
2000

VoxCPM mazwi

Default

English
Chimiro Neutral

Default Chinese

Chinese
Chimiro Neutral

VoxCPM TTS — Zvinyorwa

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← All voices