VoxCPM

VoxCPM TT TT TTT TTT T TT TT T T TTT TTT TTT

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

签名 对 5,000 字符限制的 5 000 个字符

在 SSML 标记中折行文本以精确控制 :

<speak><prosody rate="slow">Slow speech</prosody></speak>

标记选中模式的理解度 - 单击将一个输入到文本中, 发生时 :

这个模型读的是简单的文字, 所以内嵌标签会被忽略。 对于基于标签的情感, 请切换到像 Orpheus 或 Bark 这样的表达模式 。

定义自定义发音( Word = 发音) :

-12 +12
0.5x 2.0x
免费的管道、VITS、MelotTS
您生成的音频将在此显示。 选择一个模型, 输入文本, 并单击生成 。
音频生成成功
0:00
下载音频 下载.strt 24小时后链接过期
免费:个人使用。 5美元/美元商业许可证
使这个声音成为你自己的声音 30秒后打开声音
喜欢TTS.ai吗?告诉你的朋友吧!

关于 VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

最佳: High-fidelity audio, audiobooks, long-form content with voice consistency

全部浏览 VoxCPM 声音

一眼看一眼,

开发者
OpenBMB
许可证
Apache 2.0
级别
standard
速度
fast
语音克隆
语言
English, Chinese
最大字符
2000

VoxCPM 声音

Default

English
标准 Neutral

Default Chinese

Chinese
标准 Neutral

VoxCPM TTS - 常见问题

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← 所有声音