VoxCPM ТТС
A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.
Опаковка на вашия текст в SSML тагове за точен контрол:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Етикети избрания модел разбира — кликнете за да пуснете един в вашия текст, където се случва:
Този модел чете обикновен текст, така че в линия тагове се пренебрегват. За емоции, базирани на таг, превключете на експресивен модел като Orpheus или Bark.
Определяне на своите изговори (слово = произношение):
За VoxCPM
VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.
Най-добро за: High-fidelity audio, audiobooks, long-form content with voice consistency
Преглед на всички VoxCPM гласовеНа един поглед.
- Разработчик
- OpenBMB
- Лиценз
- Apache 2.0
- Ниво на равнището
- standard
- Скорост
- fast
- Гласово клониране
- Да.
- Езици
- English, Chinese
- Макс. символи
- 2000