VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Գրանցվել 5000 սանտիմետր սահմանափակում

Ձեր տեքստը SSML տեգերի մեջ տեղադրել ճշգրիտ կառավարման համար.

<speak><prosody rate="slow">Slow speech</prosody></speak>

Ընտրված մոդելի կողմից ընկալվող պիտակներ — սեղմեք դրանք ձեր տեքստում տեղադրելու համար:

Այս մոդելը կարդում է պարզ տեքստը, այնպես որ ներառված տեքստերը անտեսվում են։ Տեքստերի վրա հիմնված էմոցիաների համար փոխեք արտահայտիչ մոդել, ինչպես Orpheus կամ Bark։

Որոշել սեփական արտասանությունը (բառ = արտասանություն):

-12 +12
0.5x 2.0x
Ազատ Piper, VITS, MeloTTS-ով
Այստեղ կհայտնվի ձեր ստեղծած ձայնը։ Ընտրեք մոդել, ներդրեք տեքստ և սեղմեք Ծնվել։
Ավտոմատ ձայնագրում
0:00
Տեղադրել ձայնային Տեղադրել.srt Հղումն ավարտվում է 24 ժամ անց
Ֆրանսիա: Ֆրանսիայի ազգային ռադիո: Ֆրանսիայի ազգային ռադիո. Կազմակերպական լիցենզիա $5/ամս
Սիրում եք TTS.ai-ն? Պատմեք ձեր ընկերներին։

Ընդհանուր VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Լավագույնը: High-fidelity audio, audiobooks, long-form content with voice consistency

Ընթերցել բոլորը VoxCPM ձայներ

Հակառակորդի դիրքը

Հեղինակ
OpenBMB
Լիցենզիա
Apache 2.0
Դադար
standard
արագություն
fast
Ձայնի կլոնավորում
Այո
Լեզուներ
English, Chinese
Օգտագործված ռեժիմ
2000

VoxCPM ձայներ

Default

English
Լռելյայն Neutral

Default Chinese

Chinese
Լռելյայն Neutral

VoxCPM TTS - Հաճախ տրվող հարցեր

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Բոլոր ձայները