VoxCPM

VoxCPM የድምፅ ፋይል

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

ምዝገባ ፊደል(ሎች)

ርዕሱን በSSML መለያዎች ውስጥ ለጥሩ ቁጥጥር ይዞሩት:

<speak><prosody rate="slow">Slow speech</prosody></speak>

የተመረጠው ሞዴል የሚያውቃቸው መለያዎች - በጽሑፍዎ ውስጥ የሚከሰትበትን ቦታ ለመውሰድ ጠቅ ያድርጉ፦

ይህ ሞዴል ቀላል ጽሑፍን ያነባል፣ ስለዚህም በመስመር ውስጥ ያሉ ምልክቶች ይዘገያሉ፡፡ ለታክስ-ተኮር ስሜት እንደ ኦርፊየስ ወይም ባርክ ያሉ ግልጽ ሞዴሎችን ይለውጡ።

የራሱን ተናጋሪ ግለጽ (ቃል = ተናጋሪ):

-12 +12
0.5x 2.0x
ነጻ ከፒፐር, VITS, MeloTTS ጋር
የእርስዎ የተፈጠረ ድምፅ እዚህ ይታይ. ሞዴል ይምረጡ፣ ጽሑፍ ያስገቡ፣ እና ይፈጥሩ ላይ ጠቅ ያድርጉ
ድምፅ በፍጥነት ተፈጠረ
0:00
ድምፅ ያውርዱ ያውርዱ መገናኛ በ24 ሰዓት ውስጥ ይቋረጣል
ነጻ ደረጃ: የግል ጥቅም የኮሜርሲ ውል ከ $5/mo
TTS.aiን ወዳጅነት?

ስለ VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

ምርጥ ለ: High-fidelity audio, audiobooks, long-form content with voice consistency

ሁሉንም አጥፉ VoxCPM ድምጾች

በጥቂቱ

የድር አዘጋጅ
OpenBMB
ፈቃድ
Apache 2.0
ዐምድ
standard
ፍጥነት
fast
የድምፅ ቅጂ
አዎ
ቋንቋዎች
English, Chinese
ፊደላት
2000

VoxCPM ድምጾች

Default

English
መደበኛ Neutral

Default Chinese

Chinese
መደበኛ Neutral

VoxCPM የትርጉም መሳሪያ

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← ሁሉንም ድምጾች