VoxCPM

VoxCPM TTS

A tokenizer-free TTS model that works in continuous space, outputs 44.1kHz audio, and stays consistent across paragraphs.

Εγγραφείτε για όριο 5.000 χαρακτήρων

Τυλίξτε το κείμενο σας σε ετικέτες EEML για τον ακριβή έλεγχο:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Ετικέτες το επιλεγμένο μοντέλο καταλαβαίνει □ κάντε κλικ για να ρίξετε ένα στο κείμενο σας όπου συμβαίνει:

Αυτό το μοντέλο διαβάζει απλό κείμενο, έτσι οι ετικέτες inline αγνοούνται. Για tag-based συναίσθημα, μεταβείτε σε ένα εκφραστικό μοντέλο όπως ο Ορφέας ή Bark.

Define custom προφορές (word = εκφώνηση):

-12 +12
0.5x 2.0x
Δωρεάν με Piper, VITS, MeloTTS
Επιλέξτε ένα μοντέλο, εισάγετε το κείμενο και κάντε κλικ στη Δημιουργία.
Ο Ήχος Δημιουργήθηκε Επιτυχώς
0:00
Λήψη ήχου Κατεβάστε.srt Η σύνδεση λήγει σε 24 ώρες
Δωρεάν βαθμίδα: προσωπική χρήση. Εμπορική άδεια από $ 5 /mo
Τρέξιμο χαμηλό σε ελεύθερους χαρακτήρες Πάρτε 200K χαρακτήρες κάθε μήνα ~ $ 5 /mo ή ένα πακέτο 100K για $5
Κάνε αυτό τη δική σου φωνή. Κλώνε μια φωνή σε 30 δευτερόλεπτα
Αγάπη TTS.ai; Πες στους φίλους σου!

Σχετικά VoxCPM

VoxCPM 1.5 by OpenBMB takes an unusual approach: instead of converting speech into discrete tokens, it operates directly in continuous space, which helps it preserve fine acoustic detail. It produces high-fidelity 44.1kHz audio, supports zero-shot voice cloning from three to ten seconds of reference, and maintains a consistent voice across long passages — a common failure point for other models on multi-paragraph text. Its cross-language cloning lets an English reference voice speak Chinese and vice versa. With Apache 2.0 licensing and LoRA fine-tuning support, it is well suited to audiobooks and long-form content where voice consistency over many paragraphs is essential.

Το καλύτερο για: High-fidelity audio, audiobooks, long-form content with voice consistency

Περιήγηση σε όλα VoxCPM φωνές

Με μια ματιά.

Προγραμματιστής
OpenBMB
Άδεια
Apache 2.0
Βαθμίδα
standard
Ταχύτητα
fast
Κλωνοποίηση φωνής
Ναι.
Γλώσσες
English, Chinese
Μεγ. χαρακτήρες
2000

VoxCPM φωνές

Default

English
Πρότυπο Neutral

Default Chinese

Chinese
Πρότυπο Neutral

VoxCPM TTS □ Συχνές ερωτήσεις

Rather than discretizing speech into tokens, VoxCPM models audio in continuous space using flow matching. This helps it retain subtle acoustic detail and produce clean 44.1kHz output.

Yes. It is specifically designed to keep the voice consistent across paragraphs, which makes it well suited to audiobooks and other long passages where other models tend to drift.

Yes. It supports cross-lingual cloning between English and Chinese — for example applying an English reference voice to Chinese speech — from three to ten seconds of audio.
← Όλες οι φωνές