CosyVoice3

CosyVoice3 TTS

Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.

Aliĝi for 5, 000 character limit

Envolvu vian tekston en SSML- etikedojn por preciza kontrolo:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etikedoj kiujn la elektita modelo komprenas - klaku por meti unu en vian tekston kie ĝi okazas:

This model reads plain text, so inline tags are ignored. For tag-based emotion, switch to an expressive model like Orpheus or Bark.

Difini proprajn elparolojn (vorto = elparolo):

-12 +12
0.5x 2.0x
Libera kun Piper, VITS, MeloTTS
Via generita sono aperos tie ĉi. Elektu modelon, entajpu tekston, kaj alklaku Generi.
Sondosiero sukcese generita
0:00
Elŝuti sonon Elŝuti.srt Ligo eksvalidiĝas post 24 horoj
Libera programaro: persona uzo. Komerca licenco ekde $5/mo
Ĉu vi ŝatas TTS.ai? Diru al viaj amikoj!

Pri CosyVoice3

CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.

Plej bona por: Multilingual production TTS, real-time applications, voice cloning

Foliumi ĉiujn CosyVoice3 voĉoj

Unu rigardo

Programisto
Alibaba (FunAudioLLM)
Licenco
Apache 2.0
Tamuz
standard
Rapideco
fast
Voĉo- klonado
Jes
Lingvoj
English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
Maksimuma nombro da signoj
5000

CosyVoice3 voĉoj

Chinese Female

Chinese
Defaŭlta Female

Chinese Male

Chinese
Defaŭlta Male

English Female

English
Defaŭlta Female

English Male

English
Defaŭlta Male

French Female

French
Defaŭlta Female

German Female

German
Defaŭlta Female

Italian Female

Italian
Defaŭlta Female

Japanese Female

Japanese
Defaŭlta Female

Korean Female

Korean
Defaŭlta Female

Russian Female

Russian
Defaŭlta Female

Spanish Female

Spanish
Defaŭlta Female

CosyVoice3 TTS - FAQ

CosyVoice3 adds bi-streaming inference at around 150ms latency, instruction-based control over emotion/speed/volume, improved speaker similarity for cloning, and coverage of 9 languages plus 18 Chinese dialects, with an RL-tuned variant for state-of-the-art prosody.

Yes. It supports zero-shot voice cloning from a reference clip (around 3 seconds minimum) with improved speaker similarity over the previous generation.

Yes. CosyVoice3 is licensed under Apache 2.0, permitting commercial use.
← Ĉiuj voĉoj