CosyVoice3 TTS
Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.
Wrap wanu malemba mu SSML tags kwa kuwongolera moyenera:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags chosankhidwa chitsanzo amamvetsa - dinani kuti aphe mmodzi m'mawu anu pamene chimachitika:
Izi ndi njira yolemba malemba oyera, kotero ma tag ophatikizidwa amasiya kuganiziridwa. Kuti mupange ma tag ogwirizana ndi maganizo, gwiritsani ntchito njira yolemba malemba monga Orpheus kapena Bark.
Define custom pronunciations (word = pronunciation):
Za CosyVoice3
CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.
Best kwa: Multilingual production TTS, real-time applications, voice cloning
Pezani zonse CosyVoice3 maganizoPa mphindi
- Wopanga
- Alibaba (FunAudioLLM)
- License
- Apache 2.0
- Mtundu
- standard
- Kuyenda
- fast
- Kusintha kwa mawu
- Yes
- Zilankhulo
- English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
- Max characters
- 5000