CosyVoice 2 TTS
Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.
Amlapio' ch testun mewn tagiau SSML er mwyn cael rheoli cywir:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags y deall y model dewisiedig - cliciwch i daflu un i' ch testun lle mae' n digwydd:
Mae'r model yma yn darllen testun plaen, felly anwybyddir tagiau mewnlin. I ddelweddu teimlad yn seiliedig ar dagiau, newidiwch i ddelweddu mynegiant fel Orpheus neu Bark.
Diffinio ynganiad addasiedig (gair = ynganiad):
Am CosyVoice 2
CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.
Gorau ar gyfer: Real-time applications, streaming TTS, voice assistants
Pori Popeth CosyVoice 2 SaesnegYn syth
- Datblygwr
- Alibaba (Tongyi Lab)
- Trwydded
- Apache 2.0
- o Fawrth
- standard
- Cyflymder
- medium
- Clonio llais
- IeQShortcut
- Iaith:
- English, Chinese, Japanese, Korean, French, German, Italian, Spanish
- Uchafswm nodau
- 1000