CosyVoice 2 TT TT TTT TTT T TT TT T T TTT TTT TTT
Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.
在 SSML 标记中折行文本以精确控制 :
<speak><prosody rate="slow">Slow speech</prosody></speak>
标记选中模式的理解度 - 单击将一个输入到文本中, 发生时 :
这个模型读的是简单的文字, 所以内嵌标签会被忽略。 对于基于标签的情感, 请切换到像 Orpheus 或 Bark 这样的表达模式 。
定义自定义发音( Word = 发音) :
关于 CosyVoice 2
CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.
最佳: Real-time applications, streaming TTS, voice assistants
全部浏览 CosyVoice 2 声音一眼看一眼,
- 开发者
- Alibaba (Tongyi Lab)
- 许可证
- Apache 2.0
- 级别
- standard
- 速度
- medium
- 语音克隆
- 是
- 语言
- English, Chinese, Japanese, Korean, French, German, Italian, Spanish
- 最大字符
- 1000