CosyVoice 2 TTS
Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về CosyVoice 2
CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.
Tốt nhất cho: Real-time applications, streaming TTS, voice assistants
& Xem tất cả CosyVoice 2 giọng nóiMột cái nhìn
- Nhà phát triển
- Alibaba (Tongyi Lab)
- Giấy phép
- Apache 2.0
- Thú
- standard
- Tốc độ
- medium
- Ký âm
- Có
- Ngôn ngữ
- English, Chinese, Japanese, Korean, French, German, Italian, Spanish
- Tối đa các ký tự
- 1000