CosyVoice3 TTS
Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về CosyVoice3
CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.
Tốt nhất cho: Multilingual production TTS, real-time applications, voice cloning
& Xem tất cả CosyVoice3 giọng nóiMột cái nhìn
- Nhà phát triển
- Alibaba (FunAudioLLM)
- Giấy phép
- Apache 2.0
- Thú
- standard
- Tốc độ
- fast
- Ký âm
- Có
- Ngôn ngữ
- English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
- Tối đa các ký tự
- 5000