CosyVoice 2 TTS
Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.
Întoarceți textul în etichetele SSML pentru un control precis:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etichetele modelului selectat înțeleg — click pentru a lăsa unul în textul tău unde se întâmplă:
Acest model citește textul simplu, astfel încât etichetele inline sunt ignorate. Pentru emoții bazate pe tag, schimbați la un model expresiv cum ar fi Orpheus sau Bark.
Definiți pronunțiare personalizată (cuvânt = pronunție):
Despre CosyVoice 2
CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.
Cel mai bun pentru: Real-time applications, streaming TTS, voice assistants
Navigați toate CosyVoice 2 vociLa o privire
- Dezvoltator
- Alibaba (Tongyi Lab)
- Licență
- Apache 2.0
- Nivel
- standard
- Viteză
- medium
- Clonarea vocală
- Da.
- Limbi
- English, Chinese, Japanese, Korean, French, German, Italian, Spanish
- Caractere maxime
- 1000