CosyVoice 2 TTS
Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.
Wrap uw tekst in SSML-tags voor nauwkeurige controle:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags het geselecteerde model begrijpt
Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.
Definieer aangepaste uitspraaken (woord = uitspraak):
Info CosyVoice 2
CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.
Beste voor: Real-time applications, streaming TTS, voice assistants
Alles doorbladeren CosyVoice 2 stemmenIn een oogopslag
- Ontwikkelaar
- Alibaba (Tongyi Lab)
- Licentie
- Apache 2.0
- Niveau
- standard
- Snelheid
- medium
- Klonen van stemmen
- Ja.
- Talen
- English, Chinese, Japanese, Korean, French, German, Italian, Spanish
- Max. tekens
- 1000