CosyVoice 2

CosyVoice 2 TTS

Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.

Langganan for 5,000 characters limit

Ngresiki teks ing tag SSML kanggo kontrol presisi:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag kang dipahami model kang dipilih - klik kanggo ngethok siji ing teks sampeyan ing ngendi iku kedadeyan:

Model iki maca teks biasa, mula tag ing baris diabaikan. Kanggo emosi berbasis tag, ganti menyang model ekspresif kaya Orpheus utawa Bark.

Nyathet tembung-tembung standar (kata = tembung):

-12 +12
0.5x 2.0x
Bebas karo Piper, VITS, MeloTTS
Audio sing digawé bakal katon ing kene. Pilih modél, ketik teks, lan pencet Ngembangaké.
Audio Digawé kanthi Sukses
0:00
Unduh Audio Download.srt Link expires in 24h
Ing basa Indonésia, iku tegesé: pribadi. Lisénsi komersial saka $5/mo
TTS.ai? Nyathet kanca-kancamu!

Ngendi CosyVoice 2

CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.

Paling apik kanggo: Real-time applications, streaming TTS, voice assistants

Jajal kabeh CosyVoice 2 swara

Ing cetha

Pangembang
Alibaba (Tongyi Lab)
Lisénsi
Apache 2.0
Tanggal
standard
Kecepatan
medium
Kloning swara
Ya
Basa
English, Chinese, Japanese, Korean, French, German, Italian, Spanish
Aksara paling akèh
1000

CosyVoice 2 swara

Chinese Female

Chinese
Standar Female

Chinese Male

Chinese
Standar Male

English Female

English
Standar Female

English Male

English
Standar Male

French Female

French
Standar Female

German Female

German
Standar Female

Italian Female

Italian
Standar Female

Japanese Female

Japanese
Standar Female

Korean Female

Korean
Standar Female

Spanish Female

Spanish
Standar Female

CosyVoice 2 FAQ

Yes. CosyVoice 2 uses finite scalar quantization for streaming synthesis at very low latency, which is what makes it suitable for voice assistants and real-time applications.

Yes. It offers zero-shot voice cloning from roughly 3 seconds of reference audio, plus cross-lingual synthesis and emotion control.

Yes. CosyVoice 2 is Apache 2.0 licensed. It supports 8 languages: English, Chinese, Japanese, Korean, French, German, Italian, and Spanish.
← Sekabehing swara