CosyVoice 2

CosyVoice 2 TTS

Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.

Aliĝi for 5, 000 character limit

Envolvu vian tekston en SSML- etikedojn por preciza kontrolo:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etikedoj kiujn la elektita modelo komprenas - klaku por meti unu en vian tekston kie ĝi okazas:

This model reads plain text, so inline tags are ignored. For tag-based emotion, switch to an expressive model like Orpheus or Bark.

Difini proprajn elparolojn (vorto = elparolo):

-12 +12
0.5x 2.0x
Libera kun Piper, VITS, MeloTTS
Via generita sono aperos tie ĉi. Elektu modelon, entajpu tekston, kaj alklaku Generi.
Sondosiero sukcese generita
0:00
Elŝuti sonon Elŝuti.srt Ligo eksvalidiĝas post 24 horoj
Libera programaro: persona uzo. Komerca licenco ekde $5/mo
Ĉu vi ŝatas TTS.ai? Diru al viaj amikoj!

Pri CosyVoice 2

CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.

Plej bona por: Real-time applications, streaming TTS, voice assistants

Foliumi ĉiujn CosyVoice 2 voĉoj

Unu rigardo

Programisto
Alibaba (Tongyi Lab)
Licenco
Apache 2.0
Tamuz
standard
Rapideco
medium
Voĉo- klonado
Jes
Lingvoj
English, Chinese, Japanese, Korean, French, German, Italian, Spanish
Maksimuma nombro da signoj
1000

CosyVoice 2 voĉoj

Chinese Female

Chinese
Defaŭlta Female

Chinese Male

Chinese
Defaŭlta Male

English Female

English
Defaŭlta Female

English Male

English
Defaŭlta Male

French Female

French
Defaŭlta Female

German Female

German
Defaŭlta Female

Italian Female

Italian
Defaŭlta Female

Japanese Female

Japanese
Defaŭlta Female

Korean Female

Korean
Defaŭlta Female

Spanish Female

Spanish
Defaŭlta Female

CosyVoice 2 TTS - FAQ

Yes. CosyVoice 2 uses finite scalar quantization for streaming synthesis at very low latency, which is what makes it suitable for voice assistants and real-time applications.

Yes. It offers zero-shot voice cloning from roughly 3 seconds of reference audio, plus cross-lingual synthesis and emotion control.

Yes. CosyVoice 2 is Apache 2.0 licensed. It supports 8 languages: English, Chinese, Japanese, Korean, French, German, Italian, and Spanish.
← Ĉiuj voĉoj