CosyVoice 2

CosyVoice 2 TTS

Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.

Zarejestruj się. dla 5000 limitów znaków

Zawiń tekst w tagi SSML dla precyzyjnej kontroli:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tagi wybrany model rozumie — kliknij aby usunąć jeden do swojego tekstu, gdzie się to dzieje:

Model ten czyta tekst zwykły, więc w linii tagi są ignorowane. Dla emocji na tag, przełącz na model wyrażony jak Orfeus lub Bark.

Definiuj własny wymówki (słowo = wymówka):

-12 +12
0.5x 2.0x
Darmowe z Piper, VITS, Melotts
Tutaj pojawi się generowany dźwięk. Wybierz model, wpisz tekst i kliknij Generuj.
Pomyślnie wygenerowany dźwięk
0:00
Pobierz audio Pobierz.rt Łączność wygasa w 24h
Bezpłatny poziom: użytkowanie osobiste. Licencja handlowa od $5/mo
Powiedz znajomym!

O tematie CosyVoice 2

CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.

Najlepsze dla: Real-time applications, streaming TTS, voice assistants

Przeglądaj wszystkie CosyVoice 2 głosy

Na jedno spojrzenie

Rozwijacz
Alibaba (Tongyi Lab)
Licencja
Apache 2.0
Poziom szczelności
standard
Prędkość
medium
Klonowanie głosu
Tak.
Języki
English, Chinese, Japanese, Korean, French, German, Italian, Spanish
Maksymalna liczba znaków
1000

CosyVoice 2 głosy

Chinese Female

Chinese
Standardowe Female

Chinese Male

Chinese
Standardowe Male

English Female

English
Standardowe Female

English Male

English
Standardowe Male

French Female

French
Standardowe Female

German Female

German
Standardowe Female

Italian Female

Italian
Standardowe Female

Japanese Female

Japanese
Standardowe Female

Korean Female

Korean
Standardowe Female

Spanish Female

Spanish
Standardowe Female

CosyVoice 2 TTS — FAQ

Yes. CosyVoice 2 uses finite scalar quantization for streaming synthesis at very low latency, which is what makes it suitable for voice assistants and real-time applications.

Yes. It offers zero-shot voice cloning from roughly 3 seconds of reference audio, plus cross-lingual synthesis and emotion control.

Yes. CosyVoice 2 is Apache 2.0 licensed. It supports 8 languages: English, Chinese, Japanese, Korean, French, German, Italian, and Spanish.
← Wszystkie głosy