CosyVoice 2

CosyVoice 2 TTS

Alibaba Tongyi Lab's streaming TTS reaching human-parity naturalness with near-zero latency and zero-shot cloning.

Ojejapo 5000 caracter rehegua límite

Ojehaijey ñe'ẽnguéra etiquetas SSML-pe peteĩ control hekopete g̃uarã:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas ohechakuaáva modelo ojeporavóva - tesãirã peteĩ peteĩva ñe'ẽnguérape, oĩhápe:

Ko modelo ohai texto ndahasyivéva, upévare umi etiqueta oĩva línea ryepýpe ndojehechakuaái. Umi emoción oñemopyendáva etiqueta-pe g̃uarã, oñemoambue peteĩ modelo expresivo-pe taha'e Orfeo térã Bark.

Oñemohenda ñe'ẽnguéra ojehechapyréva (tembiapo = ñe'ẽnguéra):

-12 +12
0.5x 2.0x
Libre Piper, VITS, MeloTTS ndive
Audio-kuéra oguenohẽva ojekuaauka ko'ápe. Oñeporavo peteĩ modelo, omoĩnge ñe'ẽ ha ohesa'ỹijo Generar.
Audio oñemoheñói porã
0:00
Oñeguenohẽ marandu myambue guive Oñeguenohẽ.srt Ko enlace hi'are 24 h rire
Nivel libre: jeiporu personal. Licencia comercial $5/ha'e rupi
Ehayhuetéva TTS.ai? He'i umi iñangirũpe!

Mba'épa CosyVoice 2

CosyVoice 2, from Alibaba's Tongyi Lab, was designed to make high-quality speech viable in real time. It uses a finite scalar quantization approach combined with flow matching to support streaming synthesis at extremely low latency, while reaching human-comparable naturalness that outperforms many commercial systems in subjective tests. Beyond quality, it offers zero-shot voice cloning from about 3 seconds of audio, cross-lingual synthesis, and fine-grained emotion control. Covering 8 languages with a 1,000-character cap, it's a strong fit for voice assistants, streaming TTS, and other real-time applications.

Oñeha'ãvéva: Real-time applications, streaming TTS, voice assistants

Ojehecha opavave CosyVoice 2 ñe'ẽ

Peteĩ jehecha

Desarrollador
Alibaba (Tongyi Lab)
Licencia
Apache 2.0
Ta'ãnga
standard
Velocidad
medium
Clonación ñe'ẽnguéra rehe
Ha'e
Ñe'ẽ
English, Chinese, Japanese, Korean, French, German, Italian, Spanish
Caracteres máx.
1000

CosyVoice 2 ñe'ẽ

Chinese Female

Chinese
Estándar Female

Chinese Male

Chinese
Estándar Male

English Female

English
Estándar Female

English Male

English
Estándar Male

French Female

French
Estándar Female

German Female

German
Estándar Female

Italian Female

Italian
Estándar Female

Japanese Female

Japanese
Estándar Female

Korean Female

Korean
Estándar Female

Spanish Female

Spanish
Estándar Female

CosyVoice 2 Pregunta ojehechavéva

Yes. CosyVoice 2 uses finite scalar quantization for streaming synthesis at very low latency, which is what makes it suitable for voice assistants and real-time applications.

Yes. It offers zero-shot voice cloning from roughly 3 seconds of reference audio, plus cross-lingual synthesis and emotion control.

Yes. CosyVoice 2 is Apache 2.0 licensed. It supports 8 languages: English, Chinese, Japanese, Korean, French, German, Italian, and Spanish.
← Opaite ñe'ẽ