CosyVoice3

CosyVoice3 TTS

Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.

_Gün tertibi 5000 karakter çäk

Metini SSML taglarda dolap dogry kontrol üçin:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Saýlanan model aňlaýan taglar — birini metinde goýmak üçin basyň:

Bu model ýönekeý metin okaýar, şonuň üçin hatda taglar gözden düşürilýär. Tag-based emotions for, switch to an expression model like Orpheus or Bark.

Öz sözleriň terjimesini belli et (söz = terjime):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS bilen azat
Siziň döreden audioňyz şu ýerde görüner. Bir model saýlaň, metin girin we döred
Ses mübärek bejerildi
0:00
Ses ýükle .srt ýükle Baglanyşyk 24 sagadyň içinde gutarýar
TTS.ai-ni söýýäňmi? Dostlaryňa aýt!

Habar CosyVoice3

CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.

Muňa iň gowy: Multilingual production TTS, real-time applications, voice cloning

Ehlini _Gözle CosyVoice3 sesler

Bir seretseň

Developer
Alibaba (FunAudioLLM)
Lisenziýa
Apache 2.0
_Göçür
standard
Tizlik
fast
Ses klonlamak
Diller
English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
Maks. karakterler
5000

CosyVoice3 sesler

Chinese Female

Chinese
_Öň bellenen Female

Chinese Male

Chinese
_Öň bellenen Male

English Female

English
_Öň bellenen Female

English Male

English
_Öň bellenen Male

French Female

French
_Öň bellenen Female

German Female

German
_Öň bellenen Female

Italian Female

Italian
_Öň bellenen Female

Japanese Female

Japanese
_Öň bellenen Female

Korean Female

Korean
_Öň bellenen Female

Russian Female

Russian
_Öň bellenen Female

Spanish Female

Spanish
_Öň bellenen Female

CosyVoice3 TTS - Gynançly Soraglar

CosyVoice3 adds bi-streaming inference at around 150ms latency, instruction-based control over emotion/speed/volume, improved speaker similarity for cloning, and coverage of 9 languages plus 18 Chinese dialects, with an RL-tuned variant for state-of-the-art prosody.

Yes. It supports zero-shot voice cloning from a reference clip (around 3 seconds minimum) with improved speaker similarity over the previous generation.

Yes. CosyVoice3 is licensed under Apache 2.0, permitting commercial use.
← Ehli Sesler