CosyVoice3

CosyVoice3 TTS

Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.

_Gün tertibi 5000 karakter çäk

Metini SSML taglarda dolap dogry kontrol üçin:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Saýlanan model aňlaýan taglar — birini metinde goýmak üçin basyň:

Bu model ýönekeý metin okaýar, şonuň üçin hatda taglar gözden düşürilýär. Tag-based emotions for, switch to an expression model like Orpheus or Bark.

Öz sözleriň terjimesini belli et (söz = terjime):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS bilen azat
Siziň döreden audioňyz şu ýerde görüner. Bir model saýlaň, metin girin we döred
Ses mübärek bejerildi
0:00
Ses ýükle .srt ýükle Baglanyşyk 24 sagadyň içinde gutarýar
TTS.ai-ni söýýäňmi? Dostlaryňa aýt!

Habar CosyVoice3

CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.

Muňa iň gowy: Multilingual production TTS, real-time applications, voice cloning

Ehlini _Gözle CosyVoice3 sesler

Bir seretseň

Developer
Alibaba (FunAudioLLM)
Lisenziýa
Apache 2.0
_Göçür
standard
Tizlik
fast
Ses klonlamak
Diller
English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
Maks. karakterler
5000

CosyVoice3 sesler

Chinese Female

Chinese
_Öň bellenen Female

Chinese Male

Chinese
_Öň bellenen Male

English Female

English
_Öň bellenen Female

English Male

English
_Öň bellenen Male

French Female

French
_Öň bellenen Female

German Female

German
_Öň bellenen Female

Italian Female

Italian
_Öň bellenen Female

Japanese Female

Japanese
_Öň bellenen Female

Korean Female

Korean
_Öň bellenen Female

Russian Female

Russian
_Öň bellenen Female

Spanish Female

Spanish
_Öň bellenen Female

CosyVoice3 TTS - Gynançly Soraglar

CosyVoice3 adds bi-streaming inference at around 150ms latency, instruction-based control over emotion/speed/volume, improved speaker similarity for cloning, and coverage of 9 languages plus 18 Chinese dialects, with an RL-tuned variant for state-of-the-art prosody.

Yes. It supports zero-shot voice cloning from a reference clip (around 3 seconds minimum) with improved speaker similarity over the previous generation.

CosyVoice3 is licensed under Apache 2.0. On TTS.ai, audio you make with it is for personal use on the free tier and can be used commercially on any paid plan.
← Ehli Sesler