CosyVoice3

CosyVoice3 1 - تكنولوجيا المعلومات والاتصالات

Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.

انضم 000 5 كلمة

لف نصك في علامات SSML للتحكم الدقيق:

<speak><prosody rate="slow">Slow speech</prosody></speak>

العلامات التي يفهمها النموذج المختار — انقر لإسقاط واحدة في نصك حيث تحصل:

هذا النموذج يقرأ النص العادي، لذلك يتم تجاهل العلامات في السطر. للتعبير عن المشاعر القائمة على العلامات، انتقل إلى نموذج تعبيري مثل أورفيوس أو بارك.

تعريف النطق العادي (كلمة = نطق):

-12 +12
0.5x 2.0x
مجاني مع Piper, VITS, MeloTTS
سيظهر الصوت الذي أنتجته هنا. اختر نموذجاً، وأدخل نصاً، ثم انقر على توليد.
تم توليد الصوت بنجاح
0:00
تنزيل الصوت تنزيل.srt الرابط ينتهي بعد 24 ساعة
المستوى المجاني: الاستخدام الشخصي. ترخيص تجاري من 5 دولارات شهريا
أحب TTS.ai؟ أخبر أصدقائك!

حول CosyVoice3

CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.

أفضل لل: Multilingual production TTS, real-time applications, voice cloning

تصفح جميع CosyVoice3 الأصوات

لمحة عامة

مطوِّر
Alibaba (FunAudioLLM)
الترخيص
Apache 2.0
الرتبة
standard
السرعة
fast
استنساخ الصوت
نعم
اللغات
English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
الحد الأقصى للحروف
5000

CosyVoice3 الأصوات

Chinese Female

Chinese
المعيار Female

Chinese Male

Chinese
المعيار Male

English Female

English
المعيار Female

English Male

English
المعيار Male

French Female

French
المعيار Female

German Female

German
المعيار Female

Italian Female

Italian
المعيار Female

Japanese Female

Japanese
المعيار Female

Korean Female

Korean
المعيار Female

Russian Female

Russian
المعيار Female

Spanish Female

Spanish
المعيار Female

CosyVoice3 الأسئلة المتكررة

CosyVoice3 adds bi-streaming inference at around 150ms latency, instruction-based control over emotion/speed/volume, improved speaker similarity for cloning, and coverage of 9 languages plus 18 Chinese dialects, with an RL-tuned variant for state-of-the-art prosody.

Yes. It supports zero-shot voice cloning from a reference clip (around 3 seconds minimum) with improved speaker similarity over the previous generation.

Yes. CosyVoice3 is licensed under Apache 2.0, permitting commercial use.
← جميع الأصوات