CosyVoice3

CosyVoice3 음성 인식

Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.

가입하기 5,000자 한도

정확한 제어를 위해 SSML 태그로 텍스트를 래핑하십시오.

<speak><prosody rate="slow">Slow speech</prosody></speak>

선택한 모델이 이해하는 태그 — 텍스트에 드래그하려면 클릭하세요:

이 모델은 일반 텍스트를 읽기 때문에 인라인 태그는 무시됩니다. 태그 기반 감정을 위해서는 Orpheus 또는 Bark과 같은 표현 모델로 전환하십시오.

사용자 지정 발음 정의 (단어 = 발음):

-12 +12
0.5x 2.0x
파이퍼, VITS, MeloTTS와 무료
생성된 오디오가 여기에 나타납니다. 모델을 선택하고 텍스트를 입력한 다음 생성 을 클릭합니다.
오디오가 성공적으로 생성되었습니다
0:00
오디오 다운로드 .srt 파일 다운로드 링크는 24시간 이내에 만료됩니다.
무료 계층: 개인용. 상업용 라이센스 최저 $5/mo
이것을 당신의 목소리로 만들어라 30초만에 목소리 복제
TTS.ai가 마음에 드시나요? 친구들에게 알려주세요!

정보 CosyVoice3

CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.

최적화된 용도: Multilingual production TTS, real-time applications, voice cloning

모두 찾아보기 CosyVoice3 목소리

한눈에

개발자
Alibaba (FunAudioLLM)
라이선스
Apache 2.0
standard
속도
fast
음성 복제
언어
English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
최대 문자수
5000

CosyVoice3 목소리

Chinese Female

Chinese
표준 Female

Chinese Male

Chinese
표준 Male

English Female

English
표준 Female

English Male

English
표준 Male

French Female

French
표준 Female

German Female

German
표준 Female

Italian Female

Italian
표준 Female

Japanese Female

Japanese
표준 Female

Korean Female

Korean
표준 Female

Russian Female

Russian
표준 Female

Spanish Female

Spanish
표준 Female

CosyVoice3 TTS — 자주 묻는 질문

CosyVoice3 adds bi-streaming inference at around 150ms latency, instruction-based control over emotion/speed/volume, improved speaker similarity for cloning, and coverage of 9 languages plus 18 Chinese dialects, with an RL-tuned variant for state-of-the-art prosody.

Yes. It supports zero-shot voice cloning from a reference clip (around 3 seconds minimum) with improved speaker similarity over the previous generation.

Yes. CosyVoice3 is licensed under Apache 2.0, permitting commercial use.
← 모든 음성