CosyVoice3 TTS
Alibaba FunAudioLLM's latest multilingual model with ~150ms bi-streaming, instruction control, and zero-shot cloning.
ସଠିକ ନିୟନ୍ତ୍ରଣ ପାଇଁ SSML ଟ୍ୟାଗଗୁଡ଼ିକରେ ଆପଣଙ୍କର ପାଠ୍ୟକୁ ଲଗାନ୍ତୁ:
<speak><prosody rate="slow">Slow speech</prosody></speak>
ବଚ୍ଛିତ ନମୁନା ବୁଝିଥିବା ସୂଚକଗୁଡ଼ିକ - ଏହା ଘଟୁଥିବା ସ୍ଥାନକୁ ଆପଣଙ୍କର ପାଠ୍ୟରେ ଗୋଟିଏ ପକାଇବା ପାଇଁ କ୍ଲିକ କରନ୍ତୁ:
ଏହି ଆକାର ସରଳ ପାଠ୍ୟ ପଢ଼େ, ତେଣୁ ଅନ୍ତର୍ନିହିତ ସୂଚକଗୁଡ଼ିକୁ ଅଣଦେଖା କରାଯାଏ। ସୂଚକ ଆଧାରିତ ଅନୁଭୂତି ପାଇଁ, ଗୋଟିଏ ଅଭିବ୍ୟକ୍ତିମୂଳକ ଆକାରକୁ ପରିବର୍ତ୍ତନ କରନ୍ତୁ ଯେପରିକି Orpheus କିମ୍ବା Bark।
ଇଚ୍ଛାରୂପୀ ଉଚ୍ଚାରଣକୁ ବର୍ଣ୍ଣନା କରନ୍ତୁ (ଶବ୍ଦ = ଉଚ୍ଚାରଣ):
ବିଷୟରେ CosyVoice3
CosyVoice3 is the newest generation from Alibaba's FunAudioLLM team and a clear step up from CosyVoice 2. It introduces bi-streaming inference with roughly 150ms latency and instruction-based control, letting you steer emotion, speed, and volume through prompts. Speaker similarity for zero-shot voice cloning is improved, and coverage spans 9 languages plus 18 Chinese dialects. An RL-tuned variant pushes prosody to a state-of-the-art level. With a 5,000-character ceiling, fast generation, and strong cloning, it's geared toward multilingual production TTS and real-time applications.
ପାଇଁ ଉତ୍ତମ: Multilingual production TTS, real-time applications, voice cloning
ସମସ୍ତଙ୍କୁ ବ୍ରାଉଜ କରନ୍ତୁ CosyVoice3 ଧ୍ୱନିଗୋଟିଏ ନଜରରେ
- ବିକାଶକାରୀ
- Alibaba (FunAudioLLM)
- ଅନୁମତିପତ୍ର
- Apache 2.0
- ତିଆର
- standard
- ବେଗ
- fast
- ଧ୍ୱନି କ୍ଲୋନିଂ
- ହଁ
- ଭାଷାName
- English, Chinese, Japanese, Korean, German, Spanish, French, Italian, Russian
- ସର୍ବାଧିକ ଅକ୍ଷର
- 5000