Spark TTS ടിടിഎസ്
Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.
കൃത്യമായ നിയന്ത്രണത്തിനായി SSML തൊങ്ങലില് വാചകം പൊതിയുക:
<speak><prosody rate="slow">Slow speech</prosody></speak>
തെരഞ്ഞെടുത്ത മാതൃക മനസ്സിലാക്കുന്നത് ടാഗ് (കുടികള്) :
ഈ മോഡ് സാധാരണ പദാവലി വായിക്കുന്നു, അതുകൊണ്ട് ഇന്ലൈന് തൊങ്ങല് അവഗണിപ്പിക്കുന്നു. ടാഗ് അടിസ്ഥാനപരമായ വികാരങ്ങള്ക്കു് ഓര്ഫിയസ് അല്ലെങ്കില് ബാര്ക് പോലുള്ള ഒരു ചിത്രീകരണ മോഡില് മാറുക.
ഇഷ്ടപ്പെട്ട ഉച്ചാരണം നിര്വ്വചിക്കുക (വാക്ക് = ഉച്ചാരണം):
സംബന്ധിച്ച് Spark TTS
Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.
അതിനു വേണ്ടിയുള്ള ഏറ്റവും നല്ല സ്ഥലം.: Content creation with cloned voices and emotional control
എല്ലാം പരതുക Spark TTS ശബ്ദങ്ങള്ഒരു നോക്കുമ്പോള്
- രചയിതാവു്
- SparkAudio
- അനുമതി
- CC BY-NC-SA 4.0
- ടിയെര്
- standard
- വേഗത
- medium
- ശബ്ദമിശ്രണോപാധി
- അതെ
- ഭാഷകള്
- English, Chinese
- ഏറ്റവും കൂടിയ ക്യാരക്ടറുകള്
- 1000