Spark TTS TTS
Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.
सटीक नियन्त्रणका लागि SSML ट्यागमा तपाईँको पाठ बेर्नुहोस्:
<speak><prosody rate="slow">Slow speech</prosody></speak>
ट्याग चयन गरिएको मोडेल बुझ्छ - यो कहाँ हुन्छ आफ्नो पाठ मा एक गिर गर्न क्लिक:
यो नमूनाले सादा पाठ पढ्दछ, त्यसैले इनलाइन ट्याग उपेक्षा गरिन्छ । ट्याग-आधारित भावनाका लागि, ओर्फिस वा बारक जस्तै अभिव्यक्तिमूलक नमूनामा स्विच गर्नुहोस् ।
अनुकूल उच्चारण परिभाषित गर्नुहोस् (शब्द = उच्चारण):
यसका बारेमा Spark TTS
Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.
यसका लागि उत्तम: Content creation with cloned voices and emotional control
सबै ब्राउज गर्नुहोस् Spark TTS आवाजएक नजरमा
- विकासकर्ता
- SparkAudio
- इजाजतपत्र
- CC BY-NC-SA 4.0
- टर
- standard
- गति
- medium
- आवाज क्लोनिङ
- हो
- भाषा
- English, Chinese
- अधिकतम क्यारेक्टर
- 1000