GPT-SoVITS TTS
A few-shot voice cloning model that replicates a voice — and can even sing — from as little as five seconds of audio.
Wrap wanu malemba mu SSML tags kwa kuwongolera moyenera:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags chosankhidwa chitsanzo amamvetsa - dinani kuti aphe mmodzi m'mawu anu pamene chimachitika:
Izi ndi njira yolemba malemba oyera, kotero ma tag ophatikizidwa amasiya kuganiziridwa. Kuti mupange ma tag ogwirizana ndi maganizo, gwiritsani ntchito njira yolemba malemba monga Orpheus kapena Bark.
Define custom pronunciations (word = pronunciation):
Za GPT-SoVITS
GPT-SoVITS, created by the developer known as RVC-Boss, combines GPT-style language modeling with SoVITS (Singing Voice Conversion / synthesis) to deliver some of the most accessible voice cloning in open source. With as little as five seconds of reference audio it captures a speaker's timbre and style, and it stands out from most TTS models in handling singing as well as speech. It works across English, Chinese, Japanese, and Korean and supports cross-lingual generation, so a cloned voice can speak a language the reference clip never used. It is widely used by content creators for voice replication, dubbing, and song covers, and reaches high fidelity for a model of its size.
Best kwa: Voice cloning, singing synthesis, content creator voice replication
Pezani zonse GPT-SoVITS maganizoPa mphindi
- Wopanga
- RVC-Boss
- License
- MIT
- Mtundu
- standard
- Kuyenda
- slow
- Kusintha kwa mawu
- Yes
- Zilankhulo
- English, Chinese, Japanese, Korean
- Max characters
- 500