GPT-SoVITS

GPT-SoVITS TTS

A few-shot voice cloning model that replicates a voice — and can even sing — from as little as five seconds of audio.

Skráðu þig inn fyrir 5.000 stafa takmörk

Wrap texta í SSML tags fyrir nákvæma stjórn:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Merki sem valið líkan skilur — smelltu til að sleppa einu í textann þinn þar sem það gerist:

Þetta líkan les venjulegan texta, þannig að innlínumerki eru hunsuð. Fyrir merki sem byggja á tilfinningum, skiptu yfir í tjáningarlíkan eins og Orpheus eða Bark.

Skilgreindu sérsniðna framburð (orð = framburð):

-12 +12
0.5x 2.0x
Frjáls með Piper, VITS, MeloTTS
Hljóðskráin þín birtist hér. Veldu líkan, sláðu inn texta og smelltu á Búa til.
Hljóð búið til
0:00
Sækja hljóð Sækja.srt Tengill rennur út eftir 24 klst
Frjáls tier: persónuleg notkun. Viðskiptaleyfi frá $ 5 / mánuði
Elska TTS.ai? Segðu vinum þínum!

Um GPT-SoVITS

GPT-SoVITS, created by the developer known as RVC-Boss, combines GPT-style language modeling with SoVITS (Singing Voice Conversion / synthesis) to deliver some of the most accessible voice cloning in open source. With as little as five seconds of reference audio it captures a speaker's timbre and style, and it stands out from most TTS models in handling singing as well as speech. It works across English, Chinese, Japanese, and Korean and supports cross-lingual generation, so a cloned voice can speak a language the reference clip never used. It is widely used by content creators for voice replication, dubbing, and song covers, and reaches high fidelity for a model of its size.

Best fyrir: Voice cloning, singing synthesis, content creator voice replication

Skoða allt GPT-SoVITS raddir

Í hnotskurn

Forritari
RVC-Boss
Leyfi
MIT
Tími
standard
Hraði
slow
Raddklóðun
Tungumál
English, Chinese, Japanese, Korean
Hámarksstafir
500

GPT-SoVITS raddir

Default

Chinese
Sjálfgefið Neutral

English Default

English
Sjálfgefið Neutral

Japanese Default

Japanese
Sjálfgefið Neutral

Korean Default

Korean
Sjálfgefið Neutral

GPT-SoVITS TTS — Algengar spurningar

As little as five seconds. It uses few-shot learning, so a short reference clip is enough to capture a speaker, though a cleaner and slightly longer sample improves similarity.

Yes. Its SoVITS lineage comes from singing voice synthesis, so unlike most TTS models it can generate singing as well as spoken voice, which is why it is popular for song covers.

English, Chinese, Japanese, and Korean, with cross-lingual synthesis — a voice cloned from one language can be made to speak the others.
← Allar raddir