GPT-SoVITS

GPT-SoVITS TTS

A few-shot voice cloning model that replicates a voice — and can even sing — from as little as five seconds of audio.

Prihlásiť sa na odber Limit 5 000 znakov

Zabaliť text do SSML značiek pre presnú kontrolu:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Značky, ktorým vybraný model rozumie — kliknutím ich umiestnite do textu tam, kde sa vyskytujú:

Tento model číta obyčajný text, takže vnorené značky sa ignorujú.Pre emócie založené na značkách prejdite na expresívny model ako Orpheus alebo Bark.

Definovať vlastné výslovnosti (slovo = výslovnosť):

-12 +12
0.5x 2.0x
Zadarmo s Piper, VITS, MeloTTS
Vyberte si model, zadajte text a kliknite na tlačidlo Generovať.Generate.
Audio generované úspešne
0:00
Stiahnuť audio na stiahnutie Stiahnuť.srt súbor Platnosť odkazu vyprší za 24h
Bezplatná úroveň: osobné použitie. Komerčná licencia od 5 USD/mesiac
Láska TTS.ai? Povedzte svojim priateľom!

O nás GPT-SoVITS

GPT-SoVITS, created by the developer known as RVC-Boss, combines GPT-style language modeling with SoVITS (Singing Voice Conversion / synthesis) to deliver some of the most accessible voice cloning in open source. With as little as five seconds of reference audio it captures a speaker's timbre and style, and it stands out from most TTS models in handling singing as well as speech. It works across English, Chinese, Japanese, and Korean and supports cross-lingual generation, so a cloned voice can speak a language the reference clip never used. It is widely used by content creators for voice replication, dubbing, and song covers, and reaches high fidelity for a model of its size.

Najlepšie pre: Voice cloning, singing synthesis, content creator voice replication

Prehľadávať všetky GPT-SoVITS hlasy

Na prvý pohľad

Vývojár
RVC-Boss
Licencia
MIT
Zvieratá
standard
Rýchlosť
slow
Klonovanie hlasu
Áno
Jazyky
English, Chinese, Japanese, Korean
Max. počet znakov
500

GPT-SoVITS hlasy

Default

Chinese
Štandardné Neutral

English Default

English
Štandardné Neutral

Japanese Default

Japanese
Štandardné Neutral

Korean Default

Korean
Štandardné Neutral

GPT-SoVITS TTS — Často kladené otázky

As little as five seconds. It uses few-shot learning, so a short reference clip is enough to capture a speaker, though a cleaner and slightly longer sample improves similarity.

Yes. Its SoVITS lineage comes from singing voice synthesis, so unlike most TTS models it can generate singing as well as spoken voice, which is why it is popular for song covers.

English, Chinese, Japanese, and Korean, with cross-lingual synthesis — a voice cloned from one language can be made to speak the others.
← Všetky hlasy