GPT-SoVITS

GPT-SoVITS TT TT TTT TTT T TT TT T T TTT TTT TTT

A few-shot voice cloning model that replicates a voice — and can even sing — from as little as five seconds of audio.

签名 对 5,000 字符限制的 5 000 个字符

在 SSML 标记中折行文本以精确控制 :

<speak><prosody rate="slow">Slow speech</prosody></speak>

标记选中模式的理解度 - 单击将一个输入到文本中, 发生时 :

这个模型读的是简单的文字, 所以内嵌标签会被忽略。 对于基于标签的情感, 请切换到像 Orpheus 或 Bark 这样的表达模式 。

定义自定义发音( Word = 发音) :

-12 +12
0.5x 2.0x
免费的管道、VITS、MelotTS
您生成的音频将在此显示。 选择一个模型, 输入文本, 并单击生成 。
音频生成成功
0:00
下载音频 下载.strt 24小时后链接过期
免费:个人使用。 5美元/美元商业许可证
使这个声音成为你自己的声音 30秒后打开声音
喜欢TTS.ai吗?告诉你的朋友吧!

关于 GPT-SoVITS

GPT-SoVITS, created by the developer known as RVC-Boss, combines GPT-style language modeling with SoVITS (Singing Voice Conversion / synthesis) to deliver some of the most accessible voice cloning in open source. With as little as five seconds of reference audio it captures a speaker's timbre and style, and it stands out from most TTS models in handling singing as well as speech. It works across English, Chinese, Japanese, and Korean and supports cross-lingual generation, so a cloned voice can speak a language the reference clip never used. It is widely used by content creators for voice replication, dubbing, and song covers, and reaches high fidelity for a model of its size.

最佳: Voice cloning, singing synthesis, content creator voice replication

全部浏览 GPT-SoVITS 声音

一眼看一眼,

开发者
RVC-Boss
许可证
MIT
级别
standard
速度
slow
语音克隆
语言
English, Chinese, Japanese, Korean
最大字符
500

GPT-SoVITS 声音

Default

Chinese
标准 Neutral

English Default

English
标准 Neutral

Japanese Default

Japanese
标准 Neutral

Korean Default

Korean
标准 Neutral

GPT-SoVITS TTS - 常见问题

As little as five seconds. It uses few-shot learning, so a short reference clip is enough to capture a speaker, though a cleaner and slightly longer sample improves similarity.

Yes. Its SoVITS lineage comes from singing voice synthesis, so unlike most TTS models it can generate singing as well as spoken voice, which is why it is popular for song covers.

English, Chinese, Japanese, and Korean, with cross-lingual synthesis — a voice cloned from one language can be made to speak the others.
← 所有声音