StyleTTS 2 TTS
Reaches human-level single-speaker synthesis through style diffusion and adversarial training.
Envolva o seu texto em tags SSML para controle preciso:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etiquetas o modelo selecionado entende — clique para soltar um para o seu texto onde acontece:
Este modelo lê texto simples, por isso as etiquetas inline são ignoradas. Para emoção baseada em tags, mude para um modelo expressivo como Orpheus ou Bark.
Definir pronúncias personalizadas (palavra = pronúncia):
Sobre StyleTTS 2
StyleTTS 2, developed at Columbia University, achieves human-level text-to-speech for single-speaker synthesis by combining style diffusion with adversarial training guided by large speech language models. Its diffusion-based style modeling captures the full natural variation of human speech — subtle shifts in rhythm, emphasis, and tone — so output can rival real recordings. It is widely regarded as one of the most natural-sounding open single-speaker models, which makes it a strong choice for studio-quality narration and professional voiceover where polish matters more than cloning or multilingual range. StyleTTS 2 is English-focused and released under the permissive MIT license.
Melhor para: Studio-quality single-speaker synthesis, professional narration
Procurar todos StyleTTS 2 vozesDe uma olhada
- Desenvolvedor
- Columbia University
- Licença
- MIT
- Tier
- premium
- Velocidade
- medium
- Clonagem de voz
- Não
- Línguas
- English
- Número máximo de caracteres
- 500