Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Daftar untuk batas 5,000 karakter

Bungkus teks Anda dalam tag SSML untuk kendali yang tepat:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag yang dipilih mengerti klik °C untuk memasukkan satu ke dalam teks Anda di mana hal itu terjadi:

Model ini membaca teks biasa, sehingga tag inline diabaikan. Untuk tag berbasis emosi, beralih ke model ekspresif seperti Orpheus atau Bark.

Definisikan pengucapan ubahan (kata = pelafalan):

-12 +12
0.5x 2.0x
Free with Piper, VITS, Melotts
Audio yang Anda buat akan muncul di sini. Pilih model, masukkan teks, dan klik Generate.
Hasil Audio Berhasil
0:00
Unduh Audio Unduh.srt Sambungan berakhir dalam 24 jam
Tingkatan bebas: penggunaan pribadi. Ijin komersial dari $5/mo
Buatlah ini suara Anda sendiri Kloning suara dalam 30 detik
Beritahu teman-temanmu!

Tentang Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Terbaik untuk: Audiobooks, premium content, quality-first applications

Jelajahi semua Tortoise TTS suara

Pada sekilas

Pengembang
James Betker
Lisensi
Apache 2.0
Tier
premium
Kecepatan
slow
Penklonan Suara
Ya
Bahasa
English
Karakter maksimal
2000

Tortoise TTS suara

Random

English
Premium Neutral

Tortoise TTS TTS °F FAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Semua suara