Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Prijavite se za ograničenje od 5.000 znakova

Omotajte tekst u SSML oznake za preciznu kontrolu:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Oznake koje odabrani model razumije — kliknite da biste ih ubacili u tekst gdje se pojavljuju:

Ovaj model čita običan tekst, tako da se inline oznake ignoriraju. Za emocije zasnovane na oznakama, prebacite se na ekspresivni model poput Orpheusa ili Bark-a.

Definirajte vlastite izgovore (riječ = izgovor):

-12 +12
0.5x 2.0x
Besplatno sa Piper, VITS, MeloTTS
Ovdje će se pojaviti vaš generirani audio. Izaberite model, unesite tekst i kliknite na Generiraj.
Audio uspješno generisan
0:00
Preuzmi audio Preuzmi.srt Link istječe za 24h
Free tier: personal use. Komercijalna licenca od $5/mjesečno
Volite TTS.ai?

O meni Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Najbolje za: Audiobooks, premium content, quality-first applications

Pregledaj sve Tortoise TTS glasovi

Na prvi pogled

Programer
James Betker
Licenca
Apache 2.0
Životinje
premium
Brzina
slow
Kloniranje glasa
Da.
Jezici
English
Maksimalan broj znakova
2000

Tortoise TTS glasovi

Random

English
Premium Neutral

Tortoise TTS FAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Svi glasovi