Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Regisztrálj! 5000 karakterhatárra

Írja be a szöveget az SSML címkékbe a pontos vezérlés érdekében:

<speak><prosody rate="slow">Slow speech</prosody></speak>

A kijelölt modell tagjeiben a ~ kattintson az alábbi szövegbe:

Ez a modell egyszerű szöveget olvas, így a sorcímke figyelmen kívül marad. A tag alapú érzelem, váltson át egy expresszív modell, mint Orpheus vagy Bark.

Definiáld az egyéni kiejtéseket (szó = kiejtés):

-12 +12
0.5x 2.0x
Szabad Piper, VITS, MelotTS
A generált audio jelenik meg itt. Válasszon ki egy modellt, írja be a szöveget, és kattintson a Generate gombra.
Audio generált sikeresen
0:00
Audio letöltése Letöltés.srt A kapcsolat 24 órán belül lejár
Ingyenes szint: személyes használat. Kereskedelmi engedély 5 dollárról
Mondd el a barátaidnak!

About Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Legjobb: Audiobooks, premium content, quality-first applications

Összes böngészés Tortoise TTS hangok

Egy pillantásra

Fejlesztő
James Betker
Jogosítvány
Apache 2.0
Tier
premium
Sebesség
slow
Hang klónozása
Igen.
Nyelvek
English
Max. karakterek
2000

Tortoise TTS hangok

Random

English
Prémium Neutral

Tortoise TTS TTS - FAQCharselect unicode block name (optional, probably does not need a translation)

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Minden hang