Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Tia sahihi kwa kiwango cha tabia 5,000

Pakua maandishi yako katika tovuti ya SSML kwa ajili ya udhibiti sahihi:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag anaelewa mfano unaochaguliwa na unajibu ujumbe huu:

Mfano huu unasomeka maandishi rahisi, kwa hiyo alama za vidole hupuuzwa. Kwa hisia za ndani za watu, geukia kigezo kinachoonesha hisia kama Orfeus au Bark.

Matamshi ya desturi (neno = matamshi):

-12 +12
0.5x 2.0x
Nikiwa huru na Piper, VITS, MelloTTTS
Unaweza kuchagua mfano, maandishi, na kidofo kinachoitwa Genete.
Edio Iliyorekebishwa kwa Mafanikio
0:00
Paketi ya Audio Paketisha.srt Kiungo kinakufa mnamo 24
Safu huru: matumizi ya kibinafsi. Hati ya biashara kutoka dola 5/mo
Fanya hii sauti yako mwenyewe Chokoa sauti kwa sekunde 30
Waeleze rafiki zako kuhusu mapenzi ya TTS.ai?

Habari Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Bora kwa: Audiobooks, premium content, quality-first applications

Ng'ombe wote Tortoise TTS sauti

Kutupia jicho

Mbuni
James Betker
Lenzi
Apache 2.0
Tier
premium
Mwendo
slow
Kufanyizwa kwa Sauti
Ndiyo
Lugha
English
Wahusika wa Max
2000

Tortoise TTS sauti

Random

English
Premi Neutral

Tortoise TTS TTS ngumuSTEGAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Sauti zote