Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

תחתמי. עבור 5,000 מגבלה של תווים

לעטוף את הטקסט שלך בתגי SSML לשליטה מדויקת:

<speak><prosody rate="slow">Slow speech</prosody></speak>

שם התוויות של המודל הנבחר הוא □ לחץ כדי להפיל אחד לתוך הטקסט שלך שבו הוא קורה:

המודל הזה קורא טקסט רגיל, כך שמתעלמים מתגי ההצפנה של רגש מבוסס תג, עבור מודל אקספרסיבי כמו אורפיאוס או ברק.

הגדר הגייה מותאמת אישית (מילה = הגייה):

-12 +12
0.5x 2.0x
חינם עם פייפר, VITS, Melotts
הקול שנוצר יופיע כאן. בחר דגם, הזן טקסט, ולחץ על יצירתו.
הקול נוצר בהצלחה
0:00
הורד שמע הורדה. srt הלינק פג ב-24 שעות.
שורה חופשית: שימוש אישי. רישיון מסחרי מ-5/מו.
אוהב את ט.ט.ס.אי?

אודות Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

הטוב ביותר עבור: Audiobooks, premium content, quality-first applications

עיין בכל Tortoise TTS קולות

במבט חטוף

מפתח
James Betker
רישיון
Apache 2.0
Tier
premium
מהירות
slow
שיבוט קולי
כן.
שפות
English
תווים מרביים
2000

Tortoise TTS קולות

Random

English
פרמיום Neutral

Tortoise TTS TTS □ FAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← כל הקולות