Tortoise TTS

Tortoise TTS ТТС

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Подписывайся. для 5 000 символов

Заверните текст в SSML для точного контроля:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Помечает выбранную модель, которая понимает — нажмите, чтобы выкинуть одну в ваш текст, где она случается:

Эта модель читает простой текст, поэтому встраиваемые метки игнорируются. Для эмоций на основе метки переключайтесь на экспрессивную модель, как Орфей или Барк.

Определить традиционные произношения (слово = произношение):

-12 +12
0.5x 2.0x
Бесплатно с Пайпер, VITS, MeloTTS
Ваш генерированный звук появится здесь. Выберите модель, введите текст и нажмите на Генератор.
Аудиовизна успешно генерирована
0:00
Загрузить звук Загрузка. st Ссылка истекает в 24 ч.
Свободный уровень: личное использование. Коммерческая лицензия из 5 долл. США/мо
Нравится TTS.ai? Расскажите друзьям!

О том, что Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Лучший для: Audiobooks, premium content, quality-first applications

Просмотр Tortoise TTS голоса

Взгляните.

Разработчик
James Betker
Лицензия
Apache 2.0
Тяжелый
premium
Скорость
slow
Клонирование голоса
Выполнено
Знание языков
English
Максимум символов
2000

Tortoise TTS голоса

Random

English
Премиум Neutral

Tortoise TTS TTS - FAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Все голоса