Tortoise TTS

Tortoise TTS TTS

A quality-first autoregressive model — slow, but among the most realistic open-source speech available.

Đăng ký giới hạn 5000 ký tự

Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:

Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.

Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):

-12 +12
0.5x 2.0x
Miễn phí với Piper, VITS, MeloTTS
Âm thanh đã tạo sẽ xuất hiện ở đây. Chọn một mô hình, nhập văn bản, và nhấn vào Tạo.
Âm thanh đã được tạo thành công
0:00
Tải về âm thanh Tải về Liên kết hết hạn trong 24h
Free tier: sử dụng cá nhân. Giấy phép thương mại từ $5/mo
Cảm ơn bạn đã tin tưởng TTS.ai!

Về Tortoise TTS

Tortoise TTS, created by James Betker, deliberately trades speed for quality. It is an autoregressive multi-voice system using a DALL-E-inspired architecture, and it produces some of the most realistic synthetic speech in the open-source ecosystem, with excellent prosody and speaker similarity. The name is a nod to its pace: it is noticeably slower than most alternatives, but the payoff is studio-grade output. It supports multiple voices and voice cloning (which benefits from a longer reference, around fifteen seconds), making it a long-standing favorite for audiobooks and premium narration where wait time is acceptable. Tortoise is English-focused and released under the permissive Apache 2.0 license.

Tốt nhất cho: Audiobooks, premium content, quality-first applications

& Xem tất cả Tortoise TTS giọng nói

Một cái nhìn

Nhà phát triển
James Betker
Giấy phép
Apache 2.0
Thú
premium
Tốc độ
slow
Ký âm
Ngôn ngữ
English
Tối đa các ký tự
2000

Tortoise TTS giọng nói

Random

English
Cao cấp Neutral

Tortoise TTS TTS - FAQ

It is autoregressive and uses a DALL-E-inspired architecture that deliberately prioritizes quality over speed. The trade-off is some of the most realistic open-source speech available, which is why it remains popular for audiobooks despite the wait.

Yes. It supports multi-voice synthesis and voice cloning; results improve with a longer reference, around fifteen seconds of clean audio.

Quality-first applications — audiobooks and premium narration — where its slow but highly realistic output is worth the generation time. It is English-focused and Apache 2.0 licensed.
← Tất cả giọng nói