Dia TTS TTS
A 1.6B-parameter model purpose-built for generating natural multi-speaker dialogue, not just single-voice narration.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về Dia TTS
Dia by Nari Labs is a 1.6-billion-parameter text-to-speech model designed from the ground up for dialogue rather than monologue. It generates conversations between two speakers with realistic turn-taking, prosody, and emotional expression, producing audio that sounds like a real exchange instead of two voices read separately. Architecturally it pairs an autoregressive transformer with the Descript Audio Codec (DAC) for waveform generation. Dia is a strong fit for podcast-style content, scripted audiobook dialogue, and conversational scenes, and is released under Apache 2.0. Generations are heavier than single-voice models, so it favors quality over raw speed.
Tốt nhất cho: Podcasts, audiobook dialogues, conversational content
& Xem tất cả Dia TTS giọng nóiMột cái nhìn
- Nhà phát triển
- Nari Labs
- Giấy phép
- Apache 2.0
- Thú
- standard
- Tốc độ
- medium
- Ký âm
- Không
- Ngôn ngữ
- English
- Tối đa các ký tự
- 800