Sesame CSM TTS
A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về Sesame CSM
Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.
Tốt nhất cho: AI assistants, chatbots, conversational AI applications
& Xem tất cả Sesame CSM giọng nóiMột cái nhìn
- Nhà phát triển
- Sesame
- Giấy phép
- Apache 2.0
- Thú
- premium
- Tốc độ
- slow
- Ký âm
- Không
- Ngôn ngữ
- English
- Tối đa các ký tự
- 500