Sesame CSM TTS
A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.
Wrap ou tèks nan SSML tags pou presizyon kontwòl:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags ke modèl la chwazi konprann — klike pou mete yon nan tèks ou kote li rive:
Modèl sa a li tèks senp, se poutèt sa atik ki nan liy yo pa pran an kont. Pou efè ki baze sou atik, chanje pou yon modèl ekspresyon tankou Orpheus oswa Bark.
Define prononciations Custom (mot = prononciation):
Atik Sesame CSM
Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.
Pi bon pou: AI assistants, chatbots, conversational AI applications
Navigue tout Sesame CSM VoyYon ti gade
- Pwogramè
- Sesame
- Lisans
- Apache 2.0
- Nivo
- premium
- Vitès
- slow
- Klonaj vwa
- Non
- Lang
- English
- Karakteris maksimòm
- 500