Sesame CSM TTS
A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.
Wrap uw tekst in SSML-tags voor nauwkeurige controle:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags het geselecteerde model begrijpt
Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.
Definieer aangepaste uitspraaken (woord = uitspraak):
Info Sesame CSM
Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.
Beste voor: AI assistants, chatbots, conversational AI applications
Alles doorbladeren Sesame CSM stemmenIn een oogopslag
- Ontwikkelaar
- Sesame
- Licentie
- Apache 2.0
- Niveau
- premium
- Snelheid
- slow
- Klonen van stemmen
- Nee
- Talen
- English
- Max. tekens
- 500