Sesame CSM TTS
A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.
Enveloppez votre texte dans des balises SSML pour un contrôle précis :
<speak><prosody rate="slow">Slow speech</prosody></speak>
Mots clés le modèle sélectionné comprend — cliquez pour en déposer un dans votre texte où il se produit:
Ce modèle lit du texte clair, donc les balises en ligne sont ignorées. Pour l'émotion basée sur les tags, passer à un modèle expressif comme Orphée ou Bark.
Définir les prononciations personnalisées (mot = prononciation) :
À propos Sesame CSM
Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.
Meilleur pour: AI assistants, chatbots, conversational AI applications
Tout voir Sesame CSM voixEn un coup d'oeil
- Développeur
- Sesame
- Licence
- Apache 2.0
- Niveau
- premium
- Régime
- slow
- Closonnage de la voix
- Numéro
- Langues
- English
- Personnages maxi
- 500