Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

Ojejapo 5000 caracter rehegua límite

Ojehaijey ñe'ẽnguéra etiquetas SSML-pe peteĩ control hekopete g̃uarã:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas ohechakuaáva modelo ojeporavóva - tesãirã peteĩ peteĩva ñe'ẽnguérape, oĩhápe:

Ko modelo ohai texto ndahasyivéva, upévare umi etiqueta oĩva línea ryepýpe ndojehechakuaái. Umi emoción oñemopyendáva etiqueta-pe g̃uarã, oñemoambue peteĩ modelo expresivo-pe taha'e Orfeo térã Bark.

Oñemohenda ñe'ẽnguéra ojehechapyréva (tembiapo = ñe'ẽnguéra):

-12 +12
0.5x 2.0x
Libre Piper, VITS, MeloTTS ndive
Audio-kuéra oguenohẽva ojekuaauka ko'ápe. Oñeporavo peteĩ modelo, omoĩnge ñe'ẽ ha ohesa'ỹijo Generar.
Audio oñemoheñói porã
0:00
Oñeguenohẽ marandu myambue guive Oñeguenohẽ.srt Ko enlace hi'are 24 h rire
Nivel libre: jeiporu personal. Licencia comercial $5/ha'e rupi
Ehayhuetéva TTS.ai? He'i umi iñangirũpe!

Mba'épa Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

Oñeha'ãvéva: AI assistants, chatbots, conversational AI applications

Ojehecha opavave Sesame CSM ñe'ẽ

Peteĩ jehecha

Desarrollador
Sesame
Licencia
Apache 2.0
Ta'ãnga
premium
Velocidad
slow
Clonación ñe'ẽnguéra rehe
No
Ñe'ẽ
English
Caracteres máx.
500

Sesame CSM ñe'ẽ

Speaker 0

English
Premium Neutral

Speaker 1

English
Premium Neutral

Sesame CSM Pregunta ojehechavéva

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← Opaite ñe'ẽ