Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

Inscríbete para el límite de 5.000 caracteres

Envuelva su texto en etiquetas SSML para un control preciso:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etiquetas el modelo seleccionado entiende — haga clic para soltar uno en su texto donde sucede:

Este modelo lee texto plano, por lo que las etiquetas en línea son ignoradas. Para la emoción basada en etiquetas, cambie a un modelo expresivo como Orfeo o Bark.

Definir pronunciaciones personalizadas (palabra = pronunciación):

-12 +12
0.5x 2.0x
Libre con Piper, VITS, MeloTTS
Su audio generado aparecerá aquí. Elija un modelo, introduzca texto y haga clic en Generar.
Audio generado con éxito
0:00
Descargar audio Descargar.srt Enlace expira en 24h
Nivel libre: uso personal. Licencia comercial desde $5/mes
Haz de esto tu propia voz Clonar una voz en 30 segundos
¿Te gusta TTS.ai? ¡Cuéntaselo a tus amigos!

Acerca de Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

Lo mejor para: AI assistants, chatbots, conversational AI applications

Examinar todo Sesame CSM voces

De un vistazo

Desarrollador
Sesame
Licencia
Apache 2.0
Nivel
premium
Velocidad
slow
Clonación de voz
No
Idiomas
English
Máx. caracteres
500

Sesame CSM voces

Speaker 0

English
Prima Neutral

Speaker 1

English
Prima Neutral

Sesame CSM TTS — Preguntas más frecuentes

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← Todas las voces