Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

Înscrie-te pentru limitele de 5000 de caractere

Întoarceți textul în etichetele SSML pentru un control precis:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etichetele modelului selectat înțeleg — click pentru a lăsa unul în textul tău unde se întâmplă:

Acest model citește textul simplu, astfel încât etichetele inline sunt ignorate. Pentru emoții bazate pe tag, schimbați la un model expresiv cum ar fi Orpheus sau Bark.

Definiți pronunțiare personalizată (cuvânt = pronunție):

-12 +12
0.5x 2.0x
Gratuit cu Piper, VITS, MeloTTS
Audio generat va apărea aici. Alegeți un model, introduceți text și faceți clic pe Generați.
Audio generat cu succes
0:00
Descarcă audio Descărcare.srt Legătura expiră în 24 ore
Gratuit: utilizare personală. Licență comercială de la 5$/mo
Spune-i prietenilor tăi!

Despre Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

Cel mai bun pentru: AI assistants, chatbots, conversational AI applications

Navigați toate Sesame CSM voci

La o privire

Dezvoltator
Sesame
Licență
Apache 2.0
Nivel
premium
Viteză
slow
Clonarea vocală
Nu.
Limbi
English
Caractere maxime
500

Sesame CSM voci

Speaker 0

English
Premium Neutral

Speaker 1

English
Premium Neutral

Sesame CSM TTS – FAQ

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← Toate vocile