Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

Vpišite se. za 5000 mejnih vrednosti znakov

Za natančen nadzor zavijte svoje besedilo v oznake SSML:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Oznake izbranega modela razume – kliknite, da spustite enega v svoje besedilo, kjer se zgodi:

Ta model bere navadno besedilo, zato se v vrstici ignorirajo. Za čustva, ki temeljijo na tag, preklopite na izražen model, kot je Orfeus ali Bark.

Opredelitev posebnih izgovorov (beseda = izgovor):

-12 +12
0.5x 2.0x
Brez Piper, VITS, Melotts
Tukaj se bo pojavil vaš ustvarjeni zvok. Izberite model, vnesite besedilo in kliknite Generiraj.
Uspešno ustvarjen zvok
0:00
Prenesi zvok Prenesi.rt Povezava poteče čez 24h
Brezplačna stopnja: osebna uporaba. Trgovska licenca od 5 $/mo
Ljubi TTS.ai, povej prijateljem!

O projektu Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

Najboljše za: AI assistants, chatbots, conversational AI applications

Brskaj vse Sesame CSM glasovi

Na pogled

Razvijalec
Sesame
Licenca
Apache 2.0
Stopnja
premium
Hitrost
slow
kloniranje glasu
Ne
Jeziki
English
Največ znakov
500

Sesame CSM glasovi

Speaker 0

English
Premium Neutral

Speaker 1

English
Premium Neutral

Sesame CSM TTS – Pogosta vprašanja

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← Vsi glasovi