Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

မှတ်ပုံတင်ပါ 5,000 စာလုံးအဆုံးသတ်

တိကျသောထိန်းချုပ်မှုများအတွက် SSML tags များအတွင်းသင်၏စာသားကို Wrap:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags ရွေးချယ်ထားသောမော်ဒယ်နားလည် — ဒါဟာဖြစ်ပျက်တဲ့နေရာမှာသင့်ရဲ့စာသားထဲသို့တစ်ဦး drop ကိုကလစ်နှိပ်ပါ:

ဒီပုံစံကိုရိုးရှင်းတဲ့စာသားဖတ်ရှု, ဒါကြောင့် inline tags တွေကိုလျစ်လျူရှုနေကြသည်။ tag ကိုအခြေခံစိတ်ခံစားမှုအတွက်, Orpheus သို့မဟုတ် Bark ကဲ့သို့သောဖော်ပြမှုပုံစံသို့ switch.

custom pronunciations ကိုသတ်မှတ်ပါ (word = pronunciation):

-12 +12
0.5x 2.0x
Piper, VITS, MeloTTS နှင့်အတူအခမဲ့
သင်၏ထုတ်လုပ်အသံဒီမှာပေါ်လာလိမ့်မည်။ ရွေးချယ်ပါ, စာသားကိုထည့်သွင်း, နှင့် Generate ကိုကလစ်နှိပ်ပါ.
အသံဖိုင် အောင်မြင်စွာ ထုတ်လုပ်နိုင်ခဲ့သည်
0:00
အသံဖိုင်များ ဒေါင်းလုပ်လုပ် ဒေါင်းလုပ်လုပ်.srt Link ကို 24h တွင်ကုန်ဆုံးသည်
အခမဲ့ tier: ပုဂ္ဂလိကအသုံးပြုမှု။ စီးပွားရေးလိုင်စင်မှ $5/mo
TTS.ai ကိုချစ်ပါသလား?

အကြောင်း Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

အကောင်းဆုံး: AI assistants, chatbots, conversational AI applications

အားလုံးကို ရှာဖွေ Sesame CSM အသံများ

တစ်ချက်ကြည့်ပါ

ဖန်တီးသူ
Sesame
လိုင်စင်
Apache 2.0
အမျိုးအစား
premium
အမြန်နှုန်း
slow
အသံကို ကူးယူခြင်း
ဟုတ်ကဲ့
ဘာသာစကားများ
English
အက္ခရာအရေအတွက်
500

Sesame CSM အသံများ

Speaker 0

English
ပရီမီယံ Neutral

Speaker 1

English
ပရီမီယံ Neutral

Sesame CSM TTS — FAQ

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← အသံအားလုံး