Sesame CSM

Sesame CSM អត្ថបទ​ទៅ​សំឡេង

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

ចុះឈ្មោះ កំណត់​សម្រាប់​តួអក្សរ ៥, ០០០

រុំ​អត្ថបទ​របស់​អ្នក​ក្នុង​ស្លាក SSML សម្រាប់​ការ​ត្រួតពិនិត្យ​ជាក់លាក់ ៖

<speak><prosody rate="slow">Slow speech</prosody></speak>

ស្លាក​ដែល​ម៉ូដែល​ដែល​បាន​ជ្រើស​យល់​ — ចុច​ដើម្បី​ទម្លាក់​មួយ​ទៅ​ក្នុង​អត្ថបទ​របស់​អ្នក​នៅ​កន្លែង​ដែល​វា​កើតឡើង ៖

ម៉ូដែល​នេះ​អាន​អត្ថបទ​ធម្មតា ដូច្នេះ​ស្លាក​ក្នុង​បន្ទាត់​ត្រូវ​បាន​មិន​អើពើ ។ សម្រាប់​អារម្មណ៍​ដែល​មាន​មូលដ្ឋាន​លើ​ស្លាក ប្ដូរ​ទៅ​ម៉ូដែល​បង្ហាញ​ដូច​ជា Orpheus ឬ Bark ។

កំណត់​ការ​បញ្ចេញ​សំឡេង​ផ្ទាល់ខ្លួន (ពាក្យ = ការ​បញ្ចេញ​សំឡេង) ៖

-12 +12
0.5x 2.0x
ឥតគិតថ្លៃ​ជាមួយ Piper, VITS, MeloTTS
អូឌីយ៉ូ​ដែល​បាន​បង្កើត​របស់​អ្នក​នឹង​លេចឡើង​នៅ​ទីនេះ ។ ជ្រើស​ម៉ូដែល បញ្ចូល​អត្ថបទ ហើយ​ចុច បង្កើត ។
បាន​បង្កើត​អូឌីយ៉ូ​ដោយ​ជោគជ័យ
0:00
ទាញយក​អូឌីយ៉ូ ទាញយក.srt តំណផុតកំណត់ក្នុង 24h
កម្រិត​ឥត​គិត​ថ្លៃ ៖ ការ​ប្រើ​ផ្ទាល់​ខ្លួន ។ អាជ្ញាប័ណ្ណពាណិជ្ជកម្មពី $5/ខែ
ធ្វើ​ឲ្យ​នេះ​ជា​សំឡេង​របស់​អ្នក ក្លូន​សំឡេង​ក្នុង ៣០ វិនាទី
ស្រឡាញ់ TTS.ai? ប្រាប់មិត្តភក្តិរបស់អ្នក!

អំពី Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

ល្អបំផុត​សម្រាប់: AI assistants, chatbots, conversational AI applications

រកមើល​ទាំងអស់ Sesame CSM សំឡេង

ទិដ្ឋភាព​ទូទៅ

អ្នក​អភិវឌ្ឍន៍
Sesame
អាជ្ញាបណ្ណ
Apache 2.0
ផ្កាយ
premium
ល្បឿន
slow
ការ​ក្លូន​សំឡេង
គ្មាន
ភាសា
English
តួអក្សរ​អតិបរមា
500

Sesame CSM សំឡេង

Speaker 0

English
តម្លៃ​ខ្ពស់ Neutral

Speaker 1

English
តម្លៃ​ខ្ពស់ Neutral

Sesame CSM TTS - សំណួរ​ដែល​សួរ​ញឹកញាប់

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← សំឡេង​ទាំងអស់