Sesame CSM

Sesame CSM نووسراو بۆ بیستراو

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

تۆماربکە بۆ ٥٠٠٠ هێما

نوسراوەکەت بگۆڕە بۆ تگەکانی SSML بۆ کۆنتڕۆڵی ڕاستەقینە:

<speak><prosody rate="slow">Slow speech</prosody></speak>

تگەکان کە مۆدێلی هەڵبژێردراو تێیدەگات - بکەرەوە بۆ ئەوەی یەکێکیان بخەیتە ناو دەقەکەتەوە کە تیایدا ڕوودەدات:

ئەم مۆدێلە نوسراوێکی ئاسایی دەخوێنێتەوە، بۆیە تاگەکانی ناو ڕستە پشتگوێ دەخرێت. بۆ هەستێکی لەسەر بنەمای تاگ، بگەڕێ بۆ مۆدێلێکی دەربڕین وەک ئۆرفیۆس یان بارک.

پێناسەکردنی دەنگی خۆت (وشە = دەنگی):

-12 +12
0.5x 2.0x
بەبێ پارە لەگەڵ Piper, VITS, MeloTTS
دەنگی دروستکراوت لێرەدا دەردەکەوێت. مۆدێلێک هەڵبژێرە، نوسراوێک دابنێ، پاشان کلیک بکە لەسەر دروستکردن.
دەنگ بە سەرکەوتن دروست کرا
0:00
دابەزاندنی دەنگ دابەزاندن پەیوەستەکە دوای ٢٤ کاتژمێر کۆتایی دێت
پلەی ئازاد: بەکارهێنانی تایبەتی. مۆڵەتی بازرگانی لە ٥$/ مانگ
خۆشت دەوێت TTS.ai؟ بە هاوڕێکانت بڵێ!

دەربارەی Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

باشترین بۆ: AI assistants, chatbots, conversational AI applications

سەردانی هەموویان بکە Sesame CSM دەنگی

چاوپێکەوتن

پەرەپێدەر
Sesame
مۆڵەتی بەکارھێنەر
Apache 2.0
یه‌مه‌ن
premium
خێرایی
slow
دووبارە دروستکردنی دەنگی
نەخێر
زمان
English
زۆرترین پیت
500

Sesame CSM دەنگی

Speaker 0

English
پڕۆمیۆم Neutral

Speaker 1

English
پڕۆمیۆم Neutral

Sesame CSM پرسیاری زۆر کراوە

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← هەموو دەنگەکان