Sesame CSM

Sesame CSM 1 - تكنولوجيا المعلومات والاتصالات

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

انضم 000 5 كلمة

لف نصك في علامات SSML للتحكم الدقيق:

<speak><prosody rate="slow">Slow speech</prosody></speak>

العلامات التي يفهمها النموذج المختار — انقر لإسقاط واحدة في نصك حيث تحصل:

هذا النموذج يقرأ النص العادي، لذلك يتم تجاهل العلامات في السطر. للتعبير عن المشاعر القائمة على العلامات، انتقل إلى نموذج تعبيري مثل أورفيوس أو بارك.

تعريف النطق العادي (كلمة = نطق):

-12 +12
0.5x 2.0x
مجاني مع Piper, VITS, MeloTTS
سيظهر الصوت الذي أنتجته هنا. اختر نموذجاً، وأدخل نصاً، ثم انقر على توليد.
تم توليد الصوت بنجاح
0:00
تنزيل الصوت تنزيل.srt الرابط ينتهي بعد 24 ساعة
المستوى المجاني: الاستخدام الشخصي. ترخيص تجاري من 5 دولارات شهريا
أحب TTS.ai؟ أخبر أصدقائك!

حول Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

أفضل لل: AI assistants, chatbots, conversational AI applications

تصفح جميع Sesame CSM الأصوات

لمحة عامة

مطوِّر
Sesame
الترخيص
Apache 2.0
الرتبة
premium
السرعة
slow
استنساخ الصوت
لا
اللغات
English
الحد الأقصى للحروف
500

Sesame CSM الأصوات

Speaker 0

English
الأقساط Neutral

Speaker 1

English
الأقساط Neutral

Sesame CSM الأسئلة المتكررة

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← جميع الأصوات