Sesame CSM

Sesame CSM TTS

A 1B conversational speech model that captures natural dialogue timing, turn-taking, and backchannel responses.

ثبت نام برای حد ۵۰۰۰ کاراکتر

برای کنترل دقیق ، متن خود را در برچسبهای SSML بپیچید:

<speak><prosody rate="slow">Slow speech</prosody></speak>

برچسبهایی که مدل برگزیده می‌فهمد — برای انداختن یکی در متن خود ، جایی که اتفاق می‌افتد ، کلیک کنید:

این مدل متن ساده را می‌خواند ، بنابراین برچسب‌های خطی نادیده گرفته می‌شوند. برای احساسات مبتنی بر برچسب ، به یک مدل بیانی مانند Orpheus یا Bark تغییر دهید.

تعریف تلفظ سفارشی) کلمه = تلفظ (:

-12 +12
0.5x 2.0x
آزاد با Piper, VITS, MeloTTS
صدای تولید شده شما در اینجا ظاهر خواهد شد. یک مدل را انتخاب کنید ، متن را وارد کنید ، و تولید را فشار دهید.
صدا با موفقیت تولید شد
0:00
بارگیری صدا دانلود پیوند در ۲۴ ساعت پایان می‌یابد
1- استفاده شخصی: استفاده شخصی. مجوز تجاری از $5/mo
دوست داريد TTS.ai؟ به دوستانتون بگو!

در مورد Sesame CSM

Sesame CSM (Conversational Speech Model) is a 1-billion-parameter model from Sesame designed specifically for the rhythms of human conversation. Built on a Llama backbone paired with an audio codec, it models turn-taking timing, backchannel responses (the small acknowledgements people make while listening), emotional reactions, and overall conversational flow. The result reads less like read-aloud text and more like a real spoken exchange. It is a natural fit for AI assistants, chatbots, and conversational interfaces where the goal is speech that feels responsive and human. CSM is released under Apache 2.0, and access on TTS.ai requires a Hugging Face token at the model level.

بهترین برای: AI assistants, chatbots, conversational AI applications

مرور همۀ Sesame CSM صداها

يه نگاهي بنداز

توسعه‌دهنده
Sesame
مجوز
Apache 2.0
حیوان
premium
سرعت
slow
شبیه‌سازی صدا
نه
زبانها
English
بیشینه نویسه‌ها
500

Sesame CSM صداها

Speaker 0

English
پریمیوم Neutral

Speaker 1

English
پریمیوم Neutral

Sesame CSM FAQ - پرسش و پاسخ

Conversational speech. It models the natural patterns of dialogue — turn-taking timing, backchannel responses, and emotional reactions — so generated audio sounds like a real conversation rather than synthetic narration.

It is a 1-billion-parameter model built on a Llama backbone with an audio codec for waveform generation.

AI assistants, chatbots, and other conversational applications where responsive, human-sounding speech matters more than long-form narration.
← همه صداها