Chatterbox TTS
Resemble AI's state-of-the-art zero-shot voice cloning model with independent emotion control.
Wrap your text in SSML tags for precise control:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags the selected model understands — click to drop one into your text where it happens:
This model reads plain text, so inline tags are ignored. For tag-based emotion, switch to an expressive model like Orpheus or Bark.
Define custom pronunciations (word = pronunciation):
About Chatterbox
Chatterbox by Resemble AI is a leading open-source zero-shot voice cloning model that replicates a voice from a single audio sample, capturing not just timbre but speaking style and emotional nuance. Its distinctive feature is fine-grained emotion control that operates independently of the voice identity, so you can keep a cloned voice but shift its emotional intensity. Built around ResembleEnhance and flow matching, it targets professional-grade cloning for content creation, dubbing, and character voices. Released under the permissive MIT license, Chatterbox has become a popular foundation for derivative models — TTS.ai also runs a Saudi-Arabic fine-tune of it. It favors quality, with a modest per-request character limit.
Best for: Professional voice cloning with emotional control, content creation
Browse all Chatterbox voicesAt a glance
- Developer
- Resemble AI
- License
- MIT
- Tier
- premium
- Speed
- medium
- Voice cloning
- Yes
- Languages
- English
- Max characters
- 300