Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Sign up for 5,000 character limit

Wrap your text in SSML tags for precise control:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags the selected model understands — click to drop one into your text where it happens:

This model reads plain text, so inline tags are ignored. For tag-based emotion, switch to an expressive model like Orpheus or Bark.

Define custom pronunciations (word = pronunciation):

-12 +12
0.5x 2.0x
Free with Piper, VITS, MeloTTS
Your generated audio will appear here. Choose a model, enter text, and click Generate.
Audio Generated Successfully
0:00
Download Audio Download .srt Link expires in 24h
Free tier: personal use. Commercial license from $5/mo
Love TTS.ai? Tell your friends!

About Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Best for: Lightweight deployment, CPU-only environments, quick voice cloning

Browse all Pocket TTS voices

At a glance

Developer
Kyutai
License
CC BY 4.0 (weights)
Tier
free
Speed
fast
Voice cloning
Yes
Languages
English, French
Max characters
1000

Pocket TTS voices

Alba

English
Free Female

Azelma

English
Free Female

Cosette

English
Free Female Personal and non-commercial use only

Eponine

English
Free Female

Fantine

English
Free Female

Fantine (French)

French
Free Female

Javert

English
Free Male

Jean

English
Free Male Personal and non-commercial use only

Jean (French)

French
Free Male Personal and non-commercial use only

Marius

English
Free Male

Pocket TTS TTS — FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Mostly. Pocket TTS weights are released under CC BY 4.0 and it sits in the free tier. Its Jean and Cosette voices come from non-commercial datasets and are marked "Personal and non-commercial use only" on every plan. Audio from its other voices is for personal use on the free tier and can be used commercially on any paid plan. It supports English and French.
← All voices