Pocket TTS

Pocket TTS Ikiganiro

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Kwiyandikisha kugirango Inyuguti

Umwandiko in Itagi: ya: Igenzura:

<speak><prosody rate="slow">Slow speech</prosody></speak>

i Byahiswemo Urugero - Kanda Kuri Gukuraho Rimwe Umwandiko:

Urugero: Bisanzwe Umwandiko, Umurongo: Itagi:. Itagi: -, Hindura Kuri Urugero: Nka Cyangwa.

Kugena (Ijambo =):

-12 +12
0.5x 2.0x
Na:,,
Audio Kugaragara. A Urugero:, Injiza Umwandiko, na Kanda.
Byaremwe
0:00
Iyimura Iyimura Ihuza in
Bwite Koresha Kuva: 5 /
TTS.ai? Abayobozi!

Ikiganiro Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

kugirango: Lightweight deployment, CPU-only environments, quick voice cloning

Gushakisha byose Pocket TTS Amashusho

A

Mukoraporogaramu
Kyutai
Inyandiko y'Iyemererakoresha
CC BY 4.0 (weights)
Itariki
free
Umuvuduko
fast
Guhindura izina
Oya
Ururimi:
English, French
Inyuguti
1000

Pocket TTS Amashusho

Alba

English
Kigenga Female

Azelma

English
Kigenga Female

Cosette

English
Kigenga Female Na - Ikoresha:

Eponine

English
Kigenga Female

Fantine

English
Kigenga Female

Fantine (French)

French
Kigenga Female

Javert

English
Kigenga Male

Jean

English
Kigenga Male Na - Ikoresha:

Jean (French)

French
Kigenga Male Na - Ikoresha:

Marius

English
Kigenga Male

Pocket TTS -

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Mostly. Pocket TTS weights are released under CC BY 4.0 and it sits in the free tier. Its Jean and Cosette voices come from non-commercial datasets and are marked "Personal and non-commercial use only" on every plan. Audio from its other voices is for personal use on the free tier and can be used commercially on any paid plan. It supports English and French.
← Amashusho yose