Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Ṣẹ̀dà fun àwọn àmì-àṣírí 5,000

Fi àkọlé rẹ pamọ́ sí àwọn àmì-ìwé SSML fún ìdáràn:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Àwọn Àmì-ìwé tí àwọn ìṣàmúlò-ètò tí a yàn gbọ́ - tẹ̀ láti fi ọkan sínú àkọ́lé rẹ̀ nínú àwọn ààyè-iṣẹ́ tí o bá jẹ́:

Àwọn àwọn àkọlé àwọn ààyè-iṣẹ́ àwọn àwọn àmì-ìwé àwọn àmì-ìwé àwọn àwọn àmì-ìwé àwọn à

Àwọn àwọn ìṣàfarawé àwọn àwọn ìṣàfarawé àwọn (ọrọ = ìṣàfàlì):

-12 +12
0.5x 2.0x
Free pẹlu Piper, VITS, MeloTTS
Àwọn àwòrán tí o ti ṣẹ̀dà tí o bá han níbẹ̀. Yan àwọn àwòrán, tẹ̀lẹ̀ àkọlé, ki o si tẹ̀ Ṣẹ̀dà.
Àwọn àwọn àwòrán tí a ṣẹ̀dà
0:00
Ṣàfikún Àwọn Àmì-ìwé Ṣàfikún.srt Líǹkì náà kù nínú 24h
Ìjádé ọ̀fẹ́: ìlòjútó ara ẹni. Lisensi Iṣowo ori lati $5/mo
O fẹ́ TTS.ai? Fì sọ̀kalẹ̀ fún àwọn ọrẹ̀ rẹ̀!

Ààyè-iṣẹ́ Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Tí o dara jù fún: Lightweight deployment, CPU-only environments, quick voice cloning

Wá Gbogbo àwòrán Pocket TTS Àwọn àwòrán

Nínú àwọn ìṣàfarawé

Àwọn Àkọlé
Kyutai
Àwọn Ààyè-iṣẹ́
MIT
Àwọn àwọn ààyè-iṣẹ́
free
Ìjánu-ìṣàmúlò-ètò
fast
Ìṣàfarawé àwọn àmì-ìwé
Yà
Àwọn
English, French
Àwọn àyọkà ìpele
1000

Pocket TTS Àwọn àwòrán

Alba

English
Àìfẹ́ Female

Azelma

English
Àìfẹ́ Female

Cosette

English
Àìfẹ́ Female

Eponine

English
Àìfẹ́ Female

Fantine

English
Àìfẹ́ Female

Fantine (French)

French
Àìfẹ́ Female

Javert

English
Àìfẹ́ Male

Jean

English
Àìfẹ́ Male

Jean (French)

French
Àìfẹ́ Male

Marius

English
Àìfẹ́ Male

Pocket TTS Àwọn Àtòjọ-ẹ̀yàn

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Gbogbo àwọn ìrànwọ́