Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Langganan for 5,000 characters limit

Ngresiki teks ing tag SSML kanggo kontrol presisi:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag kang dipahami model kang dipilih - klik kanggo ngethok siji ing teks sampeyan ing ngendi iku kedadeyan:

Model iki maca teks biasa, mula tag ing baris diabaikan. Kanggo emosi berbasis tag, ganti menyang model ekspresif kaya Orpheus utawa Bark.

Nyathet tembung-tembung standar (kata = tembung):

-12 +12
0.5x 2.0x
Bebas karo Piper, VITS, MeloTTS
Audio sing digawé bakal katon ing kene. Pilih modél, ketik teks, lan pencet Ngembangaké.
Audio Digawé kanthi Sukses
0:00
Unduh Audio Download.srt Link expires in 24h
Ing basa Indonésia, iku tegesé: pribadi. Lisénsi komersial saka $5/mo
TTS.ai? Nyathet kanca-kancamu!

Ngendi Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Paling apik kanggo: Lightweight deployment, CPU-only environments, quick voice cloning

Jajal kabeh Pocket TTS swara

Ing cetha

Pangembang
Kyutai
Lisénsi
MIT
Tanggal
free
Kecepatan
fast
Kloning swara
Ya
Basa
English, French
Aksara paling akèh
1000

Pocket TTS swara

Alba

English
Bebas Female

Azelma

English
Bebas Female

Cosette

English
Bebas Female

Eponine

English
Bebas Female

Fantine

English
Bebas Female

Fantine (French)

French
Bebas Female

Javert

English
Bebas Male

Jean

English
Bebas Male

Jean (French)

French
Bebas Male

Marius

English
Bebas Male

Pocket TTS FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Sekabehing swara