Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Aliĝi for 5, 000 character limit

Envolvu vian tekston en SSML- etikedojn por preciza kontrolo:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Etikedoj kiujn la elektita modelo komprenas - klaku por meti unu en vian tekston kie ĝi okazas:

This model reads plain text, so inline tags are ignored. For tag-based emotion, switch to an expressive model like Orpheus or Bark.

Difini proprajn elparolojn (vorto = elparolo):

-12 +12
0.5x 2.0x
Libera kun Piper, VITS, MeloTTS
Via generita sono aperos tie ĉi. Elektu modelon, entajpu tekston, kaj alklaku Generi.
Sondosiero sukcese generita
0:00
Elŝuti sonon Elŝuti.srt Ligo eksvalidiĝas post 24 horoj
Libera programaro: persona uzo. Komerca licenco ekde $5/mo
Ĉu vi ŝatas TTS.ai? Diru al viaj amikoj!

Pri Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Plej bona por: Lightweight deployment, CPU-only environments, quick voice cloning

Foliumi ĉiujn Pocket TTS voĉoj

Unu rigardo

Programisto
Kyutai
Licenco
MIT
Tamuz
free
Rapideco
fast
Voĉo- klonado
Jes
Lingvoj
English, French
Maksimuma nombro da signoj
1000

Pocket TTS voĉoj

Alba

English
Libera Female

Azelma

English
Libera Female

Cosette

English
Libera Female

Eponine

English
Libera Female

Fantine

English
Libera Female

Fantine (French)

French
Libera Female

Javert

English
Libera Male

Jean

English
Libera Male

Jean (French)

French
Libera Male

Marius

English
Libera Male

Pocket TTS TTS - FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Ĉiuj voĉoj