Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Registreeru 5000 tähemärgi piir

SSML-i siltidesse teksti segamine täpseks kontrollimiseks:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Sildid valitud mudelil mõistavad ~ klõpsa ühe kukutamiseks teksti, kus see juhtub:

See mudel loeb lihtsat teksti, nii et sisemisi silte ignoreeritakse. Sildil põhinevate emotsioonide puhul lülituge ekspressiivsele mudelile nagu Orpheus või Bark.

Kohandatud häälduste määramine (sõna = hääldus):

-12 +12
0.5x 2.0x
Tasuta Piper, VITS, MeloTTS
Siin ilmub sinu loodud heli. Vali mudel, sisesta tekst ja klõpsa Genereeri.
Audio genereeritud edukalt
0:00
Audio allalaadimine Lae alla.srt Link aegub 24 tunni pärast.
Tasuta tase: isiklik kasutamine. Äriline litsents alates $5/mo
Armastus TTS.ai?

Info Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Parim: Lightweight deployment, CPU-only environments, quick voice cloning

Kõigi sirvimine Pocket TTS hääled

Põgusalt

Arendaja
Kyutai
Litsents
MIT
Määramistasand
free
Kiirus
fast
Hääle kloonimine
Jah
Keeled
English, French
Maks. märgid
1000

Pocket TTS hääled

Alba

English
Vaba Female

Azelma

English
Vaba Female

Cosette

English
Vaba Female

Eponine

English
Vaba Female

Fantine

English
Vaba Female

Fantine (French)

French
Vaba Female

Javert

English
Vaba Male

Jean

English
Vaba Male

Jean (French)

French
Vaba Male

Marius

English
Vaba Male

Pocket TTS TTS (KKK)

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Kõik hääled