Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Aanmelden voor 5.000 tekenlimiet

Wrap uw tekst in SSML-tags voor nauwkeurige controle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags het geselecteerde model begrijpt

Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.

Definieer aangepaste uitspraaken (woord = uitspraak):

-12 +12
0.5x 2.0x
Gratis met Piper, VITS, MeloTTS
Uw gegenereerde audio zal hier verschijnen. Kies een model, voer tekst in en klik op Genereren.
Audio Generated Succesvol
0:00
Audio downloaden Download.srt Link verloopt in 24 uur
Gratis niveau: persoonlijk gebruik. Commerciële licentie van $5/mo
Hou van TTS.ai? Vertel het je vrienden!

Info Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Beste voor: Lightweight deployment, CPU-only environments, quick voice cloning

Alles doorbladeren Pocket TTS stemmen

In een oogopslag

Ontwikkelaar
Kyutai
Licentie
MIT
Niveau
free
Snelheid
fast
Klonen van stemmen
Ja.
Talen
English, French
Max. tekens
1000

Pocket TTS stemmen

Alba

English
Vrij Female

Azelma

English
Vrij Female

Cosette

English
Vrij Female

Eponine

English
Vrij Female

Fantine

English
Vrij Female

Fantine (French)

French
Vrij Female

Javert

English
Vrij Male

Jean

English
Vrij Male

Jean (French)

French
Vrij Male

Marius

English
Vrij Male

Pocket TTS Veelgestelde vragen

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Alle stemmen