Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Melden Sie sich an für 5.000 Zeichen-Grenze

Verpacken Sie Ihren Text in SSML-Tags für eine präzise Kontrolle:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags, die das ausgewählte Modell versteht — klicken Sie, um einen in Ihren Text zu legen, wo es passiert:

Dieses Modell liest Text, so dass Inline-Tags ignoriert werden. Für tag-basierte Emotion, wechseln Sie zu einem ausdrucksstarken Modell wie Orpheus oder Bark.

Benutzerdefinierte Aussprachen definieren (Wort = Aussprache):

-12 +12
0.5x 2.0x
Frei mit Piper, VITS, MeloTTS
Hier erscheint Ihr generiertes Audio. Wählen Sie ein Modell, geben Sie Text ein und klicken Sie auf Generieren.
Audio-Erzeugung erfolgreich
0:00
Audio herunterladen Download.srt Link läuft in 24h aus
Freier Dienstgrad: persönlicher Gebrauch. Kommerzielle Lizenz ab $5/mo
Gefällt dir TTS.ai? Erzähl es deinen Freunden!

Über Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Das Beste für: Lightweight deployment, CPU-only environments, quick voice cloning

Alle durchsuchen Pocket TTS Stimmen

Auf einen Blick

Entwickler
Kyutai
Lizenz
MIT
Tierart
free
Geschwindigkeit
fast
Klonen der Stimme
Nein
Sprachen
English, French
Maximale Zeichen
1000

Pocket TTS Stimmen

Alba

English
Frei Female

Azelma

English
Frei Female

Cosette

English
Frei Female

Eponine

English
Frei Female

Fantine

English
Frei Female

Fantine (French)

French
Frei Female

Javert

English
Frei Male

Jean

English
Frei Male

Jean (French)

French
Frei Male

Marius

English
Frei Male

Pocket TTS TTS — FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Alle Stimmen