Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Daftar masuk had 5,000 aksara

Lilitkan teks anda dalam tag SSML untuk kawalan tepat:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag model dipilih memahami - klik untuk jatuhkan satu ke dalam teks anda di mana ia berlaku:

Model ini membaca teks biasa, jadi tag dalam baris diabaikan. Untuk emosi berdasar tag, beralih ke model ekspresif seperti Orpheus atau Bark.

Tetapkan sebutan tersendiri (perkataan = sebutan):

-12 +12
0.5x 2.0x
Bebas dengan Piper, VITS, MeloTTS
Audio yang dijana akan muncul di sini. Pilih model, masukkan teks, dan klik Janakan.
Audio Dijana Dengan Berjaya
0:00
Muat turun Audio Muat turun.srt Pautan luput dalam 24 jam
Tahap percuma: penggunaan peribadi. Lesen Komersial dari $5/mo
Cinta TTS.ai? Beritahu kawan-kawan anda!

Tentang Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Terbaik untuk: Lightweight deployment, CPU-only environments, quick voice cloning

Layari semua Pocket TTS suara

Dengan sekejap mata

Pemaju
Kyutai
Lesen
MIT
Tajuk
free
Kelajuan
fast
Klon suara
Ya
Bahasa
English, French
Aksara maksimum
1000

Pocket TTS suara

Alba

English
Bebas Female

Azelma

English
Bebas Female

Cosette

English
Bebas Female

Eponine

English
Bebas Female

Fantine

English
Bebas Female

Fantine (French)

French
Bebas Female

Javert

English
Bebas Male

Jean

English
Bebas Male

Jean (French)

French
Bebas Male

Marius

English
Bebas Male

Pocket TTS TTS - FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Semua suara