Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Daftar untuk batas 5,000 karakter

Bungkus teks Anda dalam tag SSML untuk kendali yang tepat:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag yang dipilih mengerti klik °C untuk memasukkan satu ke dalam teks Anda di mana hal itu terjadi:

Model ini membaca teks biasa, sehingga tag inline diabaikan. Untuk tag berbasis emosi, beralih ke model ekspresif seperti Orpheus atau Bark.

Definisikan pengucapan ubahan (kata = pelafalan):

-12 +12
0.5x 2.0x
Free with Piper, VITS, Melotts
Audio yang Anda buat akan muncul di sini. Pilih model, masukkan teks, dan klik Generate.
Hasil Audio Berhasil
0:00
Unduh Audio Unduh.srt Sambungan berakhir dalam 24 jam
Tingkatan bebas: penggunaan pribadi. Ijin komersial dari $5/mo
Buatlah ini suara Anda sendiri Kloning suara dalam 30 detik
Beritahu teman-temanmu!

Tentang Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Terbaik untuk: Lightweight deployment, CPU-only environments, quick voice cloning

Jelajahi semua Pocket TTS suara

Pada sekilas

Pengembang
Kyutai
Lisensi
MIT
Tier
free
Kecepatan
fast
Penklonan Suara
Ya
Bahasa
English, French
Karakter maksimal
1000

Pocket TTS suara

Alba

English
Bebas Female

Azelma

English
Bebas Female

Cosette

English
Bebas Female

Eponine

English
Bebas Female

Fantine

English
Bebas Female

Fantine (French)

French
Bebas Female

Javert

English
Bebas Male

Jean

English
Bebas Male

Jean (French)

French
Bebas Male

Marius

English
Bebas Male

Pocket TTS TTS °F FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Semua suara