Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Đăng ký giới hạn 5000 ký tự

Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:

Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.

Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):

-12 +12
0.5x 2.0x
Miễn phí với Piper, VITS, MeloTTS
Âm thanh đã tạo sẽ xuất hiện ở đây. Chọn một mô hình, nhập văn bản, và nhấn vào Tạo.
Âm thanh đã được tạo thành công
0:00
Tải về âm thanh Tải về Liên kết hết hạn trong 24h
Free tier: sử dụng cá nhân. Giấy phép thương mại từ $5/mo
Cảm ơn bạn đã tin tưởng TTS.ai!

Về Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Tốt nhất cho: Lightweight deployment, CPU-only environments, quick voice cloning

& Xem tất cả Pocket TTS giọng nói

Một cái nhìn

Nhà phát triển
Kyutai
Giấy phép
CC BY 4.0 (weights)
Thú
free
Tốc độ
fast
Ký âm
Ngôn ngữ
English, French
Tối đa các ký tự
1000

Pocket TTS giọng nói

Alba

English
Tự do Female

Azelma

English
Tự do Female

Cosette

English
Tự do Female Chỉ dùng cá nhân và không thương mại

Eponine

English
Tự do Female

Fantine

English
Tự do Female

Fantine (French)

French
Tự do Female

Javert

English
Tự do Male

Jean

English
Tự do Male Chỉ dùng cá nhân và không thương mại

Jean (French)

French
Tự do Male Chỉ dùng cá nhân và không thương mại

Marius

English
Tự do Male

Pocket TTS TTS - FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Mostly. Pocket TTS weights are released under CC BY 4.0 and it sits in the free tier. Its Jean and Cosette voices come from non-commercial datasets and are marked "Personal and non-commercial use only" on every plan. Audio from its other voices is for personal use on the free tier and can be used commercially on any paid plan. It supports English and French.
← Tất cả giọng nói