Pocket TTS

Pocket TTS TTS

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

Zarejestruj się. dla 5000 limitów znaków

Zawiń tekst w tagi SSML dla precyzyjnej kontroli:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tagi wybrany model rozumie — kliknij aby usunąć jeden do swojego tekstu, gdzie się to dzieje:

Model ten czyta tekst zwykły, więc w linii tagi są ignorowane. Dla emocji na tag, przełącz na model wyrażony jak Orfeus lub Bark.

Definiuj własny wymówki (słowo = wymówka):

-12 +12
0.5x 2.0x
Darmowe z Piper, VITS, Melotts
Tutaj pojawi się generowany dźwięk. Wybierz model, wpisz tekst i kliknij Generuj.
Pomyślnie wygenerowany dźwięk
0:00
Pobierz audio Pobierz.rt Łączność wygasa w 24h
Bezpłatny poziom: użytkowanie osobiste. Licencja handlowa od $5/mo
Powiedz znajomym!

O tematie Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

Najlepsze dla: Lightweight deployment, CPU-only environments, quick voice cloning

Przeglądaj wszystkie Pocket TTS głosy

Na jedno spojrzenie

Rozwijacz
Kyutai
Licencja
MIT
Poziom szczelności
free
Prędkość
fast
Klonowanie głosu
Tak.
Języki
English, French
Maksymalna liczba znaków
1000

Pocket TTS głosy

Alba

English
Darmowe Female

Azelma

English
Darmowe Female

Cosette

English
Darmowe Female

Eponine

English
Darmowe Female

Fantine

English
Darmowe Female

Fantine (French)

French
Darmowe Female

Javert

English
Darmowe Male

Jean

English
Darmowe Male

Jean (French)

French
Darmowe Male

Marius

English
Darmowe Male

Pocket TTS TTS — FAQ

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← Wszystkie głosy