Pocket TTS TTS
A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.
Omotajte tekst u SSML oznake za preciznu kontrolu:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Oznake koje odabrani model razumije — kliknite da biste ih ubacili u tekst gdje se pojavljuju:
Ovaj model čita običan tekst, tako da se inline oznake ignoriraju. Za emocije zasnovane na oznakama, prebacite se na ekspresivni model poput Orpheusa ili Bark-a.
Definirajte vlastite izgovore (riječ = izgovor):
O meni Pocket TTS
Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.
Najbolje za: Lightweight deployment, CPU-only environments, quick voice cloning
Pregledaj sve Pocket TTS glasoviNa prvi pogled
- Programer
- Kyutai
- Licenca
- MIT
- Životinje
- free
- Brzina
- fast
- Kloniranje glasa
- Da.
- Jezici
- English, French
- Maksimalan broj znakova
- 1000