Pocket TTS TTS
A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.
Envolva o seu texto em tags SSML para controle preciso:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etiquetas o modelo selecionado entende — clique para soltar um para o seu texto onde acontece:
Este modelo lê texto simples, por isso as etiquetas inline são ignoradas. Para emoção baseada em tags, mude para um modelo expressivo como Orpheus ou Bark.
Definir pronúncias personalizadas (palavra = pronúncia):
Sobre Pocket TTS
Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.
Melhor para: Lightweight deployment, CPU-only environments, quick voice cloning
Procurar todos Pocket TTS vozesDe uma olhada
- Desenvolvedor
- Kyutai
- Licença
- MIT
- Tier
- free
- Velocidade
- fast
- Clonagem de voz
- Sim
- Línguas
- English, French
- Número máximo de caracteres
- 1000