Pocket TTS TTS
A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.
Wrap uw tekst in SSML-tags voor nauwkeurige controle:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags het geselecteerde model begrijpt
Dit model leest platte tekst, dus inline tags worden genegeerd. Voor emotie op basis van tags, schakel naar een expressief model zoals Orpheus of Bark.
Definieer aangepaste uitspraaken (woord = uitspraak):
Info Pocket TTS
Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.
Beste voor: Lightweight deployment, CPU-only environments, quick voice cloning
Alles doorbladeren Pocket TTS stemmenIn een oogopslag
- Ontwikkelaar
- Kyutai
- Licentie
- MIT
- Niveau
- free
- Snelheid
- fast
- Klonen van stemmen
- Ja.
- Talen
- English, French
- Max. tekens
- 1000