Kokoro TTS
An 82M-parameter open model from Hexgrad that delivers studio-quality speech at nearly 100x real-time.
Wrap wanu malemba mu SSML tags kwa kuwongolera moyenera:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Tags chosankhidwa chitsanzo amamvetsa - dinani kuti aphe mmodzi m'mawu anu pamene chimachitika:
Izi ndi njira yolemba malemba oyera, kotero ma tag ophatikizidwa amasiya kuganiziridwa. Kuti mupange ma tag ogwirizana ndi maganizo, gwiritsani ntchito njira yolemba malemba monga Orpheus kapena Bark.
Define custom pronunciations (word = pronunciation):
Za Kokoro
Kokoro, built by Hexgrad, is a deliberately tiny 82-million-parameter model that punches far above its size class. It uses a StyleTTS + ISTFTNet architecture and was trained on roughly 1,200 hours of speech, yet generates audio close to 100x faster than real-time on a GPU while staying remarkably natural and expressive. It covers English, Japanese, Chinese, French, Italian, Portuguese, Spanish, and Hindi with a varied set of expressive voicepacks, and supports streaming. The combination of small footprint, low latency, and a permissive license has made Kokoro one of the most popular free models for high-volume and streaming use — it carries the largest share of traffic on TTS.ai.
Best kwa: High-quality TTS with minimal latency, streaming applications
Pezani zonse Kokoro maganizoPa mphindi
- Wopanga
- Hexgrad
- License
- Apache 2.0
- Mtundu
- free
- Kuyenda
- fast
- Kusintha kwa mawu
- Si
- Zilankhulo
- English, Japanese, Chinese, French, Italian, Portuguese, Spanish, Hindi
- Max characters
- 500