Kokoro TTS
An 82M-parameter open model from Hexgrad that delivers studio-quality speech at nearly 100x real-time.
Ojehaijey ñe'ẽnguéra etiquetas SSML-pe peteĩ control hekopete g̃uarã:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Etiquetas ohechakuaáva modelo ojeporavóva - tesãirã peteĩ peteĩva ñe'ẽnguérape, oĩhápe:
Ko modelo ohai texto ndahasyivéva, upévare umi etiqueta oĩva línea ryepýpe ndojehechakuaái. Umi emoción oñemopyendáva etiqueta-pe g̃uarã, oñemoambue peteĩ modelo expresivo-pe taha'e Orfeo térã Bark.
Oñemohenda ñe'ẽnguéra ojehechapyréva (tembiapo = ñe'ẽnguéra):
Mba'épa Kokoro
Kokoro, built by Hexgrad, is a deliberately tiny 82-million-parameter model that punches far above its size class. It uses a StyleTTS + ISTFTNet architecture and was trained on roughly 1,200 hours of speech, yet generates audio close to 100x faster than real-time on a GPU while staying remarkably natural and expressive. It covers English, Japanese, Chinese, French, Italian, Portuguese, Spanish, and Hindi with a varied set of expressive voicepacks, and supports streaming. The combination of small footprint, low latency, and a permissive license has made Kokoro one of the most popular free models for high-volume and streaming use — it carries the largest share of traffic on TTS.ai.
Oñeha'ãvéva: High-quality TTS with minimal latency, streaming applications
Ojehecha opavave Kokoro ñe'ẽPeteĩ jehecha
- Desarrollador
- Hexgrad
- Licencia
- Apache 2.0
- Ta'ãnga
- free
- Velocidad
- fast
- Clonación ñe'ẽnguéra rehe
- No
- Ñe'ẽ
- English, Japanese, Chinese, French, Italian, Portuguese, Spanish, Hindi
- Caracteres máx.
- 500