Kani TTS 2

Kani TTS 2 TTS

An ultra-lightweight 400M English model that runs in just 3GB of VRAM at a 0.2 real-time factor.

Kulembetsa for 5,000 characters limit

Wrap wanu malemba mu SSML tags kwa kuwongolera moyenera:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags chosankhidwa chitsanzo amamvetsa - dinani kuti aphe mmodzi m'mawu anu pamene chimachitika:

Izi ndi njira yolemba malemba oyera, kotero ma tag ophatikizidwa amasiya kuganiziridwa. Kuti mupange ma tag ogwirizana ndi maganizo, gwiritsani ntchito njira yolemba malemba monga Orpheus kapena Bark.

Define custom pronunciations (word = pronunciation):

-12 +12
0.5x 2.0x
Free ndi Piper, VITS, MeloTTS
Audio yanu yopangidwa idzawonekera pano. Sankhani mtundu, lemba mawu, ndipo dinani Kupanga.
Audio Yapangidwa Mofulumira
0:00
Pangani Audio Pezani.srt Kugwirizana kumatha mu 24h
Free tier: kugwiritsa ntchito kwa munthu. Lisensi yamalonda kuchokera ku $ 5 / mo
Kukonda TTS.ai? udzauza anzanu!

Za Kani TTS 2

Kani-TTS-2 by NineNineSix is an ultra-lightweight 400M-parameter text-to-speech model built on a Liquid AI LFM2 backbone with NVIDIA's NanoCodec. It runs in just 3GB of VRAM and generates roughly ten seconds of speech in about two seconds on an A100 — a real-time factor near 0.2. The current public release ships an English-only checkpoint and, unlike its predecessor, does not expose the speaker-embedding hook needed for voice cloning. Its strength is fast, low-cost English generation on modest hardware, which makes it a good fit for quick previews and high-volume English narration. It is released under Apache 2.0 and offered on the free tier.

Best kwa: Fast English generation on low-VRAM hardware, quick previews

Pezani zonse Kani TTS 2 maganizo

Pa mphindi

Wopanga
NineNineSix
License
Apache 2.0
Mtundu
free
Kuyenda
fast
Kusintha kwa mawu
Si
Zilankhulo
English
Max characters
1000

Kani TTS 2 maganizo

Default

English
Choyambirira Neutral

Kani TTS 2 TTS — Mafunso Ofala

It runs in just 3GB of VRAM and produces about ten seconds of speech in roughly two seconds on an A100 — a real-time factor near 0.2 — thanks to its 400M-parameter LFM2 backbone and NanoCodec.

No. The current v2 release removed the public speaker-embedding hook, so cloning is not available. For cloning, use Chatterbox, IndexTTS-2, or GPT-SoVITS.

English only. The public release ships a single English checkpoint; for non-English speech, use a model like Kokoro or MeloTTS.
← Mawu onse