Spark TTS

Spark TTS TTS

Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.

Kugadzwa for 5,000 character limit

Wrap yako tenzi mu SSML tags kuti zvive nyore kudzora:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags iyo yakasarudzwa model inonzwisiswa — tinya kuti utore imwe muchinyorwa chako apo inoitika:

Iyi modhi inoverenga chete mazita ezvinyorwa, saka zvinyorwa zvinonyorwa mumitsara hazvina kutariswa. Kuti uwane pfungwa dziri mumitsara, chinja kune imwe modhi inoratidza pfungwa seOrpheus kana Bark.

Define custom pronunciations (word = pronunciation):

-12 +12
0.5x 2.0x
Free with Piper, VITS, MeloTTS
Yako yakagadzirwa audio ichaonekwa pano. Choose a model, enter text, and click Generate.
Audio Yakagadzirwa Nekubudirira
0:00
Download Audio Dhawunirodha.srt Link inotanga kushanda mu 24h
Free tier: personal usage. Commercial license from $5/mo
Love TTS.ai? Tiudza shamwari dzako!

Chii Spark TTS

Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.

Yakanaka kune: Content creation with cloned voices and emotional control

Tarisa zvese Spark TTS mazwi

Mufananidzo

Developer
SparkAudio
License
CC BY-NC-SA 4.0
Tier
standard
Speed
medium
Kutaura
Yes
Zvinhu
English, Chinese
Max characters
1000

Spark TTS mazwi

Chinese Default

Chinese
Chimiro Neutral

Default

English
Chimiro Neutral

Spark TTS TTS — Zvinyorwa

It uses a prompt-based control system layered on top of voice cloning, so you can adjust emotion, speed, and speaking style while preserving the identity of the cloned voice.

About five seconds of reference audio is enough to clone a voice in English or Chinese.

Its model weights are licensed CC BY-NC-SA 4.0, which prohibits commercial use, even though the project code is Apache 2.0. Choose a permissively-licensed model for commercial work.
← All voices