Spark TTS

Spark TTS TTS

Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.

Whakawhanake mō te tepe o ngā tohu 5,000

Whāriki i tōna kupu i roto i ngā tohu SSML mō te whakahaere tika:

<speak><prosody rate="slow">Slow speech</prosody></speak>

E mōhio ana ngā tohu ki te tauira i kōwhiria - ka kōwhiria kia whakawātea tētahi ki roto i tōna kupu i reira ka puta ai:

Ka pānui tēnei tauira i te kupu noa, nā reira ka whakakāhoretia ngā tohu ā-waitara. Mō te āhua o te tohu-taihi, ka huri ki tētahi tauira whakamārama pēnei i a Orpheus, Bark rānei.

Ka tautuhia ngā tohutohu ā-ringa (wāhi = tohutohu):

-12 +12
0.5x 2.0x
Waihoki me Piper, VITS, MeloTTS
Ka puta tēnei te oro i waihangatia e koe. Ka kōwhiria tētahi tauira, ka tāurua te kupu, a, ka kōwhiria te Whakatū.
Kua angitu te whakaputanga oro
0:00
Waihoki i te oro Whakahua.srt Ka ngaro te pānga i roto i te 24h
Tauwhāinga wātea: te whakamahinga whaiaro. Whakawhiwhinga hokohoko mai i te $5/mo
E manakohia ana e TTS.ai? Whakapāpāho ki ōna hoa!

Mo Spark TTS

Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.

Pai mo: Content creation with cloned voices and emotional control

Ka tirohia katoa Spark TTS ngā oro

I te tirohanga

Ka whakawhanakehia
SparkAudio
Ka taea te whakawātea
CC BY-NC-SA 4.0
Karaka
standard
Āhuatanga
medium
Whakakōrero reo
He
reo
English, Chinese
Kāri nui rawa
1000

Spark TTS ngā oro

Chinese Default

Chinese
Paerewa Neutral

Default

English
Paerewa Neutral

Spark TTS TTS - FAQ

It uses a prompt-based control system layered on top of voice cloning, so you can adjust emotion, speed, and speaking style while preserving the identity of the cloned voice.

About five seconds of reference audio is enough to clone a voice in English or Chinese.

Its model weights are licensed CC BY-NC-SA 4.0, which prohibits commercial use, even though the project code is Apache 2.0. Choose a permissively-licensed model for commercial work.
← Ko nga oro katoa