Spark TTS

Spark TTS TTS

Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.

Langganan for 5,000 characters limit

Ngresiki teks ing tag SSML kanggo kontrol presisi:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tag kang dipahami model kang dipilih - klik kanggo ngethok siji ing teks sampeyan ing ngendi iku kedadeyan:

Model iki maca teks biasa, mula tag ing baris diabaikan. Kanggo emosi berbasis tag, ganti menyang model ekspresif kaya Orpheus utawa Bark.

Nyathet tembung-tembung standar (kata = tembung):

-12 +12
0.5x 2.0x
Bebas karo Piper, VITS, MeloTTS
Audio sing digawé bakal katon ing kene. Pilih modél, ketik teks, lan pencet Ngembangaké.
Audio Digawé kanthi Sukses
0:00
Unduh Audio Download.srt Link expires in 24h
Ing basa Indonésia, iku tegesé: pribadi. Lisénsi komersial saka $5/mo
TTS.ai? Nyathet kanca-kancamu!

Ngendi Spark TTS

Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.

Paling apik kanggo: Content creation with cloned voices and emotional control

Jajal kabeh Spark TTS swara

Ing cetha

Pangembang
SparkAudio
Lisénsi
CC BY-NC-SA 4.0
Tanggal
standard
Kecepatan
medium
Kloning swara
Ya
Basa
English, Chinese
Aksara paling akèh
1000

Spark TTS swara

Chinese Default

Chinese
Standar Neutral

Default

English
Standar Neutral

Spark TTS FAQ

It uses a prompt-based control system layered on top of voice cloning, so you can adjust emotion, speed, and speaking style while preserving the identity of the cloned voice.

About five seconds of reference audio is enough to clone a voice in English or Chinese.

Its model weights are licensed CC BY-NC-SA 4.0, which prohibits commercial use, even though the project code is Apache 2.0. Choose a permissively-licensed model for commercial work.
← Sekabehing swara