Spark TTS

Spark TTS Mga TNT

Voice cloning from five seconds of audio combined with prompt-based control over emotion, speed, and speaking style.

Mag-sign up para sa 5,000 character na limitasyon

I-wrap ang iyong teksto sa SSML tags para sa tumpak na kontrol:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags ang napili modelo nauunawaan — i-click upang ihulog ang isa sa iyong teksto kung saan ito ay nangyayari:

Ang modelong ito ay nagbabasa ng karaniwang teksto, kaya inline tags ay hindi pinapansin. Para sa tag-based na damdamin, lumipat sa isang makahulugang modelo tulad ng Orpheus o Bark.

Tukuyin ang mga pasadyang mga panlapi (word = panlapi):

-12 +12
0.5x 2.0x
Libreng may Piper, VITS, MeloTTS
Ang iyong ginawang audio ay lilitaw dito. Pumili ng modelo, ipasok ang teksto, at i-click ang Bumuo.
Audio nabuo Matagumpay
0:00
I-download ang Audio I-download ang.srt Link expires sa 24h
Free tier: personal na paggamit. Komersyal na lisensya mula sa $5/buwan
I-love TTS.ai? Ibahagi sa iyong mga kaibigan!

Tungkol sa Spark TTS

Spark TTS by SparkAudio merges voice cloning with controllable delivery in a single prompt-driven system. Using just five seconds of reference audio it clones a voice, then lets you steer emotion, speed, and speaking style while keeping that cloned identity intact. Under the hood it combines a BiCodec audio tokenizer, an LLM, and flow matching, and it supports English and Chinese. It is aimed at content creation where a single cloned voice needs to express a range of moods and pacing. Note the licensing split: Spark's code is Apache 2.0, but the model weights are released under CC BY-NC-SA 4.0, which restricts commercial use.

Pinakamahusay para sa: Content creation with cloned voices and emotional control

Mag-browse ng lahat Spark TTS Mga boses

Sa isang sulyap

Developer
SparkAudio
Lisensya
CC BY-NC-SA 4.0
Mga hayop
standard
Bilis
medium
Pag-clone ng boses
Oo
Wika
English, Chinese
Max character
1000

Spark TTS Mga boses

Chinese Default

Chinese
Pangkalahatang Neutral

Default

English
Pangkalahatang Neutral

Spark TTS Mga katanungan at sagot

It uses a prompt-based control system layered on top of voice cloning, so you can adjust emotion, speed, and speaking style while preserving the identity of the cloned voice.

About five seconds of reference audio is enough to clone a voice in English or Chinese.

Its model weights are licensed CC BY-NC-SA 4.0, which prohibits commercial use, even though the project code is Apache 2.0. Choose a permissively-licensed model for commercial work.
← Lahat ng mga boses