Darwin TTS

Darwin TTS TTS

A Qwen3-TTS variant whose talker FFN weights are blended from the Qwen3 language model for sharper cross-lingual cloning.

Bhalisa Uluhlu lwezinto zobumnini Zolwaleko...

Ulawulo oluchanekileyo:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Ii-tags imodeli ekhethiweyo iqonda - nqakraza ukushiya enye kumbhalo wakho apho isenza khona:

Le modeli ifunda umbhalo oqhelekileyo, ngoko ke i-inline tags ilahleka. Uphawu olusekelwe kwi-emotions, tshintshela kwimodeli ebonisa umbono njenge-Orpheus okanye i-Bark.

Chaza ubeko lwephepha

-12 +12
0.5x 2.0x
Ikhululekile nge Piper, VITS, MeloTTS
Isandi sakho esivelisweyo siza kuvela apha. Khetha imodeli, ngenisa umbhalo, kwaye unqakraze Yenza.
Isandi Sizaliswe Ngempumelelo
0:00
Layisha ezantsi Layisha ezantsi Ikhonkco liphelelwe lixesha kwiyure ezi-24
Inqanaba elikhululekileyo: ukusetyenziswa komuntu siqu. Ilayisensi yezorhwebo ukusuka kwi- $5/inyanga
Uthando TTS.ai? Nceda utshele abalandeli bakho!

I-About Darwin TTS

Darwin-TTS-1.7B-Cross by FINAL-Bench is a research variant of Qwen3-TTS-1.7B with an unusual construction: 84 of its talker-FFN tensors (about 8.6% of them) are blended at a 3% ratio with the matching tensors from Qwen3-1.7B-Base, all without any retraining. The result is a model that produces noticeably crisper cross-lingual voice cloning across Korean, English, Japanese, and Chinese — its four core languages. It operates in zero-shot voice-clone mode, needing only about three seconds of reference audio to capture a speaker. Darwin is best suited to transferring a single reference voice across those four languages, for example dubbing or multilingual narration with consistent speaker identity.

Elungileyo: Cross-lingual voice cloning between English / Korean / Japanese / Chinese with a single reference voice

Khangela konke Darwin TTS iilizwi

Kwingxelo

Umbhekisi phambili
FINAL-Bench
Ilayisensi
Apache 2.0
I-Tier
standard
Isantya
medium
Ukuphinda usebenzise ilizwi
Ewe
Iilwimi
English, Korean, Japanese, Chinese
Ubukhulu bamagama
2000

Darwin TTS iilizwi

Default

English
Emiselweyo Neutral

Default (Chinese)

Chinese
Emiselweyo Neutral

Default (Japanese)

Japanese
Emiselweyo Neutral

Default (Korean)

Korean
Emiselweyo Neutral

Darwin TTS TTS - Imibuzo ebuzwa rhoqo

Darwin starts from Qwen3-TTS-1.7B but blends a small fraction of its talker-FFN weights with the matching weights from the Qwen3-1.7B base language model. This training-free blend sharpens cross-lingual voice cloning rather than changing the base voices.

English, Korean, Japanese, and Chinese. The FINAL-Bench release specifically markets its cross-lingual blend for those four, and the deployed model ships voices for them.

About three seconds. It works in zero-shot mode, so no fine-tuning or training is required — you provide a short reference clip and it generates new speech in that voice.
← Zonke iingoma