Darwin TTS

Darwin TTS TTS

A Qwen3-TTS variant whose talker FFN weights are blended from the Qwen3 language model for sharper cross-lingual cloning.

Whakawhanake mō te tepe o ngā tohu 5,000

Whāriki i tōna kupu i roto i ngā tohu SSML mō te whakahaere tika:

<speak><prosody rate="slow">Slow speech</prosody></speak>

E mōhio ana ngā tohu ki te tauira i kōwhiria - ka kōwhiria kia whakawātea tētahi ki roto i tōna kupu i reira ka puta ai:

Ka pānui tēnei tauira i te kupu noa, nā reira ka whakakāhoretia ngā tohu ā-waitara. Mō te āhua o te tohu-taihi, ka huri ki tētahi tauira whakamārama pēnei i a Orpheus, Bark rānei.

Ka tautuhia ngā tohutohu ā-ringa (wāhi = tohutohu):

-12 +12
0.5x 2.0x
Waihoki me Piper, VITS, MeloTTS
Ka puta tēnei te oro i waihangatia e koe. Ka kōwhiria tētahi tauira, ka tāurua te kupu, a, ka kōwhiria te Whakatū.
Kua angitu te whakaputanga oro
0:00
Waihoki i te oro Whakahua.srt Ka ngaro te pānga i roto i te 24h
Tauwhāinga wātea: te whakamahinga whaiaro. Whakawhiwhinga hokohoko mai i te $5/mo
E manakohia ana e TTS.ai? Whakapāpāho ki ōna hoa!

Mo Darwin TTS

Darwin-TTS-1.7B-Cross by FINAL-Bench is a research variant of Qwen3-TTS-1.7B with an unusual construction: 84 of its talker-FFN tensors (about 8.6% of them) are blended at a 3% ratio with the matching tensors from Qwen3-1.7B-Base, all without any retraining. The result is a model that produces noticeably crisper cross-lingual voice cloning across Korean, English, Japanese, and Chinese — its four core languages. It operates in zero-shot voice-clone mode, needing only about three seconds of reference audio to capture a speaker. Darwin is best suited to transferring a single reference voice across those four languages, for example dubbing or multilingual narration with consistent speaker identity.

Pai mo: Cross-lingual voice cloning between English / Korean / Japanese / Chinese with a single reference voice

Ka tirohia katoa Darwin TTS ngā oro

I te tirohanga

Ka whakawhanakehia
FINAL-Bench
Ka taea te whakawātea
Apache 2.0
Karaka
standard
Āhuatanga
medium
Whakakōrero reo
He
reo
English, Korean, Japanese, Chinese
Kāri nui rawa
2000

Darwin TTS ngā oro

Default

English
Paerewa Neutral

Default (Chinese)

Chinese
Paerewa Neutral

Default (Japanese)

Japanese
Paerewa Neutral

Default (Korean)

Korean
Paerewa Neutral

Darwin TTS TTS - FAQ

Darwin starts from Qwen3-TTS-1.7B but blends a small fraction of its talker-FFN weights with the matching weights from the Qwen3-1.7B base language model. This training-free blend sharpens cross-lingual voice cloning rather than changing the base voices.

English, Korean, Japanese, and Chinese. The FINAL-Bench release specifically markets its cross-lingual blend for those four, and the deployed model ships voices for them.

About three seconds. It works in zero-shot mode, so no fine-tuning or training is required — you provide a short reference clip and it generates new speech in that voice.
← Ko nga oro katoa