Darwin TTS

Darwin TTS TTS

A Qwen3-TTS variant whose talker FFN weights are blended from the Qwen3 language model for sharper cross-lingual cloning.

Cláraigh anois! le haghaidh teorainn 5,000 carachtar

Cuir do théacs i gclibeanna SSML le haghaidh rialú beacht:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Clibeanna a thuigeann an tsamhail roghnaithe — cliceáil chun ceann a scaoileadh isteach i do théacs nuair a tharlaíonn sé:

Léann an tsamhail seo gnáth- théacs, mar sin déantar neamhaird ar chlibeanna inlíne. Chun mothúchán clibbhunaithe a chruthú, athraigh go samhail léiritheach cosúil le Orpheus nó Bark.

Sainmhínigh fuaimniú saincheaptha (focal = fuaimniú):

-12 +12
0.5x 2.0x
Saor in Aisce le Piper, VITS, MeloTTS
Taispeánfar an fhuaim a ghintear anseo. Roghnaigh samhail, iontráil téacs, agus cliceáil Giniúint.
D' éirigh le giniúint na fuaime
0:00
Íosluchtaigh Fuaim Íoslódáil.srt Téann an nasc in éag i 24h
Ciseal saor in aisce: úsáid phearsanta. Ceadúnas Tráchtála ó $ 5 / mo
Leabaigh an Tweet Ag tabhairt freagra ar friends!

Eolas Faoi Darwin TTS

Darwin-TTS-1.7B-Cross by FINAL-Bench is a research variant of Qwen3-TTS-1.7B with an unusual construction: 84 of its talker-FFN tensors (about 8.6% of them) are blended at a 3% ratio with the matching tensors from Qwen3-1.7B-Base, all without any retraining. The result is a model that produces noticeably crisper cross-lingual voice cloning across Korean, English, Japanese, and Chinese — its four core languages. It operates in zero-shot voice-clone mode, needing only about three seconds of reference audio to capture a speaker. Darwin is best suited to transferring a single reference voice across those four languages, for example dubbing or multilingual narration with consistent speaker identity.

Is Fearr le haghaidh: Cross-lingual voice cloning between English / Korean / Japanese / Chinese with a single reference voice

Brabhsáil Uile Darwin TTS guthanna

Ag Sracfhéachaint

Forbróir
FINAL-Bench
Ceadúnas
Apache 2.0
Tír
standard
Luas
medium
Clónáil gutha
Teangacha
English, Korean, Japanese, Chinese
Carachtair Uasta
2000

Darwin TTS guthanna

Default

English
Caighdeán Neutral

Default (Chinese)

Chinese
Caighdeán Neutral

Default (Japanese)

Japanese
Caighdeán Neutral

Default (Korean)

Korean
Caighdeán Neutral

Darwin TTS TTS - Ceisteanna Coitianta

Darwin starts from Qwen3-TTS-1.7B but blends a small fraction of its talker-FFN weights with the matching weights from the Qwen3-1.7B base language model. This training-free blend sharpens cross-lingual voice cloning rather than changing the base voices.

English, Korean, Japanese, and Chinese. The FINAL-Bench release specifically markets its cross-lingual blend for those four, and the deployed model ships voices for them.

About three seconds. It works in zero-shot mode, so no fine-tuning or training is required — you provide a short reference clip and it generates new speech in that voice.
← Gach guth