Darwin TTS

Darwin TTS TTS

A Qwen3-TTS variant whose talker FFN weights are blended from the Qwen3 language model for sharper cross-lingual cloning.

אַרײַנשרײַבן פֿאַר 5,000 שריפֿטצײכן

איבער־פֿאַרקער דעם טעקסט אין SSML הענטלעך פֿאַר אַ פּשוטער קאָנטראָל:

<speak><prosody rate="slow">Slow speech</prosody></speak>

הענטלעך װאָס דער אויסגעקליבן מאָדעל פֿאַרשטײט — קליק צו אַרײַנשרײַבן אײן אין װײַז־טעקסט װוּ עס פּאַסט זיך:

דאָס מאָדעל לייענט נאָרמאַלן טעקסט, אַזוי אַרײַנגעלייגטע הענטלעך ווערן איגנאָרירט. פֿאַר הענטלעך־באזירטע װײַב־איגנאָרירונג, װײַז צו אַ װײַז־מאָדעל װי אורפיאָ אָדער װאַרק

װײַז פֿאָרױסװײַז

-12 +12
0.5x 2.0x
פֿרײַ מיט Piper, VITS, MeloTTS
די אױדיו־טעקע װעט דאָ װײַזן זיך. קלײַב אַ מאָדעל אױס, אַרײַנשרײַב דעם טעקסט און קליק אױף אױסשרײַבן
אוודיאָ אױסגעגרײט
0:00
דאַונלאָוד אַרײַנשטעלן.srt פֿאַרבינדונג ענדיקט זיך אין 24 שעה
פרייע מדרגה: פּערזענלעכער ניצן קאָממערשעל ליסענסע פֿון $5/חודש
ליבע TTS.ai? זאָגן דיין פריינט

אױף Darwin TTS

Darwin-TTS-1.7B-Cross by FINAL-Bench is a research variant of Qwen3-TTS-1.7B with an unusual construction: 84 of its talker-FFN tensors (about 8.6% of them) are blended at a 3% ratio with the matching tensors from Qwen3-1.7B-Base, all without any retraining. The result is a model that produces noticeably crisper cross-lingual voice cloning across Korean, English, Japanese, and Chinese — its four core languages. It operates in zero-shot voice-clone mode, needing only about three seconds of reference audio to capture a speaker. Darwin is best suited to transferring a single reference voice across those four languages, for example dubbing or multilingual narration with consistent speaker identity.

בעסטער פֿאַר: Cross-lingual voice cloning between English / Korean / Japanese / Chinese with a single reference voice

בלעטער Darwin TTS שריפֿטצײכן

אין אַ שריט

אַנטוויקלער
FINAL-Bench
לינקס
Apache 2.0
װײַז
standard
שאַטירונג
medium
שפּראַך
יאָ
שפּראַכן
English, Korean, Japanese, Chinese
גרעסטע שריפֿטצײכן
2000

Darwin TTS שריפֿטצײכן

Default

English
סטאַנדאַרד Neutral

Default (Chinese)

Chinese
סטאַנדאַרד Neutral

Default (Japanese)

Japanese
סטאַנדאַרד Neutral

Default (Korean)

Korean
סטאַנדאַרד Neutral

Darwin TTS FAQ

Darwin starts from Qwen3-TTS-1.7B but blends a small fraction of its talker-FFN weights with the matching weights from the Qwen3-1.7B base language model. This training-free blend sharpens cross-lingual voice cloning rather than changing the base voices.

English, Korean, Japanese, and Chinese. The FINAL-Bench release specifically markets its cross-lingual blend for those four, and the deployed model ships voices for them.

About three seconds. It works in zero-shot mode, so no fine-tuning or training is required — you provide a short reference clip and it generates new speech in that voice.
← אַלע שפּראַכן