Chinese (Mandarin) Text-zu-Sprooch

Turn Chinese (Mandarin) Text an natierlech Sprooch mat AI Stimmen ëmwandelen. 25 Stimmen. D'Lidd ass gratis ze downloaden an ass als MP3 oder WAV verfügbar.

Anmelden Limit fir 5. 000 Zeichen

Wrap your text in SSML tags for precise control:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Tags déi d'gewielt Modell verstinn - klickt fir eng an Ärem Text ze setzen wou se geschitt:

Dëse Modell liest einfache Text, sou datt Inline-Tags ignoréiert ginn. Fir Tag-baséiert Emotiounen, wielt e expressiven Modell wéi Orpheus oder Bark.

Eegen Aussproochen definéieren (Wuert = Aussprooch):

-12 +12
0.5x 2.0x
Free mat Piper, VITS, MeloTTS
Äert generéiert Audio wäert hei erscheinen. Wielt e Modell, gitt Text an a klickt op Generéieren.
Audio gouf erfollegräich generéiert
0:00
Audio erofgelueden Lëscht vu lëtzebuergesche Schrëftsteller Link expires in 24h
Den Haaptuert ass Personnes. Kommerziell Lizenz vun $5/mo
Maacht dat Är eege Stëmm Klonen eng Stëmm an 30 Sekonnen
Liewe TTS.ai? Erzielt Är Frënn!

Iwwer Chinese (Mandarin) Text- op- Sprooch

Mandarin text-to-speech lives or dies on tone: it has four lexical tones plus a neutral tone, and getting the contour wrong turns "mā" (mother) into "mǎ" (horse), so the model must predict pitch per syllable, not just per sentence. Tone sandhi adds another layer — for example two third tones in a row shift the first to a rising tone, and the words "一" (yī) and "不" (bù) change tone depending on what follows. Because Hanzi carry no spaces and many characters are polyphonic (多音字), high-quality Chinese synthesis depends heavily on word segmentation and grapheme-to-phoneme disambiguation from context.

Sample — 中文(普通话)

“今天天气很好,我们一起去公园散步,顺便买点水果回家吧。”

Natierlecher Numm
中文(普通话)
Lautsprecher
about 1.1 billion speakers (roughly 920 million native Mandarin)
Sproochen
Sinitic branch of Sino-Tibetan
Skript
Chinese characters (Hanzi) — Simplified and Traditional
Spuenesch
Mainland China, Taiwan, Singapore, Malaysia, Hong Kong, global Chinese diaspora

25 Chinese (Mandarin) Stimmen

Chinese Speaker 1

Bark
Standard Neutral

Chinese Speaker 2

Bark
Standard Neutral

Chinese Speaker

Bark Small
Standard Neutral

Chinese Female

CosyVoice 2
Standard Female

Chinese Male

CosyVoice 2
Standard Male

Chinese Female

CosyVoice3
Standard Female

Chinese Male

CosyVoice3
Standard Male

Default (Chinese)

Darwin TTS
Standard Neutral

Default

GPT-SoVITS
Standard Neutral

Chinese Default

IndexTTS-2
Standard Neutral

Xiaobei

Kokoro
Fräi Female

Xiaoni

Kokoro
Fräi Female

Xiaoxiao

Kokoro
Fräi Female

Yunjian

Kokoro
Fräi Male

Chinese

MeloTTS
Fräi Female

Default (Chinese)

Ming-Omni TTS
Fräi Neutral

Chinese

MOSS-TTS Nano
Standard Neutral

Default (Chinese)

MOSS-TTSD
Standard Neutral

Chinese

OpenVoice
Premium Neutral

Huayan (Chinese)

Piper
Fräi Female

Uncle Fu

Qwen3 TTS
Standard Male

Chinese Default

Spark TTS
Standard Neutral

Speaker 1 (Chinese)

VibeVoice
Standard Neutral

Speaker 2 (Chinese)

VibeVoice
Standard Neutral

Default Chinese

VoxCPM
Standard Neutral

Wat Leit benotzen Chinese (Mandarin) Text zu Sprooch fir

E-learning and Mandarin language-teaching narration
Short-video (Douyin/Bilibili) and livestream voiceover
Navigation and in-car voice prompts
Customer-service IVR and chatbot voices
News and audiobook narration for the diaspora

Chinese (Mandarin) Text zu Sprooch

Yes. You can paste either Simplified (mainland/Singapore) or Traditional (Taiwan/Hong Kong) text; both are read in Mandarin pronunciation.

The model predicts each syllable's tone contour from context and applies tone sandhi rules, so sequences like third-tone pairs and the special cases of 一 and 不 come out naturally.

Mostly yes. Multi-reading characters such as 行 (xíng vs háng) or 长 (cháng vs zhǎng) are disambiguated from surrounding words, though rare proper nouns can still be ambiguous.

These voices are Standard Mandarin (Putonghua). Cantonese uses a different tone system and pronunciation and is not the same as Mandarin TTS.

Sproochen