Bark

Bark TTS

Suno's transformer-based text-to-audio model that generates speech plus laughter, sighs, music, and sound effects.

Ṣẹ̀dà fun àwọn àmì-àṣírí 5,000

Fi àkọlé rẹ pamọ́ sí àwọn àmì-ìwé SSML fún ìdáràn:

<speak><prosody rate="slow">Slow speech</prosody></speak>

Àwọn Àmì-ìwé tí àwọn ìṣàmúlò-ètò tí a yàn gbọ́ - tẹ̀ láti fi ọkan sínú àkọ́lé rẹ̀ nínú àwọn ààyè-iṣẹ́ tí o bá jẹ́:

Àwọn àwọn àkọlé àwọn ààyè-iṣẹ́ àwọn àwọn àmì-ìwé àwọn àmì-ìwé àwọn àwọn àmì-ìwé àwọn à

Àwọn àwọn ìṣàfarawé àwọn àwọn ìṣàfarawé àwọn (ọrọ = ìṣàfàlì):

-12 +12
0.5x 2.0x
Free pẹlu Piper, VITS, MeloTTS
Àwọn àwòrán tí o ti ṣẹ̀dà tí o bá han níbẹ̀. Yan àwọn àwòrán, tẹ̀lẹ̀ àkọlé, ki o si tẹ̀ Ṣẹ̀dà.
Àwọn àwọn àwòrán tí a ṣẹ̀dà
0:00
Ṣàfikún Àwọn Àmì-ìwé Ṣàfikún.srt Líǹkì náà kù nínú 24h
Ìjádé ọ̀fẹ́: ìlòjútó ara ẹni. Lisensi Iṣowo ori lati $5/mo
O fẹ́ TTS.ai? Fì sọ̀kalẹ̀ fún àwọn ọrẹ̀ rẹ̀!

Ààyè-iṣẹ́ Bark

Bark comes from Suno and takes a different approach from most TTS systems: it is a GPT-style transformer trained as a text-to-audio model rather than a pure text-to-speech one. Because it generates raw audio tokens, it can produce nonverbal sounds — laughing, sighing, crying — as well as background music and sound effects alongside the spoken words. It ships with 100+ speaker presets and handles 13+ languages including English, Chinese, French, German, Hindi, Japanese, and Korean. The trade-off is speed and length: at 350M parameters it runs slowly (~15s per clip) and caps at 200 characters, so it shines for short, emotive, creative audio rather than long narration.

Tí o dara jù fún: Creative audio content, audiobooks with emotion, sound effects

Wá Gbogbo àwòrán Bark Àwọn àwòrán

Nínú àwọn ìṣàfarawé

Àwọn Àkọlé
Suno
Àwọn Ààyè-iṣẹ́
MIT
Àwọn àwọn ààyè-iṣẹ́
standard
Ìjánu-ìṣàmúlò-ètò
slow
Ìṣàfarawé àwọn àmì-ìwé
Àwọn àwọn àgbéwọlé
Àwọn
English, Chinese, French, German, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Turkish
Àwọn àyọkà ìpele
200

Bark Àwọn àwòrán

Chinese Speaker 1

Chinese
Àwọn ìpéwọ̀n Neutral

Chinese Speaker 2

Chinese
Àwọn ìpéwọ̀n Neutral

English Female 1

English
Àwọn ìpéwọ̀n Female

English Female 2

English
Àwọn ìpéwọ̀n Female

English Female 3

English
Àwọn ìpéwọ̀n Female

English Female 4

English
Àwọn ìpéwọ̀n Female

English Male 1

English
Àwọn ìpéwọ̀n Male

English Male 2

English
Àwọn ìpéwọ̀n Male

English Male 3

English
Àwọn ìpéwọ̀n Male

English Male 4

English
Àwọn ìpéwọ̀n Male

English Male 5

English
Àwọn ìpéwọ̀n Male

English Male 6

English
Àwọn ìpéwọ̀n Male

French Speaker 1

French
Àwọn ìpéwọ̀n Neutral

French Speaker 2

French
Àwọn ìpéwọ̀n Neutral

German Speaker 1

German
Àwọn ìpéwọ̀n Neutral

German Speaker 2

German
Àwọn ìpéwọ̀n Neutral

Hindi Speaker 1

Hindi
Àwọn ìpéwọ̀n Neutral

Italian Speaker 1

Italian
Àwọn ìpéwọ̀n Neutral

Japanese Speaker 1

Japanese
Àwọn ìpéwọ̀n Neutral

Japanese Speaker 2

Japanese
Àwọn ìpéwọ̀n Neutral

Korean Speaker 1

Korean
Àwọn ìpéwọ̀n Neutral

Korean Speaker 2

Korean
Àwọn ìpéwọ̀n Neutral

Polish Speaker 1

Polish
Àwọn ìpéwọ̀n Neutral

Portuguese Speaker 1

Portuguese
Àwọn ìpéwọ̀n Neutral

Russian Speaker 1

Russian
Àwọn ìpéwọ̀n Neutral

Spanish Speaker 1

Spanish
Àwọn ìpéwọ̀n Neutral

Spanish Speaker 2

Spanish
Àwọn ìpéwọ̀n Neutral

Turkish Speaker 1

Turkish
Àwọn ìpéwọ̀n Neutral

Bark Àwọn Àtòjọ-ẹ̀yàn

Yes. Bark is a text-to-audio model, so beyond speech it can generate nonverbal cues like laughing, sighing and crying, plus music and background sound effects — one of its defining capabilities.

Yes. Bark is MIT-licensed, which permits commercial use.

Bark caps at 200 characters per request and is on the slower side (around 15 seconds per clip), so it is best suited to short, expressive snippets rather than long-form audio. It does not support voice cloning.
← Gbogbo àwọn ìrànwọ́