E mōhio ana ngā tohu ki te tauira i kōwhiria - ka kōwhiria kia whakawātea tētahi ki roto i tōna kupu i reira ka puta ai:
Ka pānui tēnei tauira i te kupu noa, nā reira ka whakakāhoretia ngā tohu ā-waitara. Mō te āhua o te tohu-taihi, ka huri ki tētahi tauira whakamārama pēnei i a Orpheus, Bark rānei.
Ko ngā tohu whakawhiwhinga reo māori. I whakawhiwhia e Qwen3-TTS i tēnei rā; Ka tae mai ētahi atu tauira whakamāramatanga.
Kāore tēnei tauira i te tautoko i ngā tohutohu kāhua — ka huri ki tētahi tauira e mahi ana (hei tauira, Qwen3-TTS).
Ka tautuhia ngā tohutohu ā-ringa (wāhi = tohutohu):
-12
+12
Ka whakahaeretia te whakahuahua (0 = pūmau, 1 = tino whakahuahua)
Te taumahatanga tohu whakawātea-whakahaere (nui ake = nui ake te whakahau-whakahaere)
E whakaahua ana i te kāhua reo i roto i te reo māori (ka whakamahia e te Pāpāho tēnei i te wāhi o ngā reo i whakaritea i mua)
He pūāhua taupānga te taupānga taupānga: Ka whakamahia ngā tohu [S1] me [S2] hei tohu i ngā kaikōrero rerekē. Hei tauira: [S1] Hei hei! [S2] Hei, he pēhea koe?
FreyaTTS-small is a 183-million-parameter model built for one language and built well. It is a non-autoregressive conditional flow-matching diffusion transformer that reads Turkish directly at the character level — 92 symbols, no phonemizer and no grapheme-to-phoneme stage, which removes a whole class of mispronunciation that pronunciation dictionaries introduce. It generates in a frozen AudioVAE2 latent space and decodes to 48 kHz mono, more than double the sample rate of the piper Turkish voice, so the output carries treble detail that a 22 kHz model simply cannot represent. On the Freya-TR-Eval benchmark it reaches 8.0% word error rate, placing it ahead of both XTTS-v2 and F5-TTS among open sub-billion-parameter Turkish systems, and it runs fast enough for real-time use at roughly a tenth of real time.
Pai mo: Turkish narration, voice agents, and any Turkish audio that needs high sample-rate output
It was trained from scratch on Turkish speech alone rather than adapted from a multilingual model. Specialising lets a 183M-parameter model compete with far larger multilingual systems on Turkish, but it means the model has no ability to read other languages — requests in another language are rejected rather than mispronounced.
Most open TTS models emit 22.05 kHz or 24 kHz, which caps reproducible audio at around 11-12 kHz and audibly dulls sibilants. FreyaTTS decodes to 48 kHz, the standard sample rate for video and broadcast, so its output drops into a production timeline without upsampling.
No. FreyaTTS has a single fixed speaker and does not support cloning. For a cloned Turkish voice use one of our zero-shot cloning models instead.
Many TTS systems first convert text into phonemes using a pronunciation dictionary, and anything missing from that dictionary — new words, names, loanwords — gets guessed. FreyaTTS reads the characters themselves, so Turkish spelling, which is highly regular, maps to sound without that lossy middle step.
E tāuru ana i tōna mairā hei haere tonu i te whakaputa kōrero mō te wāteatanga, hei whakaingoatanga rānei mō ngā pūāhua tāpiri 15,000.
Kāore he kupu whakarongo e hiahiatia ana — ka tukuna mai e tātau he pātahitanga ki te tautuhi i tētahi i muri ake nei.
Kua tae te tepe taumata wātea
Kua whakamahia e koe ōna pūāhua wātea 5,000 i ia rā. Whakatū i tētahi kāwanatanga wātea kia whiwhi ai ki ngā pūāhua tāpiri 15,000 tae atu ki ngā pūāhua wātea 10,000 i ia marama.