VieNeu-TTS-v2

VieNeu-TTS-v2 Ikiganiro

A Vietnamese-first, CPU-only model with en-vi code-switching, 7 regional preset voices, and zero-shot cloning.

Kwiyandikisha kugirango Inyuguti

Umwandiko in Itagi: ya: Igenzura:

<speak><prosody rate="slow">Slow speech</prosody></speak>

i Byahiswemo Urugero - Kanda Kuri Gukuraho Rimwe Umwandiko:

Urugero: Bisanzwe Umwandiko, Umurongo: Itagi:. Itagi: -, Hindura Kuri Urugero: Nka Cyangwa.

Kugena (Ijambo =):

-12 +12
0.5x 2.0x
Na:,,
Audio Kugaragara. A Urugero:, Injiza Umwandiko, na Kanda.
Byaremwe
0:00
Iyimura Iyimura Ihuza in
Bwite Koresha Kuva: 5 /
TTS.ai? Abayobozi!

Ikiganiro VieNeu-TTS-v2

VieNeu-TTS-v2 is a 300M-parameter Vietnamese-first model built on a Qwen3 backbone and trained on more than 10,000 hours of bilingual data. It handles seamless English-Vietnamese code-switching, ships 7 preset voices spanning Northern and Southern accents, and clones a voice instantly from just 3-5 seconds of reference audio. Notably it runs entirely on CPU — using GGUF Q4 inference plus an ONNX audio decoder — with no GPU required, finishing a generation in about 7 seconds. It's purpose-built for Vietnamese content and bilingual en-vi narration, an underserved niche in open TTS.

kugirango: Vietnamese content and bilingual en-vi narration

Gushakisha byose VieNeu-TTS-v2 Amashusho

A

Mukoraporogaramu
Phạm Nguyễn Ngọc Bảo
Inyandiko y'Iyemererakoresha
Apache 2.0
Itariki
standard
Umuvuduko
fast
Guhindura izina
Oya
Ururimi:
Vietnamese, English
Inyuguti
1000

VieNeu-TTS-v2 Amashusho

Bích Ngọc (North, Female)

Vietnamese
Bisanzwe Female

Phạm Tuyên (North, Male)

Vietnamese
Bisanzwe Male

Thanh Bình (North, Male)

Vietnamese
Bisanzwe Male

Thái Sơn (South, Male)

Vietnamese
Bisanzwe Male

Thục Đoan (South, Female)

Vietnamese
Bisanzwe Female

Trúc Ly (North, Female)

Vietnamese
Bisanzwe Female

Xuân Vĩnh (South, Male)

Vietnamese
Bisanzwe Male

VieNeu-TTS-v2 -

Yes. VieNeu-TTS-v2 runs entirely on CPU via GGUF Q4 inference and an ONNX audio decoder — no GPU needed — and completes a generation in around 7 seconds.

It is Vietnamese-first with English support and seamless en-vi code-switching. It ships 7 preset voices spanning Northern and Southern Vietnamese accents.

Yes. It supports instant zero-shot voice cloning from just 3-5 seconds of reference audio. It is Apache 2.0 licensed and free to use commercially.
← Amashusho yose