VITS

VITS 음성 인식

The end-to-end TTS architecture that combines a variational autoencoder, normalizing flows, and adversarial training.

가입하기 5,000자 한도

정확한 제어를 위해 SSML 태그로 텍스트를 래핑하십시오.

<speak><prosody rate="slow">Slow speech</prosody></speak>

선택한 모델이 이해하는 태그 — 텍스트에 드래그하려면 클릭하세요:

이 모델은 일반 텍스트를 읽기 때문에 인라인 태그는 무시됩니다. 태그 기반 감정을 위해서는 Orpheus 또는 Bark과 같은 표현 모델로 전환하십시오.

사용자 지정 발음 정의 (단어 = 발음):

-12 +12
0.5x 2.0x
파이퍼, VITS, MeloTTS와 무료
생성된 오디오가 여기에 나타납니다. 모델을 선택하고 텍스트를 입력한 다음 생성 을 클릭합니다.
오디오가 성공적으로 생성되었습니다
0:00
오디오 다운로드 .srt 파일 다운로드 링크는 24시간 이내에 만료됩니다.
무료 계층: 개인용. 상업용 라이센스 최저 $5/mo
이것을 당신의 목소리로 만들어라 30초만에 목소리 복제
TTS.ai가 마음에 드시나요? 친구들에게 알려주세요!

정보 VITS

VITS — Variational Inference with adversarial learning for end-to-end Text-to-Speech — was introduced by Jaehyeon Kim and collaborators in 2021 and became a foundational architecture for modern neural speech. Rather than the older two-stage pipeline, it synthesizes audio in a single parallel end-to-end pass, pairing a variational autoencoder with normalizing flows and a GAN-style adversarial training process to lift naturalness. At about 25M parameters and trained on ~585 hours, it produces natural prosody at fast inference speeds and supports multiple speakers. It serves as a solid general-purpose, free baseline and underpins many later models such as Piper and MeloTTS.

최적화된 용도: General-purpose text-to-speech with natural prosody

모두 찾아보기 VITS 목소리

한눈에

개발자
Jaehyeon Kim et al.
라이선스
MIT
free
속도
fast
음성 복제
아니요
언어
English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, Polish
최대 문자수
2000

VITS 목소리

CSS10 (Dutch)

Dutch
자유 Neutral

CSS10 (Finnish)

Finnish
자유 Neutral

CSS10 (French)

French
자유 Neutral

CSS10 (German)

German
자유 Neutral

CSS10 (Hungarian)

Hungarian
자유 Neutral

CSS10 (Spanish)

Spanish
자유 Neutral

Common Voice (Bulgarian)

Bulgarian
자유 Neutral

Common Voice (Portuguese)

Portuguese
자유 Neutral

Default

English
자유 Neutral

MAI (Polish)

Polish
자유 Female

MAI (Ukrainian)

Ukrainian
자유 Neutral

VITS TTS — 자주 묻는 질문

VITS means Variational Inference with adversarial learning for end-to-end Text-to-Speech. It generates audio in a single parallel pass using a variational autoencoder, normalizing flows, and adversarial (GAN) training, rather than a two-stage pipeline.

Yes. VITS is MIT-licensed and in the free tier, so it can be used commercially.

On TTS.ai, VITS covers 11 languages including English, German, Spanish, French, Portuguese, Dutch, Finnish, Hungarian, Bulgarian, Japanese, and Polish, with multi-speaker support. It does not do voice cloning.
← 모든 음성