VibeVoice TTS
Microsoft's multi-speaker long-form model that generates up to 90 minutes with 4 distinct speakers.
Lập vòng văn bản trong thẻ SSML để kiểm soát chính xác:
<speak><prosody rate="slow">Slow speech</prosody></speak>
Thẻ mà mô hình đã chọn hiểu — nhấn để thả một trong văn bản của bạn nơi nó xảy ra:
Mô hình này đọc văn bản đơn giản, vì vậy các thẻ trong dòng sẽ bị bỏ qua. Đối với cảm xúc dựa trên thẻ, hãy chuyển sang mô hình biểu cảm như Orpheus hay Bark.
Định nghĩa cách phát âm tùy chỉnh (từ = phát âm):
Về VibeVoice
VibeVoice from Microsoft is built for long-form, multi-speaker audio. Its 1.5B model can generate up to 90 minutes of speech with as many as 4 simultaneous speakers, using speaker tags to drive multi-turn dialogue — a strong fit for podcasts, audiobooks, and conversations that need speaker consistency across long passages. A separate Realtime 0.5B variant reaches roughly 300ms latency for interactive use. On TTS.ai it covers English and Chinese and accepts up to 50,000 characters per request, so an entire episode can be scripted in one pass.
Tốt nhất cho: Podcasts, dialogues, long-form narration, multi-speaker content
& Xem tất cả VibeVoice giọng nóiMột cái nhìn
- Nhà phát triển
- Microsoft
- Giấy phép
- MIT
- Thú
- standard
- Tốc độ
- fast
- Ký âm
- Không
- Ngôn ngữ
- English, Chinese
- Tối đa các ký tự
- 50000