TTS Arena-AI 声音模型板

与20+文本对语音模型相比,官方基准、社区评级和并肩比较。

逐边比较

文本类型、选择两个模型并比较结果。 自由级模型不需要记账 。

免费的模特儿在没有账户的情况下工作。 签名 比较溢价模型。

模型板

# 型 正式官方 社区 你的评分 速度 级别
1
Kokoro
Kokoro
Lightweight 82M parameter model delivering studio-quality speech with blazing-fast inference.
82M 1200h 2024
4.8 /5 5.0 /5
1 表决票、票票票
fast Free
2
CosyVoice 2
CosyVoice 2
Alibaba's scalable streaming TTS with human-parity naturalness and near-zero latency.
300M 200000h 2024
4.26 /5 尚未表决
medium Standard
3
Chatterbox
Chatterbox
State-of-the-art zero-shot voice cloning with emotion control from Resemble AI.
300M 2025
4.25 /5 尚未表决
medium Premium
4
StyleTTS 2
StyleTTS 2
Human-level text-to-speech through style diffusion and adversarial training.
100M 585h 2024
4.23 /5 尚未表决
medium Premium
5
Piper
Piper
A fast, local neural text to speech system optimized for Raspberry Pi and embedded devices.
15M 2023
4.15 /5 尚未表决
fast Free
6
MeloTTS
MeloTTS
High-quality multilingual text-to-speech that runs on CPU with minimal latency.
25M 2024
4.13 /5 尚未表决
fast Free
7
Dia TTS
Dia TTS
Multi-speaker dialog generation model that creates natural conversations between speakers.
1.6B 2024
4.09 /5 尚未表决
medium Standard
8
VITS
VITS
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.
25M 585h 2021
4.0 /5 尚未表决
fast Free
9
Orpheus
Orpheus
Human-level emotional TTS model trained on 100K hours of speech data.
3B 100000h 2025
4.0 /5 尚未表决
medium Standard
10
OpenVoice
OpenVoice
Instant voice cloning with granular control over style, emotion, and accent.
300M 2024
4.0 /5 尚未表决
medium Premium
11
IndexTTS-2
IndexTTS-2
Zero-shot TTS with fine-grained emotion control and high expressiveness.
300M 2025
3.91 /5 尚未表决
medium Standard
12
Spark TTS
Spark TTS
Voice cloning TTS with controllable emotion and speaking style via prompts.
500M 2025
3.9 /5 尚未表决
medium Standard
13
Parler TTS
Parler TTS
Describe the voice you want in natural language and Parler generates matching speech.
880M 45000h 2024
3.83 /5 尚未表决
medium Standard
14
Tortoise TTS
Tortoise TTS
Multi-voice text-to-speech focused on quality with autoregressive architecture.
400M 50000h 2022
3.7 /5 尚未表决
slow Premium
15
Bark
Bark
Transformer-based text-to-audio model that generates realistic speech, music, and sound effects.
350M 100000h 2023
3.57 /5 尚未表决
slow Standard
16
Bark Small
Bark Small
Lighter version of Bark with faster inference and lower memory usage.
150M 100000h 2023
— 尚未表决
medium Standard
17
Indic Parler TTS
Indic Parler TTS
High-quality speech for 8+ Indian languages with natural-language voice control.
900M 8000h 2024
— 尚未表决
slow Standard
18
KhanomTan TTS
KhanomTan TTS
Thai-first text-to-speech with a choice of speaker voices.
85M 100h 2023
— 尚未表决
fast Standard
19
GPT-SoVITS
GPT-SoVITS
Few-shot voice cloning TTS that replicates any voice from just 5 seconds of audio.
200M 2024
— 尚未表决
slow Standard
20
Qwen3 TTS
Qwen3 TTS
Alibaba's multilingual TTS with preset voices and voice design from text.
1.7B 2025
— 尚未表决
medium Standard
21
VieNeu-TTS-v2
VieNeu-TTS-v2
Vietnamese + English code-switching TTS with 7 preset voices and zero-shot voice cloning. CPU-only, no GPU required.
0.3B 10000h 2026
— 尚未表决
fast Standard
22
Sesame CSM
Sesame CSM
Conversational speech model generating natural dialogue with appropriate timing and emotion.
1B 2025
— 尚未表决
slow Premium
23
Chatterbox Turbo
Chatterbox Turbo
Faster Chatterbox with sub-200ms latency and paralinguistic tags for laughs, coughs, and more.
350M 2025
— 尚未表决
fast Standard
24
VoxCPM
VoxCPM
Tokenizer-free TTS producing 44.1kHz audio with context-aware paragraph consistency.
500M 1800000h 2025
— 尚未表决
fast Standard
25
Kani TTS 2
Kani TTS 2
Ultra-lightweight 400M English TTS model running in just 3GB VRAM.
400M 10000h 2026
— 尚未表决
fast Free
26
OuteTTS
OuteTTS
LLM-based TTS that runs on CPU, GPU, or browser via llama.cpp and Transformers.js.
1B 5000h 2025
— 尚未表决
slow Free
27
VibeVoice
VibeVoice
Microsoft's multi-speaker long-form TTS generating up to 90 minutes with 4 distinct speakers.
1.5B 100000h 2025
— 尚未表决
fast Standard
28
Pocket TTS
Pocket TTS
Lightweight 100M parameter model by Kyutai with voice cloning from a single sample.
100M 50000h 2025
— 尚未表决
fast Free
29
Kitten TTS
Kitten TTS
Ultra-lightweight TTS under 80MB. Runs on CPU without GPU.
80M 2025
— 尚未表决
fast Free
30
CosyVoice3
CosyVoice3
Next-generation multilingual TTS with bi-streaming, emotion control, and zero-shot voice cloning.
500M 200000h 2025
— 尚未表决
fast Standard
31
NAMAA Saudi TTS
NAMAA Saudi TTS
First open Saudi-Arabic TTS. Native Saudi dialect with Chatterbox-quality voice cloning.
300M 2026
— 尚未表决
medium Standard
32
Darwin TTS
Darwin TTS
Cross-modal Qwen3-TTS variant with FFN weights blended from the Qwen3-1.7B language model for sharper multilingual cloning.
2.1B 2026
— 尚未表决
medium Standard
33
MOSS-TTSD
MOSS-TTSD
Multi-speaker dialogue continuation model — generate podcast-style conversations with up to 5 speakers and 60 minutes of coherent audio.
7B 2026
— 尚未表决
medium Standard
34
Ming-Omni TTS
Ming-Omni TTS
Compact 0.5B omni-modal speech model from inclusionAI with high-fidelity 44.1kHz output and zero-shot voice cloning.
500M 2026
— 尚未表决
medium Free
35
MOSS-TTS Nano
MOSS-TTS Nano
Tiny 100M MOSS-TTS variant — same architecture, 80x smaller, free-tier latency.
100M 500000h 2026
— 尚未表决
fast Free
36
FreyaTTS
FreyaTTS
Turkish-only speech at 48 kHz from a compact flow-matching model.
183M 2026
— 尚未表决
fast Free

详细基准分数

官方TTS.ai个基准分数分分三个方面:自然性、准确性和速度。

KokoroKokoro

Free
自然性质 4.8/5
准确性 4.7/5
速度 4.9/5
总体 4.8/5

CosyVoice 2CosyVoice 2

Standard
自然性质 4.5/5
准确性 4.4/5
速度 3.8/5
总体 4.26/5

ChatterboxChatterbox

Premium
自然性质 4.7/5
准确性 4.5/5
速度 3.4/5
总体 4.25/5

StyleTTS 2StyleTTS 2

Premium
自然性质 4.5/5
准确性 4.3/5
速度 3.8/5
总体 4.23/5

PiperPiper

Free
自然性质 3.5/5
准确性 4.2/5
速度 4.95/5
总体 4.15/5

MeloTTSMeloTTS

Free
自然性质 3.8/5
准确性 4.1/5
速度 4.6/5
总体 4.13/5

Dia TTSDia TTS

Standard
自然性质 4.6/5
准确性 4.3/5
速度 3.2/5
总体 4.09/5

VITSVITS

Free
自然性质 3.4/5
准确性 4.0/5
速度 4.8/5
总体 4.0/5

OrpheusOrpheus

Standard
自然性质 4.3/5
准确性 4.1/5
速度 3.5/5
总体 4.0/5

OpenVoiceOpenVoice

Premium
自然性质 4.0/5
准确性 4.1/5
速度 3.9/5
总体 4.0/5

IndexTTS-2IndexTTS-2

Standard
自然性质 4.3/5
准确性 4.1/5
速度 3.2/5
总体 3.91/5

Spark TTSSpark TTS

Standard
自然性质 4.2/5
准确性 4.0/5
速度 3.4/5
总体 3.9/5

Parler TTSParler TTS

Standard
自然性质 4.1/5
准确性 3.9/5
速度 3.4/5
总体 3.83/5

Tortoise TTSTortoise TTS

Premium
自然性质 4.6/5
准确性 4.4/5
速度 1.8/5
总体 3.7/5

BarkBark

Standard
自然性质 4.2/5
准确性 3.8/5
速度 2.5/5
总体 3.57/5

基准方法

测试设置

  • 硬件 : NVIDIA Tesla P40(各24GBVRAM),96GB总计
  • 测试文本 : 5个标准化段落,涵盖不同的演讲模式(叙述、对话、技术、情感、多语言)
  • 评价: 自动度量(MOS估计、WER、RTF)与人体听觉测试相结合
  • 运行 : 每个模型每通过每次测试10次,平均分数d

标分标准

  • 自然(40%): 欢乐、共鸣、节奏、情感——它听起来如何?
  • 准确性(30%): 发音正确性、字出错率、可知性
  • 速度(30%): 实时因数(音量秒/每一代秒)更高=更快。
  • 总体情况: 加权平均数:0.4x自然+0.3x精确度+0.3x速度

注: 基准反映了我们具体硬件和测试文本的绩效。 现实世界的质量可能因输入文本、语言和语音选择而不同。 社区评级提供了基于不同实际用途的补充信号。

常问问题

TTS竞技场是一个以官方基准测试和社区评级为基础的AI文本对语音模型排名的领头板。 比较模型,边边听样板,并投票支持最适合你的那些模型。

我们使用相同的文本段落、硬件和评估标准对每种模型进行标准化测试。 计分包括自然性( 人的声音 ) 、 准确性( 发音和智能 ) 、 速度( 生成时间 ) 。 所有测试都使用 NVIDIA Tesla P40 GPUs的 GPU服务器。

是的! 单击任何模型旁边的恒星, 从 1 到 5 进行评分 。 您需要签名才能投票 。 您的评分有助于显示在头板上的社区平均值 。 您可以随时改变评分 。

键入任何文本, 选择两个模式, 并单击“ 比较 ” 。 两种模式同时从同一个文本中生成语音。 倾听两者并投票, 听声音会更好。 这种盲目比较有助于识别最适合您具体需要的最佳模式 。

自然度测量了语言声音(假音、内向、节奏)如何像人的声音。精确度测量了发音的正确性和洞察力。速度测量了模型相对于实时的音频生成速度。总体来说,是所有计量的加权平均值。

没有基准分数的模型要么是新添加的,等待测试,要么需要特殊设置(例如门票),这些模型仍然有社区评级。

当模型得到重大更新或新模型被添加时,正式基准就会更新。社区评级随着用户投票而实时更新。头版数据会缓存5分钟用于业绩。

自由级模型(Kokoro、Piper、VITS、MeloTTS)不收取附加费,而抽取免费额度。标准模型使用2x字符(例如,1 000个文本字符需要2,000个字符)。优先级模型使用4x字符,一般提供最优质或独特的功能,如语音克隆。

对于大多数使用的案例,Kokoro(免费级)提供极好的质量。对于语音克隆,请尝试Chatterbox或CosyVoice 2。对于多语言内容,请尝试MelotTS或CosyVoice 2。对于表达式解说、巴克或Dia,请使用比较工具来测试具体文本。

是的, 您可以生成和比较任何两种模型的音频, 没有使用自由级模型的账户就可以生成和比较音频。 对模型的投票需要一个免费账户。 优先级模型的比较需要字符 。

我们努力通过在所有模式中使用标准化测试文本、相同的硬件和一致的评价标准来实现客观性,社区评级提供了另一个独立信号,我们的方法见下文基准方法一节。

模式主要按官方基准总得分排列,然后按社区平均评分划为断线符,没有基准的模型的排名低于有基准的模型,按社区评分顺序排列。
5.0/5 (1)

我们能改进什么?您的反馈帮助我们解决问题。

找到完美的声音

使用 Kokoro、 Piper、 VTS 或 MelotTS 尝试任意模式。 不需要账户 。