The release pace of open‑source text‑to‑speech (TTS) models has been extraordinary: on the Hugging Face Hub there were over 8,000 TTS models available as of September 30, 2026. Evaluation practices, however, have not scaled at the same rate and remain fragmented and nonstandardized.
Human preference measures (e.g., MOS or MUSHRA) remain the gold standard. Arena‑style leaderboards such as TTS Arena v2, Artificial Analysis and Voice Arena have become community reference points: they present outputs from two models to users, collect pairwise votes and compute Elo‑like rankings (commonly using the Bradley–Terry model).
Arena evaluations have limits: they do not scale easily with the flood of new TTS releases. Practical factors contribute to an underrepresentation of open‑weights models on these lists — commercial API models can be added with an API key, while open models must be hosted and served by the arena operator. As of September 30, 2026, only 16 of the 92 models listed on Artificial Analysis were open‑weights, and Voice Arena shows a similar skew. Another challenge is voter consistency: an arena cannot ensure the same voters apply the same criteria over time.
To address these limitations, the Open TTS Leaderboard was developed to evaluate models using objective metrics that capture complementary aspects of performance:
- Intelligibility: WER and CER computed between the prompt and the ASR transcript of the generated audio using Qwen3 ASR (the top open‑source ASR on the Open ASR Leaderboard).
- Speed: inverse real‑time factor (RTFx) for batched offline inference on an H200 GPU, and time‑to‑first‑audio (TTFA) to measure streaming/batch=1 latency on H200 GPU and CPU.
- Speaker similarity: cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and a reference clip.
Relying on objective metrics shortens evaluation time from weeks (to collect human votes) to hours using automated runs.
Important caveat: the Open TTS Leaderboard does not replace human preference rankings. ASR‑based WER is a proxy for intelligibility and SIM estimates voice identity preservation, but neither directly measures naturalness, expressiveness, or listener preference. Still, these metrics can inform which models arena‑style leaderboards should include for human evaluations.
Multilingual and voice‑cloning evaluation
By default, models are ranked by macro‑average WER on the English splits of Seed TTS Eval (paper) and CV3 Eval (zero shot) (paper). On these English splits, hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead in average WER. Pareto plots visualize trade‑offs between WER, batched inference (RTFx) and model size.
English performance does not necessarily generalize to other languages. The leaderboard allows toggling multiple languages for multilingual rankings. Since Seed TTS Eval provides audio only for English and Chinese, scores for other languages rely on CV3 Eval (zero shot). Chinese, Japanese and Korean are character‑based languages here, so CER is reported; the “Average WER” across languages is a macro‑average over the selected languages.
Strong multilingual models include k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512.
Enabling the “Voice cloning” option filters to models that support reference‑based synthesis in the selected languages. A SIM column for speaker similarity is shown in the table and additional Pareto plots illustrate trade‑offs between SIM, batched inference and model size. Some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, show improved average WER when voice cloning (i.e., when a reference audio is provided).
Compare outputs and vote
Numbers only tell part of the story and human preference remains the ultimate decider. The leaderboard’s “Listen” tab exposes the generated outputs behind the metrics so users can compare actual audio. Users can select language/dataset, choose whether to compare voice cloning, and pick specific models or sample random outputs.
The “Listen” tab fills a gap in existing TTS leaderboards by providing a space to explore model outputs. Users can also give feedback on generated clips; if the project collects enough community votes, these may be incorporated into the leaderboard. The team asks voters to log in with their Hugging Face account to reduce spam and bot activity.
Streaming performance
The “Streaming” tab compares streaming capabilities: models are ranked by TTFA (time‑to‑first‑audio), the delay from probing a model to receiving the first playable audio chunk — a critical metric for voice agents and interactive applications.
For streaming‑enabled models (marked with ✅ under “Streaming API”), TTFA is the time until the first audio chunk arrives. For non‑streaming models, TTFA equals the time until the complete utterance is generated because playback cannot start earlier. Each model is run one audio at a time (batch size 1) on the same 50 English CV3‑Eval prompts, on identical hardware and using its default voice. The first three runs are dropped as warm‑up and the median TTFA across the remainder is reported.
The default comparisons show H200 GPU results, and CPU results are available for a small (but growing) set of models. The kyutai/pocket-tts model stands out for strong streaming performance on both GPU and CPU.
Conclusion and community participation
The Open TTS Leaderboard aims to keep pace with the rapid stream of TTS model releases while being shaped by the community: the team requests feedback to keep evaluations relevant and informative. Current emphases are on open‑source models, which have been neglected by arena‑style evaluations, and on multilingual performance, since English is not a reliable proxy for other languages.
The project plans to open‑source its evaluation scripts soon, similar to the Open ASR Leaderboard repository, enabling direct feedback and contributions via GitHub Issues and PRs. Community input will guide which datasets, models and metrics should be added next.



