Open TTS Leaderboard: A Scalable Benchmark For Speech And Voice Cloning
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Open TTS Leaderboard: A Scalable Benchmark For Speech And Voice Cloning on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has introduced the Open TTS Leaderboard, which compares text-to-speech models using speech-recognition error rates, generation speed and speaker similarity. The project says automated evaluations can run in a few hours, but the scores do not measure naturalness or replace listener preferences.

Hugging Face has launched the Open TTS Leaderboard, a tool for comparing text-to-speech models on speech accuracy, generation speed and speaker similarity. The project is intended to make evaluations faster and more repeatable as the Hugging Face Hub held more than 8,000 TTS models as of September 30, 2026, according to the company. Its scores do not assess every quality listeners care about: Hugging Face says the leaderboard does not replace human preference judgments.

The leaderboard estimates speech accuracy by comparing transcripts generated from model audio with the original text prompts. It uses Qwen3 automatic speech recognition to calculate word error rate and character error rate. For generation speed, it reports offline speed using inverse real-time factor on an H200 GPU, and streaming responsiveness using time-to-first-audio on both an H200 GPU and a CPU.

A separate voice-cloning measure estimates how closely generated speech matches reference audio. The system compares WavLM embeddings from the output and reference recordings to produce a speaker-similarity score. The default ranking uses macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval; users can select other languages and switch to a voice-cloning view.

Hugging Face identifies k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leaders. These are results under the leaderboard’s chosen measures, not an overall judgment that those models sound best. A separate Listen tab lets users compare audio and submit preferences.

At a glance
announcementWhen: Launched in September 2026; Hugging Fac…
The developmentHugging Face launched an open leaderboard for comparing text-to-speech and voice-cloning models with automated metrics and a separate listening feature.
At a glance
announcementWhen: Announced in material dated September 3…
The developmentHugging Face has launched an open leaderboard that evaluates text-to-speech models across multiple languages using objective performance metrics and offers audio comparisons for community feedback.

Faster Comparisons for Open TTS

Text-to-speech developers and users often have to weigh several different needs: accurate pronunciation, low delay, voice similarity and audio that sounds natural. The leaderboard places some of those measurable tradeoffs in one view. Its time-to-first-audio scores, for example, may help teams evaluating voice-agent systems where a slow start can affect how responsive a conversation feels. Accuracy and cloning scores can also help identify candidates for more detailed listening tests.

The project’s speed claim matters for teams tracking a quickly changing model field. Hugging Face says objective evaluations can be completed in a couple of hours, compared with weeks for arena voting. That could make it easier to test new or less widely used models, though the source material does not provide an independent comparison of evaluation times or results. The scores are best treated as complementary evidence: they can narrow a field, but cannot establish which voice listeners will prefer.

Hugging Face says open-weight systems are underrepresented in some existing rankings. It counted 16 open-weight models among 92 on Artificial Analysis as of September 30, 2026, and said Voice Arena showed a similar skew. The company attributes the imbalance in part to the work needed to host open models and to commercial providers’ stronger incentives to seek placement. The count is Hugging Face’s characterization; the leaderboard may offer another route to compare open models without relying only on hosted, vote-based arenas.

Amazon

AI voice cloning software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Arena Votes to Metrics

Existing TTS comparisons often use arena-style evaluations. Listeners hear two outputs, choose a preferred one and contribute votes used to rank models, sometimes with an Elo score calculated using a Bradley–Terry model. Hugging Face points to TTS Arena v2, Artificial Analysis and Voice Arena as examples of such comparison systems. Listener votes capture subjective reactions directly, but collecting enough of them takes time, and operating an arena can require hosting the models being tested.

The new leaderboard takes a different approach for much of its evaluation: it uses standardized datasets and automated metrics to produce comparisons more quickly. Its listening feature retains a role for direct human judgment. Users can select a language and dataset, choose whether to compare voice cloning, listen to outputs and vote. Hugging Face asks users to sign in to reduce spam and bot submissions.

The automated measures each cover a limited part of performance. Recognition error rates estimate how well an ASR system can recover the prompted words from generated audio; speaker similarity estimates preservation of voice identity. Neither directly measures naturalness, expressiveness or listener preference. Hugging Face describes the leaderboard as a way to complement, not replace, human preference rankings.

“The Open TTS Leaderboard does not replace human preference ranking.”

— Hugging Face

Amazon

text-to-speech model hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Published Rankings

The source announcement does not provide the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not show how closely the automated metrics match listener judgments across languages, accents and speaking styles. Those gaps make it difficult to judge how stable small differences between model scores are or how well they predict real-world preference.

Coverage varies by language. The leaderboard uses character error rate for Chinese, Japanese and Korean. For languages beyond English and Chinese, Hugging Face says Seed TTS Eval has no audio and scores come from CV3 Eval alone. The announcement does not explain the effect of that difference on cross-language comparisons. Nor does it specify how often the rankings will be refreshed or how changes to model versions and evaluation data will be tracked.

Hugging Face says community votes may be incorporated as feedback accumulates, but has not stated a schedule or threshold for doing so. It is also unclear how those votes would be combined with objective scores. Until those details are published, readers should interpret the rankings as metric-specific snapshots, not definitive ordering of the best-sounding TTS systems.

Amazon

speech recognition device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Listening Tests and Ranking Updates

Users can try the leaderboard’s Listen tab, select supported languages and datasets, and submit preferences after signing in with a Hugging Face account. The project says it may use accumulated votes in future leaderboard results, but has not announced when or how voting will affect rankings. No refresh schedule or next evaluation milestone is specified in the source material.

Further updates would be useful if they publish the model roster, evaluation sample counts, score uncertainty and the process for handling new model versions. Those details could help users distinguish meaningful performance differences from results tied to a particular dataset or evaluation setup. For now, the confirmed next step is continued listening and community feedback; changes to the leaderboard and the timing of any vote-based ranking remain unannounced.

Amazon

voice synthesis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the Open TTS Leaderboard measure?

It reports speech-recognition word and character error rates, offline generation speed, streaming time-to-first-audio and a speaker-similarity score for voice cloning. These measures cover selected aspects of model performance, not overall listener preference.

Does a high leaderboard position mean a model sounds the most natural?

No. The listed metrics do not directly measure naturalness, expressiveness or human preference. Hugging Face provides a separate Listen tab so users can compare audio and submit preferences.

Which models does Hugging Face highlight?

For multilingual performance, Hugging Face names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512. For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro. These are leaderboard results under particular metrics, not a general verdict on sound quality.

How quickly can the evaluations run?

Hugging Face says its objective-metric evaluations can take a couple of hours, compared with weeks for arena voting. The supplied material does not give an independent timing study or enough detail to compare evaluation conditions.

Will community votes affect the rankings?

Hugging Face says votes may be incorporated as more feedback arrives, but it has not announced a schedule, threshold or method for combining votes with automated scores.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Minerva. The opposite path.

Italy’s Minerva LLM, trained from scratch on 2.5 trillion tokens, shows impressive architecture but underperforms on Italian benchmarks, raising questions about native-language investment.

Decoding ByteDance’s AI Reshuffle: A Sign Of Industry Confidence

ByteDance is restructuring its AI division, with co-founder Zhang Yiming emphasizing long-term development over shortcuts, signaling industry confidence.

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers present a framework outlining pathways from human-level AI to superintelligence, emphasizing scaling, paradigm shifts, and multi-agent systems.

Getting 50 GB/S Back From The Apple Neural Engine

Apple’s M3 Neural Engine reportedly reaches 50 GB/s throughput after addressing a performance issue, boosting AI processing capabilities significantly.