Research

New benchmark Real World VoiceEQ measures human-quality aspects of voice AI

Real World VoiceEQ is a new benchmark designed to assess how well voice AI systems capture and act on acoustic cues that transcripts omit—tone, emotion, speaker identity and background context.

New benchmark Real World VoiceEQ measures human-quality aspects of voice AI

Real World VoiceEQ is a benchmark created to evaluate whether voice AI systems can detect, produce and act on the acoustic information that transcripts omit — from tone and emotion to speaker identity and background context.

Why a new measurement was needed

Voice is rapidly becoming a primary interface for AI across customer support, healthcare, education, entertainment and personal assistants. Although technical measures such as word error rate (WER) and latency have improved markedly, real-world conversations reveal shortcomings: voice models may not "hear" or interpret paralinguistic cues the way humans do.

Traditional benchmarks often focus on latency and word accuracy, so they can miss prosodic and contextual signals such as pitch, pauses, hesitation, emphasis and background noise. These cues influence whether a system is perceived as listening, responding appropriately and remaining natural and reliable in live interactions.

Design and dataset of Real World VoiceEQ

  • The benchmark evaluates more than 40 leading proprietary and open-source voice models.
  • It covers 15+ evaluation dimensions and over 60 metrics across Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Speech-to-Speech (S2S) and Speech Understanding.
  • Development relied on more than 1 million individual human ratings collected across diverse demographics, speaking styles and acoustic environments. The current dataset includes 785,000 TTS ratings and 48,000 STS ratings, making it among the largest human evaluations of voice AI to date.
  • All evaluations were conducted on Kairos, a flexible, voice-native evaluation platform. The same infrastructure enables AI labs and enterprises to run custom evaluations, identify granular failure modes in production voice systems, gather human preference data, and iterate models with reinforcement learning and human feedback.

Key findings

  • Progress is growing more specialized. Rather than a single "best" model, systems now optimize different strengths: technical accuracy, emotional understanding, conversational intelligence, expressiveness and robustness. A model that reliably repeats booking reference numbers or pharmaceutical names may underperform on emotionally expressive speech, while a very natural-sounding system may be less reliable on precision tasks.

  • Models have become better at speaking than actually listening. Speech-to-Speech models showed the widest variation. Some systems recognized emotion well but did not respond naturally. Access to audio did not guarantee that agents used paralinguistic information — many remained largely transcript-driven and overlooked cues such as tone, pacing, hesitation, emphasis and volume.

  • Traditional benchmarks can overestimate real-world performance. Models still struggle with accented speech, overlapping speakers, emotional speech, background noise and longer conversations. Performance varied much more across leading open-source and proprietary models than conventional benchmarks suggest. For example, transcription WER on noise-backed speech was roughly four times higher than on music-backed speech, indicating a single background-audio score can hide important failure modes.

  • Human evaluation remains essential. Preliminary research found signs that some models may have been optimized for public benchmarks: reproducing known errors in reference transcripts, following arbitrary spelling conventions, or reconstructing masked words not present in the audio.

  • Use of speech-language models (SLMs) for evaluation requires caution. While LLMs are widely used for text-based model evaluation, the Real World VoiceEQ team found stronger agreement between SLMs and trained human raters only on clear, verifiable tasks such as pronunciation accuracy. Agreement dropped for subjective judgments — for example whether a voice fit an acting role or maintained a consistent identity — and SLMs sometimes inferred emotion from text-context rather than acoustic cues. Automated evaluators can be useful for well-defined tasks, but they are not yet a substitute for human listeners when judgments depend on acoustic context, perception and social interpretation.

Why a new measurement layer matters

As voice becomes one of AI's defining interfaces, speed and technical accuracy alone will not determine which systems succeed. Users will prefer systems that truly understand, express and respond like humans — not only in ideal benchmark conditions but across the diversity of real conversations.

For decades, speech AI advanced by optimizing against quantitative metrics on standardized benchmarks (from WER for transcription to objective perceptual metrics like PESQ and DNSMOS for speech quality). Real World VoiceEQ aims to extend that paradigm with a human-grounded metric that evaluates the components of synthetic voice interaction.

What’s next

The full technical report and public leaderboards are available, and Hume offers to evaluate voice models or agents using Real World VoiceEQ or to design custom evaluations for specific use cases. The benchmark and Kairos infrastructure support detailed failure analysis, collection of human preference data, and continual model improvement informed by real user feedback.

The results from Real World VoiceEQ indicate that advancing voice AI further will require measurement layers focused on human perception and acoustic-context, particularly for phenomena that traditional text-based or purely technical metrics do not capture.