Skip to content

Humyn Labs Benchmarks Voice AI Models: Reveals The Gaps and Failures In Human-Robot Communication

Humyn Labs Benchmarks Voice AI Models: Reveals The Gaps and Failures In Human-Robot Communication
Humyn Labs voice AI benchmark revealing gaps in human-robot communication

SUMMARY

BRIDGE evaluates how leading AI models like Sarvam v3, Gemini 3 Pro, ElevenLabs among others perform across the world languages for humans to interact with robots effectively

Bengaluru, 17 September 2026: Humyn Labs, a Physical AI research lab dedicated to accelerating the deployment-readyness of robots in the real world, today launched the second edition of BRIDGE, a benchmark report that reveals the gap between human speech and AI voice models. This gap must be addressed for humans and robots to effectively interact and perform tasks in the real world. The report considered 23 voice AI models across 23 languages in real-world noisy conversations to evaluate how leading voice models perform across the world’s languages, accents and dialects. The benchmark maps and compares models like Sarvam v3, Gemini 3 Pro, ElevenLabs among others. 

BRIDGE, the global independent ASR benchmark maps voice models using seven core metrics, including overlapping speech, conversational density, code-switching and other conditions that mirror real-world environments. This edition of the benchmark report covers Indic languages alongside Latin American Spanish, Brazilian Portuguese and Vietnamese, and is built on more than 200 hours of human-verified real world audio collected across two to three districts per language. With 5.5 billion people speaking languages other than English, the ability to understand speech across linguistic and conversational contexts is becoming a critical test of how inclusive and reliable AI can be. 

The findings reveal that overlapping speech alone pushed the average error rate up from 41.2% to 45.2%. Dialect made things worse: Bengali scored 42.4% in its standard form but 51.0% in a regional dialect outside Kolkata, a gap of nearly nine points caused by dialect alone. The dialect gap extends well beyond Indic languages, with Spanish showing the same pattern: Argentinian Spanish recorded a 7.85% word error rate against 16.04% for Venezuelan Spanish, more than doubling across dialects of the same language and confirming the effect first seen in Bengali is not Indic-specific. 

See also  Top 10 Private Equity Firms in USA

Model choice remains equally decisive, with the top performer across five non-Indic languages, ElevenLabs, averaging 5.8% error compared to 24.6% for the widely used GPT-4o-mini-transcribe, a gap of more than 4x on identical audio that makes provider selection one of the largest levers on real-world accuracy. The silence problem also replicated internationally, with Brazilian Portuguese calls showing an 18.8% error rate for gaps exceeding 150 seconds versus 12.4% for shorter gaps of around 35 seconds, confirming that models losing the thread across long pauses is not limited to the Indic languages.

With an aim to address the gap, Manish Agarwal, Co-Founder, Humyn Labs, stated, “Voice is a critical interface for Physical AI, and therefore voice accuracy becomes a business imperative, not just a technical metric. If Voice AI models cannot understand overlapping speech, interruptions, code-switching, pauses and the diversity of languages people use every day, that gap ultimately impacts customer experience, automation and trust. BRIDGE is designed to help Physical AI & Voice AI builders and enterprises evaluate models against the complexity of real-world conversations and understand whether they are truly ready to scale across markets.”

BRIDGE separates script choice from genuine transcription error; a distinction standard scoring misses. In Bengali, Soniox and Sarvam v3 recorded raw error rates of around 20% to 21%, roughly half of which came from writing English loanwords in a different script rather than from mishearing them. Gemini 3 Pro recorded 9.2%, of which only 1.6 points were script mismatch.

Models were found to fail in structurally different ways at similar overall scores. 19 out of the 23 models tested, mishear a word and substitute the wrong one. A smaller group fails by omission instead: OpenAI’s transcribe models, Speechmatics, and Gnani Vachana disproportionately drop words rather than mistranscribe them, with deletions accounting for 38–39% of their total errors. Gemini Flash fails a third way, by fabrication: it introduces words that were never spoken, adding invented content equal to 9.3% of the reference transcript’s length. The report highlighted that the distinction mattered commercially, since a workflow that can absorb a missing word may not tolerate a fabricated one.

See also  Iztri secured ₹10 crore in a seed funding round led by All In Capital and Suashish Group

BRIDGE also tested the economics of model routing. The best single model reached 10.7% loanword-adjusted error, the best model per language reached 9.7%, and a theoretical best model per call reached 8.9%. One model won 78.6% of files outright.

Ishank Gupta, Co-Founder, Humyn Labs, adds, “Physical AI cannot learn the real world through vision alone. Sound carries information about people, actions, distance, environment and intent and for a robot operating alongside humans, being able to interpret that signal reliably is fundamental. BRIDGE provides the evaluation layer that has been missing for this modality: testing speech models not just on words, but across conditions and context that mirror real-world environments.”

The full BRIDGE report and dataset are available here.

About Humyn Labs

Humyn Labs is a physical AI research engine. The company is building data enrichment pipelines from collecting human actions “in the wild” to building training ready multi modality enriched signals for machines to read thus acting as an accelerator to the deployment-readyness of the robots in the real world. It pioneers source-first data collection, labelling, and data enrichment across voice, vision, motion and touch to help machines to understand the complexity of the real world and evolve into a functioning brain for the robots.

Working across India, Southeast Asia, LATAM and the Middle East, Humyn Labs brings together human intelligence, technology and deep research to create high-quality, diverse and contextual data for frontier AI and Physical AI. Its mission is to shape the intelligence layer that enables AI to work more accurately, inclusively and intelligently in the real world.  

About BRIDGE

BRIDGE is a global independent ASR benchmark that evaluates commercial and open-source models on field-collected, human-verified conversational audio using a seven-metric scoring stack. It maps the gap between human speech and voice models including overlapping speech, conversational density, code-switching and other conditions that mirror real-world environments. The benchmark report covers Indic languages alongside Latin American Spanish, Brazilian Portuguese and Vietnamese, and is built on more than 200 hours of human-verified real world audio collected across two to three districts per language. 

See also  Yuma Energy secured $35 million in a Series A funding round led by Magna International Inc.

For more information:

Saheli Chatterjee | +91 9163323848 | Saheli.chatterjee@4wdtechpr.com

Note: We at scoopearth take our ethics very seriously. More information about it can be found here.