ChatPaper.aiChatPaper

VIBE: Stemgeïnduceerde open-einde bias evaluatie voor grote audio-taalmodellen via spraak uit de echte wereld

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

July 3, 2026
Auteurs: Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee
cs.AI

Samenvatting

Grote Audio-Taalmodellen (LALM's) worden steeds vaker geïntegreerd in dagelijkse toepassingen, maar hun generatieve vertekeningen blijven onderbelicht. Bestaande spraakrechtvaardigheidsbenchmarks maken gebruik van synthetische spraak en meerkeuzevragen (MCV's), die beide een gefragmenteerd beeld van rechtvaardigheid bieden. Wij stellen VIBE voor, een raamwerk dat generatieve vertekening evalueert via open taken zoals gepersonaliseerde aanbevelingen, gebruikmakend van door mensen opgenomen spraak. In tegenstelling tot MCV's laat onze methode stereotiepe associaties organisch tot uiting komen zonder vooraf gedefinieerde opties, waardoor deze eenvoudig uitbreidbaar is naar nieuwe taken. Evaluatie van 12 state-of-the-art LALM's onthult systematische vertekeningen in realistische scenario's. Zowel geslachts- als accentaanwijzingen veroorzaken statistisch significante distributieverschuivingen, en de vertekeningsgrootte is sterk taakafhankelijk.
English
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.