ChatPaper.aiChatPaper

VIBE:通過真實世界語音對大型音頻-語言模型進行語音引發的開放式偏見評估

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

July 3, 2026
作者: Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee
cs.AI

摘要

大型音頻語言模型(LALMs)日益融入日常應用,但其生成性偏誤仍未被充分探討。現有的語音公平性基準依賴合成語音與選擇題(MCQs),兩者皆僅提供片面的公平性觀點。我們提出 VIBE 框架,透過開放式任務(如個人化推薦)並使用真人錄製語音評估生成偏誤。與選擇題不同,我們的方法讓刻板聯想能在無預設選項的情況下自然浮現,且易於擴展至新任務。評估 12 個最先進的大型音頻語言模型後,結果顯示在真實情境中存在系統性偏誤:性別與口音線索皆引發統計上顯著的分布偏移,而偏誤幅度與任務類型高度相關。
English
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.