ChatPaper.aiChatPaper

VIBE: 実世界音声による大規模音声言語モデルの音声誘発型開放バイアス評価

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

July 3, 2026
著者: Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee
cs.AI

要旨

大規模音声言語モデル(LALMs)は日常的なアプリケーションにますます統合されつつあるが、その生成バイアスは未だ十分に解明されていない。既存の音声公平性ベンチマークは合成音声と多肢選択式質問(MCQs)に依存しており、いずれも公平性の断片的な評価しか提供しない。本稿では、人間が録音した音声を用いて、パーソナライズドレコメンデーションなどの自由回答形式タスクを通じて生成バイアスを評価するフレームワークVIBEを提案する。MCQsとは異なり、本手法では事前定義された選択肢なしにステレオタイプ的な関連付けが自然に現れるため、新しいタスクへの拡張が容易である。12の最先端LALMsを評価した結果、現実的なシナリオにおいて系統的なバイアスが明らかになった。性別とアクセントの手がかりは共に統計的に有意な分布シフトを引き起こし、バイアスの大きさはタスクに強く依存することが示された。
English
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.