VIBE: 실제 음성을 통한 대규모 오디오-언어 모델의 음성 유도 개방형 편향 평가
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
July 3, 2026
저자: Yi-Cheng Lin, Yusuke Hirota, Sung-Feng Huang, Hung-yi Lee
cs.AI
초록
대규모 오디오-언어 모델(LALM)은 일상적인 응용 프로그램에 점점 더 통합되고 있지만, 이들의 생성적 편향은 아직 충분히 탐구되지 않았다. 기존의 음성 공정성 벤치마크는 합성 음성과 객관식 질문(MCQ)에 의존하며, 두 방식 모두 공정성에 대한 단편적인 시각을 제공한다. 본 논문에서는 인간이 녹음한 음성을 사용하여 개인화 추천과 같은 개방형 과제를 통해 생성적 편향을 평가하는 프레임워크인 VIBE를 제안한다. MCQ와 달리, 본 방법은 고정관념적 연관성이 미리 정의된 선택지 없이 자연스럽게 드러나도록 하여 새로운 과제로 쉽게 확장 가능하다. 12개의 최첨단 LALM을 평가한 결과, 현실적인 시나리오에서 체계적인 편향이 발견되었다. 성별 및 억양 신호 모두 통계적으로 유의미한 분포 변화를 유발하며, 편향의 정도는 과제에 크게 의존적이다.
English
Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.