ContextBias: 텍스트-이미지 모델에서 맥락 변화에 따른 편향 지속성의 통제된 평가
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
August 30, 2026
저자: Shaghayegh Kolli, Sina Emami, Moreno D'Incà, Pouyan Nejadi, Nicu Sebe, Massimiliano Mancini, Jana Diesner
cs.AI
초록
텍스트-이미지 모델은 개념들, 즉 이 논문에서 역할이라고 부르는 사람들의 직업과 시각적 속성 사이의 연관성을 학습한다. 이러한 연관성은 관찰되는 많은 형태의 고정관념적 편향을 뒷받침할 수 있다. 이 분야의 핵심 미해결 질문은 이러한 연관성이 안정적인지, 아니면 직업적 역할을 가진 사람들의 시각적 표상이 서로 다른 프롬프트 맥락에 놓일 때 변화하는지 여부이다. 우리는 통제된 평가 프레임워크인 ContextBias와 92개 역할과 1,656개의 의미론적으로 통제된 프롬프트를 포괄하는 벤치마크인 ContextBench를 도입하여 역할과 연계된 시각적 표상에 대한 맥락 변화의 영향을 분리한다. 66,240개의 생성 이미지에서 4개의 최첨단 모델을 평가한 결과, 역할을 의미론적으로 무관한 맥락에 배치하더라도 역할 연계 속성이 억제되지 않으며, 오히려 역할 간 속성 집중도가 증가함을 발견했다(통합 BI +0.047). 인구통계학적 단서, 특징적인 의복, 역할 고유 도구는 맥락이 없는 조건, 관련 조건, 무관한 조건 모두에서 높은 빈도로 유지되며, 의미론적 프롬프트 재구성에도 견고하다. 장면 구성과 카메라 구도는 가장 큰 맥락 민감성을 보인다. 이러한 발견은 맥락이 없는 평가에서는 대체로 드러나지 않는 고정관념적 지속성의 한 형태를 밝혀내며, 편향 벤치마킹에서 통제된 맥락 변화의 필요성을 강조한다. 코드 및 데이터셋: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias
English
Text-to-image models learn associations between concepts - in the case of this paper, people's professions, which we refer to as roles - and visual attributes. These associations can underpin many observed forms of stereotypical bias. A key open question in this area is whether these associations are stable or change when visual representations of people in professional roles are placed in different prompted contexts. We introduce ContextBias, a controlled evaluation framework, and ContextBench, a benchmark spanning 92 roles and 1,656 semantically controlled prompts, designed to isolate the effect of contextual variation on role-linked visual representations. Evaluating four state-of-the-art models on 66,240 generated images, we find that placing a role in a semantically unrelated context does not suppress role-linked attributes; instead, cross-role attribute concentration increases (pooled BI +0.047). Demographic cues, characteristic garments, and role-specific tools remain highly prevalent across context-free, related, and unrelated conditions, and are robust to semantic prompt reformulation. Scene composition and camera framing show the greatest context-sensitivity. These findings reveal a form of stereotypical persistence that remains largely invisible to context-free evaluations, highlighting the need for controlled contextual variation in bias benchmarking. Code and dataset: https://huggingface.co/datasets/shaghayegh/ContextBias , https://github.com/Sina-Emami/ContextBias