ChatPaper.aiChatPaper

시각적 유사성을 넘어: 지식 기반 시각 질의응답을 위한 엔티티 정렬 검색

Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering

August 19, 2026
저자: Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen, Rui Cong, Xiangwen Deng, Zhifang Liu, Peng Jiao, Haoqian Wang
cs.AI

초록

지식 기반 시각 질의 응답(KB-VQA)은 롱테일 엔티티와 관련된 질의에 답하기 위해 외부 정보 검색에 의존한다. 그러나 기존 검색 파이프라인은 대부분 CLIP 방식의 이중 인코더를 사용하며, 이는 엔티티 수준의 의미 정렬보다 표면적인 시각적 유사성을 우선시한다. 이러한 패러다임은 의미적으로 동일한 개념이 큰 시각적 차이를 보이거나, 서로 다른 엔티티가 시각적으로 유사하게 나타나는 경우 종종 실패한다. 이를 해결하기 위해, 본 논문은 KB-VQA에 특화된 최초의 MLLM 기반 임베딩 검색기인 KBMR을 제안한다. KBMR은 MLLM의 강력한 자기회귀 능력을 활용하여 이미지를 개념 정체성을 더 잘 보존하는 의미 공간으로 매핑한다. 또한 위키피디아 규모 검색에서의 노이즈가 있는 지도(supervision) 문제를 해결하기 위해, 연속적인 엔티티 일관성 가중치를 생성하는 MLLM 기반 의미 판별기를 도입한다. 이러한 가중치는 새로운 연속 의미 증류 목적 함수를 안내함으로써, 경직된 이진 레이블을 넘어 효과적인 하드 네거티브 샘플링과 소프트 지도를 가능하게 한다. 광범위한 실험을 통해 KBMR은 CLIP 베이스라인을 크게 능가하며, 검색 Recall@1에서 최대 14.7%의 향상과 종단 간 VQA 정확도에서 9.4%의 개선을 달성한다. 코드는 https://github.com/realHarryX/KBMR에서 제공된다.
English
Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, existing retrieval pipelines predominantly employ CLIP-style dual encoders, which prioritize surface-level visual similarity over entity-level semantic alignment. This paradigm often fails when semantically identical concepts exhibit large visual variations or when distinct entities appear visually similar. To address this, we propose KBMR, the first MLLM-based embedding retriever tailored for KB-VQA. Leveraging the robust autoregressive capabilities of MLLMs, KBMR maps images into a semantic space that better preserves concept identity. To tackle the challenge of noisy supervision in Wikipedia-scale retrieval, we introduce an MLLM-based semantic discriminator that generates continuous entity-consistency weights. These weights guide a novel continuous semantic distillation objective, enabling effective hard negative sampling and soft supervision beyond rigid binary labels. Extensive experiments demonstrate that KBMR significantly outperforms CLIP baselines, yielding up to a 14.7% improvement in retrieval Recall@1 and a 9.4% gain in end-to-end VQA accuracy. Code is available at https://github.com/realHarryX/KBMR.