ChatPaper.aiChatPaper

CORE: 리랭커 증류를 통한 MLLM 임베딩의 구성적 추론 향상

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

September 3, 2026
저자: Tingyu Song, Mingxin Li, Yanzhao Zhang, Dingkun Long, Chu Liu, Pengjun Xie, Yilun Zhao, Shu Wu
cs.AI

초록

MLLM 기반 임베딩 모델은 구성적 검색에서 여전히 제한적이며, 동일한 개념을 포함하되 서로 다른 속성-객체 결합을 지닌 장면들을 구분하지 못하는 경우가 많다. 그러나 동일한 백본(backbone)은 교차 주의 재순위화기(cross-attentive reranker)로 사용될 때 그러한 구분을 해결할 수 있으며, 이는 임베딩 모델로 구성적 판단을 증류하도록 하는 동기를 제공한다. 본 연구에서는 CORE를 제안한다. CORE는 다섯 가지 구성적 정합 수준을 포괄하는 후보 목록을 합성하고, 재순위화기의 세밀한 순위를 임베딩 모델이 재현하도록 훈련시키는 Rank-KL 목적 함수를 도입한다. 또한 등급화된 평가 프로토콜을 제안하고, 동일한 데이터 및 튜닝 예산 하에서 대조 학습, 쌍별 CoSENT, 목록별 Rank-KL을 비교한다. 비교 결과, CoSENT와 Rank-KL은 모두 대조 학습보다 다중 수준 지도 정보를 더 효과적으로 활용하며, Rank-KL이 전반적으로 가장 우수한 성능을 달성한다. 세 가지 구성적 추론 벤치마크(COLA, SUGARCREPE++, NEGBENCH)에서 CORE-RERANKER-8B는 총평균 82.7%를 기록하여 Jina-Reranker를 10.7포인트 능가하며, CORE-EMBED-8B는 평가된 모든 임베딩 모델 중 최고 총평균(0.666)을 달성한다. 이러한 개선은 COCO 및 Flickr30K에서의 검색 성능 저하 없이 MCMR 벤치마크로 전이된다.
English
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.