EviRank: 멀티모달 이미지 재순위화를 위한 구조화된 관련성 증거
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
August 21, 2026
저자: Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
cs.AI
초록
실제 이미지 검색 질의는 다중 양식(multimodal)이며 구성적(compositional)이다. "이 셔츠를 분홍색으로 찾아라"는 유지할 개체, 수정할 속성, 무시할 맥락을 명시한다. 그러나 기존 재순위화(re-ranking) 모델은 이러한 다면적 관련성을 불투명한 임베딩으로 압축하거나, 미세한 제약 조건을 쉽게 누락하거나 환각(hallucination)하는 자유 형식의 체인오브소트(chain-of-thought)에 의존한다. 우리는 NLP의 루브릭(rubric) 및 체크리스트 기반 평가에서 착안하여 다중 양식 이미지 재순위화를 의미적 제약 충족 문제로 재정의하고, 텍스트 전용, 이미지 전용, 또는 복합 질의를 통합된 증거 패키지로 파싱하는 EviRank를 제안한다. 증거 패키지는 6개의 의미적 슬롯(예: 개체, 속성, 관계)에 걸친 유형화된 기준으로 구성되며, 각 기준은 필수(required), 금지(forbidden), 무시 가능(ignorable)으로 라벨링된다. 재순위화는 이후 결정적 루브릭 점수 계산과 증거 기반 리스트 단위 비교를 단일 학습 불필요 절차에서 결합한 증거 조건부 검증(evidence-conditioned verification)으로 축소된다. 명시적 증거는 선택적으로 경량 학생 모델을 증류하기 위한 구조적 감독으로도 활용될 수 있다. 텍스트-이미지, 이미지-이미지, 복합 이미지 검색을 포괄하는 5개 벤치마크에서 EviRank는 최고 수준의 성능을 달성하며, 증류된 학생 모델은 교사 모델 성능의 90% 이상을 유지하면서 훨씬 낮은 비용을 보인다.
English
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.