EviRank:多模態影像重排序的結構化相關性證據
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
August 21, 2026
作者: Enjun Du, Siyi Liu, Zirong Chen, Xinyu Zuo, Jinwen Luo, Ruiwen Tao, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
cs.AI
摘要
真實世界的影像搜尋查詢具有多模態與組合式特性:「找一件粉紅色的這款襯衫」指定了要保留的實體、要修改的屬性,以及要忽略的上下文。然而,現有的重排序器要嘛將這種多面向的相關性壓縮成不透明的嵌入,要嘛依賴自由形式的思維鏈,而思維鏈容易省略或幻覺化精細的約束。借鑒自然語言處理中基於評分規程與檢查清單的評估,我們將多模態影像重排序重新建構為一個語意約束滿足問題,並提出EviRank。它能將任何查詢——僅文字、僅影像或組合式——解析為統一的證據包:涵蓋六個語意槽位(例如實體、屬性、關係)的類型化準則,每個準則標記為必需、禁止或可忽略。重排序因此簡化為證據條件化驗證,在單一免訓練流程中結合確定性評分規程計分與基於證據的清單式比較。這些明確的證據更可作為結構化監督,可選擇性地蒸餾出輕量級學生模型。在涵蓋文字到影像、影像到影像及組合式影像檢索的五個基準上,EviRank達到了最先進的效能,且蒸餾後的學生模型以大幅更低的成本保留了教師模型超過90%的能力。
English
Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction problem and propose EviRank, which parses any query - text-only, image-only, or composed - into a unified evidence package: typed criteria across six semantic slots (e.g., entities, attributes, relations), each labelled required, forbidden, or ignorable. Re-ranking then reduces to evidence-conditioned verification, combining deterministic rubric scoring and evidence-grounded listwise comparison in a single training-free procedure. The explicit evidence can further serve as structured supervision for optionally distilling a lightweight student. Across five benchmarks spanning text-to-image, image-to-image, and composed image retrieval, EviRank achieves state-of-the-art performance, and the distilled student preserves over 90% of the teacher's capability at substantially lower cost.