ChatPaper.aiChatPaper

관련성 중심 검색을 넘어서: 루브릭 중심 문서 집합 선택 및 순위화

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

July 22, 2026
저자: Kailin Jiang, Lei Liu, Jian Xi, Hui Xu, Junlin Liu, Baochen Fu, Shaoqing Ren, Bin Li, Vichwang, Yu Lu, Haibo Shi
cs.AI

초록

대규모 언어 모델과 AI 에이전트가 검색 결과의 주요 소비자가 되면서, 문서 집합의 품질이 하위 생성(다운스트림 생성)의 상한선을 결정한다. 그러나 기존 평가 시스템은 문서를 독립적으로 평가하고 nDCG를 통해 집계하는 데 그쳐, 문서 간 상호작용(중복, 충돌, 보완성)을 무시하며, 어떤 문서 집합이 다른 집합보다 우수한지에 대한 답을 제시하지 못한다. 이러한 문제를 해결하기 위해 우리는 완전한 평가-진단-최적화 프레임워크를 제안한다. 단문 및 장문 시나리오를 모두 포괄하는 세 수준, 아홉 차원의 문서 집합 평가 벤치마크인 SetwiseEvalKit을 설계하였으며, 약 28,000개의 고품질 평가 루브릭을 포함한다. 우리는 12개의 재순위화기를 체계적으로 평가한 결과, 최고 성능의 방법조차 50% 이하의 커버리지를 보였고, 교차 문서 조정 차원에서 전반적으로 취약했으며, 어떤 단일 방법도 두 설정에서 모두 최고 성능을 유지하지 못했다. 이를 바탕으로, 루브릭 기반 평가 기준을 문서 집합 선택 신호로 변환하는 훈련 불필요 방식인 Rubric4Setwise를 제안한다. 이 방법은 더 적은 문서와 검색 라운드로 최상의 하위 생성 성능을 달성하며, 두 시나리오 모두에서 최첨단 결과를 유지하는 유일한 방법으로, 평가에서 최적화로 이어지는 폐쇄 루프의 효과성을 입증한다.
English
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.