ChatPaper.aiChatPaper

증거-분류체계 검색을 위한 분해 가설 탐색

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

August 6, 2026
저자: Linhai Ma, Ethan F. Wei, Xueqing Peng, Yan Wang, Lingfei Qian, Víctor Gutiérrez-Basulto
cs.AI

초록

대규모 분류 체계 검색은 종종 입력이 이미 대상 개념을 표현한다고 가정한다. 그러나 많은 상황에서 입력은 간접 증거에 불과하다. 예를 들어 표 셀의 의미는 해당 셀이 속한 행, 열, 데이터 유형, 그리고 맥락에 따라 달라진다. 우리는 이러한 불일치를 검색 준비도 격차(retrieval readiness gap)라고 부른다. 우리의 분석은 현재 인덱스가 의미가 명시적으로 드러날 때는 대상을 안정적으로 검색하지만, 원시 증거가 주어질 때는 대상이 순위 깊숙이 머무는 경우가 많다는 것을 보여준다. 우리는 명명된 의미 차원들에 걸쳐 여러 부분 해석을 유지하는 분해 가설 검색(Factorized Hypothesis Search, FHS)을 제안한다. 이러한 가설들은 구조화된 질의 렌더링, 다중 가설 검색, 차원 수준 후보 검증을 지원한다. 금융 분류 체계 태깅 및 CodiEsp 임상 코딩 작업 모두에서 FHS는 비-오라클 방법들 중 최고의 Recall@1, MRR, 그리고 최종 정확도를 달성한다. 분해 가설 경로를 자유 텍스트 앙상블로 대체하면 상위 순위 성능이 가장 크게 하락하며, 순차적 정제는 FHS의 강력한 병렬 첫 라운드에 비해 추가적인 이득을 제공하지 않는다.
English
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.