ChatPaper.aiChatPaper

分解式假設搜尋用於證據至分類樹檢索

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

August 6, 2026
作者: Linhai Ma, Ethan F. Wei, Xueqing Peng, Yan Wang, Lingfei Qian, Víctor Gutiérrez-Basulto
cs.AI

摘要

大規模分類檢索通常假設輸入已經表達了目標概念。然而,在許多情境中,輸入是間接證據,例如一個表格儲存格,其含義取決於所在列、欄、資料型別與上下文。我們將這種不一致稱為「檢索就緒差距」。我們的分析顯示,當語意明確時,現有索引能可靠地檢索到目標;但原始證據往往使其在排名中位居深處。我們提出「因子化假說搜尋」(Factorized Hypothesis Search, FHS),它在具名的語意維度上維護多個部分詮釋。這些假說支援結構化查詢生成、多假說檢索,以及維度層級的候選驗證。在金融分類標註與 CodiEsp 臨床編碼兩項任務中,FHS 在非 oracle 方法中取得了最佳的 Recall@1、MRR 與最終準確率。以自由文本整合方法取代因子化假說路徑,會導致頂部排序效能下降幅度最大;而序列式精煉相較於 FHS 強大的平行第一輪並未帶來額外增益。
English
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.