エビデンスからタクソノミーへの検索のための因子分解型仮説探索
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
August 6, 2026
著者: Linhai Ma, Ethan F. Wei, Xueqing Peng, Yan Wang, Lingfei Qian, Víctor Gutiérrez-Basulto
cs.AI
要旨
大規模タクソノミー検索は、入力がすでに対象概念を明示していることを前提とすることが多い。しかし、多くの設定では、入力は間接的証拠にすぎない。例えば、表のセルの意味は、その行、列、データ型、文脈に依存する。我々はこの不一致を検索準備性ギャップ(retrieval readiness gap)と呼ぶ。我々の分析は、現在のインデックスは入力の意味が明示的な場合には対象を確実に検索する一方で、生の証拠はしばしば対象をランキングの深い位置に残すことを示している。そこで我々は、名前付き意味次元にわたって複数の部分解釈を保持するFactorized Hypothesis Search(FHS:因子化仮説検索)を提案する。これらの仮説により、構造化クエリの生成、複数仮説検索、次元レベルの候補検証が可能になる。金融タクソノミータグ付けとCodiEsp臨床コーディングタスクの両方において、FHSは非オラクル手法の中で最高のRecall@1、MRR、最終精度を達成する。因子化仮説経路を自由テキストアンサンブルに置き換えると、上位ランキング性能の低下が最も大きい。一方、逐次精緻化はFHSの強力な並列第1ラウンドに対して追加の利得をもたらさない。
English
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose meaning depends on its row, column, datatype, and context. We call this mismatch the retrieval readiness gap. Our analysis shows that the current index retrieves the target reliably when its semantics are explicit, while raw evidence often leaves it deep in the ranking. We propose Factorized Hypothesis Search (FHS), which maintains multiple partial interpretations over named semantic dimensions. These hypotheses support structured query rendering, multi-hypothesis retrieval, and dimension-level candidate verification. On both financial taxonomy tagging and CodiEsp clinical coding tasks, FHS achieves the best Recall@1, MRR, and final accuracy among the non-oracle methods. Replacing the factorized hypothesis path with a free-text ensemble causes the largest drop in head-ranking performance, while sequential refinement provides no additional gain over FHS's strong parallel first round.