ChatPaper.aiChatPaper

語り手の評価:マルチエージェント知識システムにおける主張レベルの来歴のためのイスナード・リジャールの枠組み

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

July 27, 2026
著者: Ali Zahid Raja
cs.AI

要旨

現代のマルチエージェント知識システムは、直接的な検索ではなく、自律的な変換の連鎖を通じて知識を蓄積する傾向が強まっている。既存の来歴研究は、実行トレース、ツール呼び出し、証拠リンクといった「何が起こったか」を記録しており、ソース信頼性の推定(真実発見、評判システム)も長年にわたり確立されている。しかし不足しているのは、主張レベルの伝達連鎖に対して、ドメイン別に段階付けられた送信者信頼性を付与し、完全性の意味論、変換タイプ別の集約、内容批判の分離、配信/レビュー/隔離ルーティングを備えた、運用可能なフレームワークである。 古典的なイスラム聖訓学(ハディース学)は、構造的に類似した問題に直面していた。すなわち、人間の伝承者の連鎖を通じて伝達された知識を受け入れるべきかどうかを判断することである。何世紀にもわたって、厳格な方法論が発展してきた。すなわち、イスナード(あらゆる主張に付随する完全な伝達連鎖)、リジャール(各伝承者の誠実性と正確性を体系的に等級付け)、最弱リンクによる連鎖評価、独立した連鎖による裏付け、そしてマトン批判(連鎖の品質とは独立に内容を評価する)である。本稿は、この方法論をAIシステム設計に応用する。 我々の貢献は、ハディース学の概念からマルチエージェントパイプラインへの形式的マッピング、主張連鎖と等級付き伝承者レジストリを実装する関係スキーマ、連鎖等級と内容批判を組み合わせた決定マトリックス、および実際の物理学の教科書から得られた20,000件の主張に対する評価である。評価では、最弱リンク隔離と独立した連鎖による裏付けを検証した。また、最も欠陥の多い伝承者を見逃した等級回復ループの部分的な失敗を報告し、フレームワークが参照用の内容批判器との一致範囲比較に到達できなかったことを含む2つの分析を結論不確定として報告する。本稿では、証拠が裏付ける主張とまだ裏付けていない主張について、一貫して明示的に述べている。
English
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.