ChatPaper.aiChatPaper

「評級敘述者:多智能體知識系統中基於聲明層級溯源之伊斯納德-里賈爾框架」

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

July 27, 2026
作者: Ali Zahid Raja
cs.AI

摘要

現代多智能體知識系統日益透過自主轉換鏈而非直接檢索來累積知識。現有的溯源工作記錄了「發生了什麼」——執行軌跡、工具呼叫、證據連結——而來源可靠性評估(如真相發現、聲譽系統)早已確立。目前所欠缺的,是一個具備操作性的框架,能為主張層級的傳播鏈附加依領域劃分的可細粒度傳輸者可靠性評級,並納入完整性語義、經轉換類型聚合的機制、去耦的內容批評,以及提供、審查、隔離的流程管理。 古典伊斯蘭聖訓學曾面對結構類似的問題:判斷透過人類敘述者鏈條傳遞的知識是否應被接受。經過數個世紀的發展,它建立了一套嚴謹的方法論——伊斯納德(每一主張附帶完整傳播鏈)、里賈爾(系統評定每位敘述者的誠信與精確度)、最弱鏈評估機制、透過獨立鏈進行佐證,以及馬特恩批評(不依賴鏈品質而獨立評估內容)。本文將該方法論轉移至人工智慧系統設計。 我們提出的貢獻包括:從聖訓學概念到多智能體管線的形式化映射、實現主張鏈與評級敘述者註冊表的關聯式結構、結合鏈評級與內容批評的決策矩陣,以及在來自真實物理教科書的兩萬條主張上的評估。評估結果驗證了最弱鏈隔離與獨立鏈佐證的有效性;同時報告了評級恢復迴圈的部分失敗——該機制未能識別出最高錯誤率的敘述者;另有兩項分析結果無法得出結論,其中包括一項該框架因無法觸及對照內容批評者而未能完成的匹配覆蓋比較。本文自始至終明確指出哪些主張有證據支持,哪些尚無證據支持。
English
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing. Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design. We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.