对叙述者进行分级:多智能体知识系统中声明级溯源的伊斯纳德-里贾尔框架
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
July 27, 2026
作者: Ali Zahid Raja
cs.AI
摘要
现代多智能体知识系统日益通过自主转换链而非直接检索来积累知识。现有的溯源工作记录了"发生了什么"——执行轨迹、工具调用、证据链接——而信源可靠性评估(真值发现、声誉系统)早已成熟。目前缺失的是一种可操作的框架,该框架能为论断级传输链附加分领域、分级别的传输者可信度,配备完备性语义、转换类型化的聚合机制、解耦的内容批评,以及"服务/审查/隔离"路由机制。
古典伊斯兰圣训学曾面临结构相似的问题:判定通过人际传述链传播的知识是否应当接受。历经数世纪发展,它形成了一套严谨的方法论——伊斯纳德(每条论断附带的完整传述链)、里贾勒(系统化评价每位传述人诚信度与精确度)、最弱环节链评估、通过独立链进行佐证,以及马特恩批评(独立于链质量评估内容)。本文将该方法论迁移至人工智能系统设计。
我们贡献了以下成果:圣训学概念到多智能体管线的形式化映射;实现论断链与分级传述人注册表的关系型模式;结合链评级与内容批评的决策矩阵;以及对来自真实物理教科书的20,000条论断的评估。评估验证了最弱环节隔离机制与独立链佐证的效果;报告了评级恢复循环的部分失败——未能识别出最高失误率的传述人;并指出两项分析结论不明确,包括框架无法与参考内容批评者实现匹配覆盖率的对比。论文全程明确说明了证据支持与尚未支持的论断边界。
English
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing.
Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design.
We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.