내레이터 등급 평가: 다중 에이전트 지식 시스템에서 주장 수준 출처 추적을 위한 이스나드-리잘 프레임워크
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
July 27, 2026
저자: Ali Zahid Raja
cs.AI
초록
현대의 다중 에이전트 지식 시스템은 직접 검색보다는 자율적 변환 체인을 통해 점점 더 지식을 축적합니다. 기존의 출처 추적 연구는 실행 추적, 도구 호출, 증거 링크 등 발생한 일을 기록하며, 출처 신뢰도 추정(진실 발견, 평판 시스템)은 오래전부터 확립되어 있습니다. 부족한 것은 등급화된 도메인별 전송자 신뢰도를 주장 수준 전송 체인에 연결하고, 완전성 의미론, 변환 유형별 집계, 분리된 내용 비판, 제공/검토/격리 라우팅을 갖춘 운영 프레임워크입니다.
고전 이슬람 하디스 학문은 구조적으로 유사한 문제에 직면했습니다: 인간 전승자들의 체인을 통해 전달된 지식을 받아들일지 여부를 결정하는 것이었습니다. 수세기에 걸쳐 엄격한 방법론을 발전시켰습니다 - 이스나드(모든 주장에 첨부된 완전한 전송 체인), 리잘(각 전승자의 청렴성과 정밀성에 대한 체계적 등급 매기기), 최약 고리 체인 평가, 독립적 체인을 통한 입증, 그리고 마트 비판(체인 품질과 독립적으로 평가된 내용). 이 논문은 이 방법론을 AI 시스템 설계에 이전합니다.
우리는 하디스 학문 개념에서 다중 에이전트 파이프라인으로의 형식적 매핑, 주장 체인과 등급화된 전승자 등록부를 구현하는 관계형 스키마, 체인 등급과 내용 비판을 결합한 의사 결정 행렬, 그리고 실제 물리학 교과서에서 가져온 20,000개의 주장에 대한 평가를 제시합니다. 평가는 최약 고리 격리와 독립적 체인 입증을 검증하며; 최고 결함 전승자를 놓친 등급 복구 루프의 부분적 실패를 보고하며; 참조 내용 비판과 프레임워크가 도달할 수 없었던 일치 커버리지 비교를 포함하여 두 가지 분석이 결론이 나지 않았음을 보고합니다. 이 논문은 증거가 지지하는 주장과 아직 지지하지 않는 주장이 무엇인지 전반적으로 명시적입니다.
English
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems). What is missing is an operational framework that attaches graded, per-domain transmitter reliability to claim-level transmission chains, with completeness semantics, transformation-typed aggregation, decoupled content criticism, and serve/review/quarantine routing.
Classical Islamic hadith science confronted a structurally similar problem: deciding whether knowledge transmitted through chains of human narrators should be accepted. Over centuries it developed a rigorous methodology - isnad (a complete transmission chain attached to every claim), rijal (systematic grading of each narrator's integrity and precision), weakest-link chain evaluation, corroboration through independent chains, and matn criticism (content evaluated independently of chain quality). This paper transfers that methodology to AI system design.
We contribute a formal mapping from hadith-science concepts to multi-agent pipelines, a relational schema implementing claim chains and a graded narrator registry, a decision matrix combining chain grade with content criticism, and an evaluation on 20,000 claims from real physics textbooks. The evaluation validates weakest-link quarantine and independent-chain corroboration; reports a partial failure of the grade-recovery loop, which missed the highest-fault narrator; and reports two analyses as inconclusive, including a matched-coverage comparison the framework could not reach with the reference content critic. The paper is explicit throughout about which claims the evidence does and does not yet support.