ChatPaper.aiChatPaper

大型語言模型時代下的聖訓計算科學:一項批判性敘事回顧

Hadith computational science in the age of large language models: a critical narrative review

June 18, 2026
作者: Md. Ashraful Haque, Riasat Islam
cs.AI

摘要

我們探討聖訓計算科學如何被Transformer模型、基於檢索的管線與大型語言模型(LLMs)所重塑。近期回顧文獻記錄了相關研究的增長,但尚未批判性地說明哪些進展在方法論上穩健、哪些仍受制於基準評測,以及哪些未解決的問題仍限制學術應用。我們透過一項批判性敘事回顧來填補此缺口,該回顧結合對既有回顧的評論、對具代表性原始研究的逐篇評析,以及綜合伊斯蘭學者與領域專家對真實性、權威性與負責任使用的觀點。我們發現進展並不均衡。資料資源已擴展,分段任務已趨成熟,傳述人與來源驗證問題得到更完善的形式化,而LLM輔助的工作流程現已支援語料庫規模的豐富化、多語言存取與具依據的評估。與此同時,進展仍受限於狹窄的語料庫、薄弱的基準可比較性、合成資料到真實資料的遷移落差、傳述人身份解析、前處理的脆弱性、有限的可重現性,以及稀少的專家依據驗證。我們表明,重要的缺口存在於主流基準之外:非正典與冷門語料庫、注疏與解說文獻、與古蘭經和聖傳的跨來源連結,以及面向伊斯蘭法學的證據支援。我們主張,聖訓計算不應僅被評估為孤立的模型效能,而應視為一項需要知識整合、來源追溯與專家監督的證據基礎設施問題。在此基礎上,我們定義了一項研究議程,旨在使該領域在方法論上更加穩健,並對伊斯蘭學術更具實用性。
English
We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language models (LLMs). Recent reviews document growth in the literature, but they do not yet provide a critical account of which advances are methodologically robust, which remain benchmark-bound, and which unresolved problems still limit scholarly use. We address this gap through a critical narrative review that combines critique of existing reviews, paper-level appraisal of representative original studies, and synthesis of Islamic scholar and domain-expert perspectives on authenticity, authority, and responsible use. We find uneven progress. Data resources have expanded, segmentation tasks have matured, narrator and source-verification problems are better formalized, and LLM-assisted workflows now support corpus-scale enrichment, multilingual access, and grounded evaluation. At the same time, progress remains constrained by narrow corpora, weak benchmark comparability, synthetic-to-real transfer gaps, narrator identity resolution, preprocessing fragility, limited reproducibility, and sparse expert-grounded validation. We show that important gaps lie beyond dominant benchmarks: non-canonical and obscure corpora, commentary and explanatory literature, cross-source links with Qur'an and seerah, and fiqh-facing evidence support. We argue that hadith computation should be assessed less as isolated model performance than as an evidence infrastructure problem requiring knowledge integration, provenance, and expert supervision. On this basis, we define a research agenda for making the field methodologically stronger and more useful to Islamic scholarship.