經由檢索增強大型語言模型利用外部知識進行歷史文獻修復
Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
July 24, 2026
作者: Gabeen Kim, Kyeongpil Kang
cs.AI
摘要
歷史文獻作為寶貴的知識檔案,常因物理劣化與損毀而難以辨識。現有的修復方法雖基於遮蔽語言建模有效利用局部語境,但在處理需外部歷史知識的命名實體時力有未逮。為解決此限制,我們提出一新穎的歷史文獻修復框架,結合大型語言模型與檢索增強生成技術。透過融合預訓練語言模型的內隱知識與明確檢索的外在語境,我們的模型ARI有效克服推論語境依賴專有名詞的挑戰。廣泛的韓國歷史文獻實驗顯示,本方法顯著優於基準模型,在還原一般字符與命名實體上均取得重大進展。此外,包含專家評估的全面性評量證實ARI可作為領域專家的實用工具,有望加速歷史記錄的分析。
English
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.