基于检索增强大语言模型利用外部知识的历史文档修复
Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
July 24, 2026
作者: Gabeen Kim, Kyeongpil Kang
cs.AI
摘要
历史文献作为宝贵的知识档案,常因物理退化和损坏导致字迹模糊难以辨识。现有基于掩码语言建模的修复方法虽能有效利用局部上下文,但在需要外部历史知识的名词实体复原方面存在局限。为克服这一缺陷,我们提出一种基于检索增强生成的大型语言模型历史文献修复框架。通过融合预训练语言模型的内隐知识与显式检索获取的外延上下文,我们的ARI模型有效解决了依赖上下文之专有名词推断的难题。在朝鲜半岛历史文献上的大量实验表明,本方法显著优于基线模型,在通用字符与命名实体修复方面均实现显著提升。此外,包含专家评估在内的综合评测证实,ARI可作为领域专家的实用工具,有望加速历史档案的分析进程。
English
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.