検索拡張型大規模言語モデルを用いた歴史的文書復元のための外部知識の活用
Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models
July 24, 2026
著者: Gabeen Kim, Kyeongpil Kang
cs.AI
要旨
歴史文書は貴重な知識の宝庫であるが、物理的な劣化や損傷により判読不能になることが多い。既存のマスク言語モデリングに基づく修復手法は局所的な文脈を効果的に利用するものの、外部の歴史的知識を必要とする固有表現の修復には困難を伴う。この制限に対処するため、我々は検索拡張生成(RAG)を用いた大規模言語モデルを活用する歴史文書修復のための新しい枠組みを紹介する。事前学習済みLLMの暗黙的知識と明示的に検索された外部文脈を組み合わせることで、我々のモデルARIは文脈依存の固有名詞を推論するという課題を効果的に軽減する。韓国歴史文書に関する広範な実験により、我々の手法がベースラインを大幅に上回り、一般文字と固有表現の両方の修復において顕著な改善を達成することが示された。さらに、専門家による評価を含む包括的な評価により、ARIが領域専門家にとって実用的なツールとして機能し、歴史記録の分析を加速することが期待されることが確認された。
English
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.