ChatPaper.aiChatPaper

검색 증강 대형 언어 모델을 통한 역사 문서 복원을 위한 외부 지식 활용

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

July 24, 2026
저자: Gabeen Kim, Kyeongpil Kang
cs.AI

초록

역사 문서는 귀중한 지식 아카이브 역할을 하지만, 물리적 노후화와 손상으로 인해 판독이 어려운 경우가 많다. 기존의 마스킹 언어 모델링 기반 복원 방법은 국소적 맥락을 효과적으로 활용하지만, 외부 역사 지식이 필요한 명명된 개체를 복원하는 데는 어려움을 겪는다. 이러한 한계를 해결하기 위해, 본 연구는 검색 증강 생성(Retrieval-Augmented Generation, RAG)을 활용한 대규모 언어 모델 기반의 새로운 역사 문서 복원 프레임워크를 제안한다. 사전 학습된 LLM의 암묵적 지식과 명시적으로 검색된 외부 맥락을 결합함으로써, 제안 모델 ARI는 맥락 의존적 고유 명사 추론의 어려움을 효과적으로 완화한다. 한국 역사 문서에 대한 광범위한 실험 결과, 제안 방법은 일반 문자와 명명된 개체 복원 모두에서 기준 모델들을 현저히 능가하는 성능 향상을 달성함을 보여준다. 또한 전문가 평가를 포함한 종합적 평가는 ARI가 도메인 전문가를 위한 실용적 도구로 기능하여 역사 기록 분석을 가속화할 수 있음을 확인한다.
English
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.