ChatPaper.aiChatPaper

기관 신문 파이프라인: 역사적 신문에서 수십억 개의 고품질 토큰 추출

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

August 19, 2026
저자: Matteo Cargnelutti, Catherine Brobston, Eben English, Jake Sadow, Kacie Bailey, Greg Leppert, Amanda Watson, Jessica Chapel, Jonathan Zittrain
cs.AI

초록

역사적 신문은 공공 생활에 대한 풍부한 기록이지만, 밀도 높고 불규칙하며 때로는 노이즈가 많은 레이아웃으로 인해 이러한 자료에 대한 컴퓨팅 접근은 어렵고 제한적이다. 우리는 보스턴 공공도서관과 공동으로 설계한 모듈식 시스템인 Institutional Newspapers Pipeline을 제시한다. 이 시스템은 역사적 신문 스캔에서 고품질의 구조화된 데이터셋을 추출한다. 이 파이프라인은 각 단계가 해석 가능하고 사용자 정의가 가능하며, 파이프라인 전체가 워크스테이션급 하드웨어에서 실행될 수 있을 정도로 계산 효율적으로 유지되도록 설계되었다. 파이프라인은 각 스캔을 다단계 프로세스로 처리한다. 즉, 스캔을 유형에 구애받지 않는 개별 크롭으로 분할하고, 각 결과 세그먼트에 대해 OCR을 수행한 후, 모든 크롭에 대해 텍스트 분석, 유형 분류, 판독 순서 감지, 개체명 인식, 주제 분류, 언어 감지 및 사전 계산된 임베딩 생성을 수행한다. 우리는 보스턴 공공도서관 소장 자료의 일부에 이 파이프라인을 적용하고, 그 결과를 개방형 데이터셋으로 공개했다. 광학 문자 인식(OCR) 출력은 1795년부터 1930년 사이에 출판된 1,473,635건의 퍼블릭 도메인 신문 스캔에서 추출한 8,310만 개의 개별 크롭에 걸쳐 총 163억 개의 o200k_base 토큰으로 구성된다. 본 보고서는 각 처리 단계의 방법론, 우리가 학습시킨 소형 모델, 그리고 그 과정에서 수집한 평가 결과와 데이터셋 규모 측정치를 설명한다. 본 보고서는 파이프라인, 모델 및 데이터셋의 공개와 함께 제공된다. 우리는 이 작업을 수천만 건의 신문 스캔에서 고품질 데이터를 확보하기 위한 중대한 진전으로서 자리매김한다.
English
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers Pipeline, a modular system we jointly designed with Boston Public Library to extract high-quality, structured datasets from historical newspaper scans. It was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware. The pipeline runs each scan through a multi-step process: it segments scans into individual type-agnostic crops and performs OCR on each resulting segment before then performing text analysis, type classification, reading order detection, named entities recognition, subject classification, language detection, and pre-computed embeddings generation on every crop. We ran this pipeline against a portion of Boston Public Library's holdings and released the results as an open dataset. The optical character recognition (OCR) output represents 16.3 billion o200k_base tokens across 83.1 million individual crops, extracted from 1,473,635 public domain newspaper scans published between 1795 and 1930. This report describes our methods for each processing step, the small models we trained, as well as the evaluation results and dataset-scale measurements we collected in the process. It accompanies the release of the pipeline, models, and dataset. We position this work as a substantial step towards unlocking high-quality data from tens of millions of newspaper scans.