ChatPaper.aiChatPaper

NaviDC-OCR: 디지털 및 카메라 촬영 문서 전반의 문서 파싱 탐색

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

August 13, 2026
저자: Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun
cs.AI

초록

문서 파싱은 비정형 문서를 구조화되고 기계가 읽을 수 있는 표현으로 변환하는 것을 목표로 한다. 최근 비전-언어 모델(VLM)의 발전은 문서 파싱을 크게 진전시켰다. 그러나 기존 접근 방식은 여전히 두 가지 주요 과제에 직면해 있다. 첫째, 분리된 VLM 기반 방법은 정확한 레이아웃 분석에 크게 의존하며, 카메라로 촬영된 문서의 기하학적 왜곡은 연쇄 오류를 유발할 수 있다. 둘째, 엔드투엔드 VLM 기반 방법은 명시적 레이아웃 탐지에 대한 의존성을 완화하지만, 고해상도 시나리오에서 중복 생성, 환각, 불충분한 구조 추론으로 인해 어려움을 겪는 경우가 많다. 이러한 과제를 해결하기 위해 우리는 문서 파싱을 위한 통합 프레임워크인 NaviDC-OCR을 제안한다. NaviDC-OCR은 변형 인지 학습을 도입하여 기하학적 인식을 VLM에 통합하고, 복잡한 레이아웃 표현을 위한 적응형 샘플링 메커니즘을 제안한다. 또한 콘텐츠-구조 분리 학습 전략을 개발하여 수식 문법과 표 구조를 명시적으로 모델링함으로써 보다 효과적인 구조화된 표현 학습을 가능하게 한다. 광범위한 실험을 통해 NaviDC-OCR이 다양한 문서 파싱 벤치마크에서 최첨단 성능을 달성함을 입증한다. OmniDocBench v1.6, Wild-OmniDocBench, PureDocBench에서 각각 96.87, 88.53, 78.41의 전체 점수를 획득하고, ICDAR 2026 Sci-ImageMiner Challenge에서 1위를 차지한다. 이러한 결과는 복잡한 문서 파싱 시나리오에서 NaviDC-OCR의 효율성과 일반화 능력을 검증한다.
English
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.