NaviDC-OCR: デジタル文書とカメラ撮影文書を横断する文書解析
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
August 13, 2026
著者: Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang, Jiawei Tang, Kaicheng Yang, Meng Tong, Zhongjiang He, Hao Sun
cs.AI
要旨
文書解析は、非構造化文書を構造化された機械可読な表現に変換することを目的とする。近年の視覚言語モデル(VLM)の進歩により、文書解析は大幅に進展している。しかしながら、既存手法には依然として2つの大きな課題がある。第一に、分離型VLMベース手法は正確なレイアウト解析に大きく依存しており、カメラ撮影文書における幾何学的歪みが連鎖的な誤りを引き起こす可能性がある。第二に、エンドツーエンドのVLMベース手法は明示的なレイアウト検出への依存を軽減するものの、高解像度のシナリオにおいて冗長な生成、幻覚、および構造推論の不足に悩まされることが多い。これらの課題に対処するため、我々は文書解析のための統合フレームワークであるNaviDC-OCRを提案する。NaviDC-OCRは、変形を考慮した学習を導入して幾何学的認識をVLMに組み込み、複雑なレイアウト表現のための適応的サンプリング機構を提案する。さらに、数式文法と表構造を明示的にモデル化する内容・構造分離学習戦略を開発し、より効果的な構造化表現学習を可能にする。大規模な実験により、NaviDC-OCRが多様な文書解析ベンチマークで最先端の性能を達成することが示された。具体的には、OmniDocBench v1.6、Wild-OmniDocBench、PureDocBenchにおいてそれぞれ96.87、88.53、78.41の総合スコアを獲得し、ICDAR 2026 Sci-ImageMinerチャレンジで第1位を獲得した。これらの結果は、複雑な文書解析シナリオにおけるNaviDC-OCRの有効性と汎化能力を実証するものである。
English
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.