SynthDocBench: 장문맥 시각 문서 이해를 위한 제어된 벤치마크
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
July 11, 2026
저자: Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba
cs.AI
초록
비전 언어 모델(VLM)은 DocVQA, ChartQA, MMLongBench-Doc과 같은 시각 문서 이해 벤치마크에서 강력한 성능을 보여주었다. 그러나 실제 문서는 길이, 레이아웃 복잡성, 양식, 질문 난이도 등의 여러 요소가 결합되어 있어, 모델 실패의 원인을 특정하기 어렵다. 본 연구에서는 문서 길이, 레이아웃 구조, 양식 구성, 질문 유형을 체계적으로 통제하는 완전 합성 기반의 장문 맥락 시각 문서 이해 벤치마크인 SynthDocBench를 제안한다. 이 벤치마크는 조합적 설계를 통해 구축되었으며, 각 요소는 생성된 문서 전반에 걸쳐 독립적으로 변이되어 모델 행동의 통제된 분석을 가능하게 한다. 문서는 6가지 레이아웃 원형에 걸쳐 LLM 파이프라인을 사용해 종단간 생성되며, 40%의 무작위 재정의를 적용하여 모델이 허위 상관관계를 활용하는 것을 방지한다. 또한 SynthDocBench는 기존 벤치마크보다 훨씬 더 긴 길이와 구조적 다양성을 지닌 장문 맥락 문서를 포함한다. 일곱 개의 최첨단 VLM을 평가한 결과, 기존 벤치마크로는 포착할 수 없는 세 가지 실패 모드가 드러났다. 즉, 문서 길이에 따른 급격한 성능 저하, 6개 모델 중 5개에서 문서의 중간 3분의 1이 가장 어려운 체계적 위치 민감성(가장 가파른 하락 폭: 8.3 퍼센트 포인트)과 6개 모델 중 5개가 음의 초반-후반 추세를 보이는 현상, 그리고 장문 문서 환경에서의 차트 이해 붕괴이다. 이러한 결과는 현재 모델이 강건한 장문 맥락 시각 문서 이해를 달성하기보다는 벤치마크 인공물에 과적합하고 있을 가능성을 시사한다.
English
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.