SynthDocBench: 長文脈ビジュアル文書理解のための制御されたベンチマーク

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

July 11, 2026
著者: Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba
cs.AI

要旨

視覚言語モデル(VLM)は、DocVQA、ChartQA、MMLongBench-Docなどの視覚的文書理解ベンチマークにおいて高い性能を達成している。しかし、現実世界の文書は、長さ、レイアウトの複雑さ、モダリティ、質問の難易度など複数の要素を組み合わせており、モデルの失敗を特定の原因に帰属させることが困難である。我々はSynthDocBenchを紹介する。これは、長文脈の視覚的文書理解のための完全に合成されたベンチマークであり、文書の長さ、レイアウト構造、モダリティ構成、質問タイプなどの要素を体系的に制御する。このベンチマークは組み合わせ設計を用いて構築されており、生成された文書間で各要素が独立に変動するため、モデルの振る舞いの制御された分析が可能となる。文書は、6つのレイアウト原型にわたってLLMパイプラインを用いてエンドツーエンドで生成され、モデルが疑似相関を利用するのを防ぐために40%のランダム上書きが適用される。さらに、SynthDocBenchは既存のベンチマークよりもはるかに長い長さと構造的多様性を持つ長文脈文書を網羅している。7つの最先端VLMを評価した結果、既存のベンチマークでは明らかにできない3つの失敗モードを発見した:文書の長さによる急激な性能低下、6モデルのうち5モデルで文書の中間3分の1が最も困難であるという体系的な位置感度、そして6モデルのうち5モデルが負の初期から後期への傾向を示すこと(最大低下:8.3パーセントポイント)、および長文書設定におけるグラフ理解の崩壊である。これらの結果は、現在のモデルがロバストな長文脈視覚的文書理解を達成するのではなく、ベンチマークのアーティファクトに過学習している可能性を示唆している。
English
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.
PDF521July 16, 2026