SynthDocBench:長上下文視覺文件理解之受控基準測試

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

July 11, 2026
作者: Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba
cs.AI

摘要

視覺語言模型(VLM)已在 DocVQA、ChartQA 和 MMLongBench-Doc 等視覺文檔理解基準測試中展現出優異表現。然而,真實世界的文檔往往同時包含長度、版面複雜度、模態以及問題難度等多種因素,使得難以將模型失敗歸因於特定原因。為此,我們提出了 SynthDocBench——一個完全合成的長上下文視覺文檔理解基準,能夠系統性地控制文檔長度、版面結構、模態組成及問題類型等因素。該基準採用組合設計建構,每個因素在生成的文檔中獨立變化,從而實現對模型行為的可控分析。文檔透過 LLM 流程端到端生成,涵蓋六種版面原型,並內建 40% 的隨機覆蓋機制,以防止模型利用虛假相關性。此外,SynthDocBench 包含的文檔在長度和結構多樣性上遠超現有基準。通過評估七個前沿 VLM,我們發現了現有基準無法揭示的三種失敗模式:文檔長度增加時性能急遽下降;系統性的位置敏感性——五分之六的模型在文檔中間三分之一處表現最差,且五分之六的模型呈現負向的「早期至晚期」趨勢(最大下降幅度達 8.3 個百分點);以及在長文檔場景下圖表理解能力的崩潰。這些結果表明,當前模型可能過度擬合於基準測試的虛假特徵,而非真正實現了穩健的長上下文視覺文檔理解。
English
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.
PDF521July 16, 2026