SynthDocBench: kontrollierte Benchmark für das Verständnis visueller Dokumente mit langem Kontext

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

July 11, 2026
Autoren: Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, Sagar Davasam, Vikas Yadav, Sai Rajeswar Mudumba
cs.AI

Zusammenfassung

Vision-Language-Modelle (VLMs) haben auf Benchmarks zum visuellen Dokumentverständnis wie DocVQA, ChartQA und MMLongBench-Doc starke Ergebnisse erzielt. Allerdings kombinieren reale Dokumente mehrere Faktoren wie Länge, Layoutkomplexität, Modalität und Frageschwierigkeit, was es erschwert, Modellfehler auf spezifische Ursachen zurückzuführen. Wir stellen SynthDocBench vor, einen vollständig synthetischen Benchmark für das visuelle Dokumentverständnis im Langkontext, der systematisch Faktoren wie Dokumentlänge, Layoutstruktur, modale Zusammensetzung und Fragetyp kontrolliert. Der Benchmark wird mittels eines kombinatorischen Designs erstellt, wobei jeder Faktor über die generierten Dokumente hinweg unabhängig variiert wird, was eine kontrollierte Analyse des Modellverhaltens ermöglicht. Dokumente werden Ende-zu-Ende mit einer LLM-Pipeline über sechs Layout- Archetypen hinweg generiert, wobei eine 40-prozentige zufällige Überschreibung verhindert, dass Modelle von Scheinkorrelationen profitieren. Darüber hinaus umfasst SynthDocBench Langkontext-Dokumente mit wesentlich größerer Länge und struktureller Vielfalt als bestehende Benchmarks. Bei der Evaluierung von sieben führenden VLMs decken wir drei Fehlermodi auf, die bestehende Benchmarks nicht sichtbar machen können: eine starke Verschlechterung mit der Dokumentlänge, eine systematische positionsbezogene Empfindlichkeit, bei der das mittlere Drittel eines Dokuments für fünf von sechs Modellen am schwierigsten ist und fünf von sechs Modellen einen negativen Früh-zu-Spät-Trend aufweisen (stärkster Rückgang: 8,3 Prozentpunkte), sowie ein Versagen des Diagrammverständnisses in Langkontext-Settings. Diese Ergebnisse deuten darauf hin, dass aktuelle Modelle möglicherweise eher auf Benchmark-Artefakte überangepasst sind, als dass sie ein robustes visuelles Dokumentverständnis im Langkontext erreichen.
English
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documents combine multiple factors such as length, layout complexity, modality, and question difficulty, which makes it difficult to attribute model failures to specific causes. We introduce SynthDocBench, a fully synthetic benchmark for long-context visual document understanding that systematically controls factors including document length, layout structure, modality composition, and question type. The benchmark is constructed using a combinatorial design, each factor is varied independently across generated documents, enabling controlled analysis of model behavior. Documents are generated end to end using an LLM pipeline across six layout archetypes, with a 40 percent random override to prevent models from exploiting spurious correlations. Additionally, SynthDocBench spans long-context documents with substantially greater length and structural diversity than existing benchmarks. Evaluating seven frontier VLMs, we uncover three failure modes that existing benchmarks cannot surface: sharp degradation with document length, a systematic positional sensitivity in which the middle third of a document is hardest for five of six models and five of six models show a negative Early-to-Late trend (steepest decline: 8.3 percentage points), and breakdown of chart comprehension in long-document settings. These results suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context visual document understanding.
PDF521July 16, 2026