ChatPaper.aiChatPaper

MonkeyOCRv2: Ein Bild-Text-Basismodell für die Dokumenten-KI

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

July 13, 2026
Autoren: Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, Qiang Liu, Jiajun Song, Zidun Guo, Xinhan Wang, Handong Zheng, Yang Liu, Dongliang Luo, Zhiyin Ma, Jiarui Zhang, Xiang Bai
cs.AI

Zusammenfassung

Mainstream-visuelle-Encoder werden auf natürlichen Bildern vortrainiert und können ohne dokumentorientierte Anpassung nicht effektiv auf Dokumentbilder angewendet werden, da dichter Text und feinkörnige Zeichenstriche eine zeichenebene visuelle Wahrnehmung erfordern. Wir stellen MonkeyOCRv2 vor, ein visuell-textuell vortrainiertes Modell für Dokumenten-KI. Zunächst erstellen wir MonkeyDoc v2, nach unserem Wissen das größte Dokumentenbild-Vortrainingskorpus, das 113 Millionen Bilder in 17 Sprachen umfasst. Zweitens schlagen wir eine Vortrainingsstrategie vor, die gemeinsam Bild-zu-Text-Generierung und pixelebene Dokumentenrekonstruktion lernt: Ersteres gleicht visuelle Repräsentationen mit textuellen Inhalten ab, während Letzteres Zeichenstriche und Layoutdetails bewahrt. Umfangreiche Experimente werden an fünf repräsentativen Dokumentenanalyseaufgaben durchgeführt, darunter Texterkennung, Formelerkennung, Textdetektion, Dokumentenmanipulationserkennung und Segmentierung überlappender Texte. Das Ersetzen der ursprünglichen Encoder durch MonkeyOCRv2 verbessert die Leistung durchgängig bei allen fünf Aufgaben. Schließlich validieren wir seine Wirksamkeit als visueller Encoder von multimodalen großen Sprachmodellen bei den anspruchsvolleren Aufgaben der Dokumentenparsung und des Dokumentenverständnisses. Eingefroren und mit einem leichten Sprachmodell gekoppelt, ergibt es ein 0,7B-Dokumentenparsungsmodell, das auf MDPBench einen neuen Open-Source-State-of-the-Art setzt – einem aktuellen Benchmark, der digital erstellte und fotografierte Dokumente in 17 Sprachen umfasst – und das bisher beste 3B dots.mocr um absolute 2,8% übertrifft, bei einem etwa 11-mal kleineren visuellen Encoder. Der eingefrorene Encoder treibt auch ein Dokumentenverständnismodell an, das unter identischen Trainingsbedingungen auf acht Benchmarks die auf CLIP, DINO und SAM basierenden Gegenstücke übertrifft. Diese Ergebnisse legen nahe, dass dokumentorientiertes visuelles Vortraining als eigenständige Grundlage für die Dokumentenintelligenz dienen kann.
English
Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11times smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.