言語知能のためのスケーラブルな視覚的事前学習

Scalable Visual Pretraining for Language Intelligence

July 10, 2026
著者: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
cs.AI

要旨

大規模基盤モデルの急速な進歩は、主に大規模なテキストコーパスによる事前学習によって推進されてきた。しかし、多くの知識は視覚的表現を通じて伝達され、図、組版された数式、ページレイアウトなどは、テキストだけでは正確かつ完全に捉えられない豊かな情報を内包している。にもかかわらず、現在の事前学習手法では、文書やウェブページなどの視覚的に豊かな情報源をプレーンテキストに変換することでこれらの視覚的手がかりを捨象し、言語知能の学習を行っている。本論文では、言語モデルはテキストのみの表現で訓練されなければならないというデフォルトの前提に疑問を投げかけ、視覚的事前学習が基盤モデル知能のためのスケーラブルな学習手法であることを示す。この目的のため、我々はテキスト抽出を行わずに視覚的文書を直接活用する、教師なし視覚的事前学習パラダイムの体系的な研究を行う。複数のバックボーンとベンチマークにおいて、同一の基礎コーパスに対する視覚的事前学習は、テキストのみの事前学習を一貫して上回り、スケーラブルな言語知能への効率的な経路を提供する。
English
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.
PDF410July 14, 2026