可擴展視覺預訓練於語言智能

Scalable Visual Pretraining for Language Intelligence

July 10, 2026
作者: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
cs.AI

摘要

大型基础模型的快速进步主要归因于在大规模文本语料库上的预训练。然而,许多形式的知识是通过视觉表征传递的,其中图表、排版公式和页面布局蕴含了文本单独无法忠实或完整捕捉的丰富信息。然而,当前的预训练方法通过将视觉丰富的资源(例如文档和网页)转换为纯文本来进行语言智能学习,从而丢弃了这些视觉线索。本文挑战了语言模型必须在纯文本表征上训练的默认假设,并展示了视觉预训练是基础模型智能的一种可扩展学习器。为此,我们对直接利用视觉文档而无需文本提取的无监督视觉预训练范式进行了系统性研究。在多个骨干网络和基准测试中,基于相同底层语料库的视觉预训练始终优于纯文本预训练,为可扩展的语言智能提供了一条高效路径。
English
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.
PDF410July 14, 2026