언어 지능을 위한 확장 가능한 시각적 사전 학습
Scalable Visual Pretraining for Language Intelligence
July 10, 2026
저자: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
cs.AI
초록
대규모 기반 모델의 급속한 발전은 주로 대규모 텍스트 말뭉치에 대한 사전 학습에 의해 주도되어 왔다. 그러나 많은 형태의 지식은 시각적 표현을 통해 전달되며, 그림, 조판된 수식, 페이지 레이아웃은 텍스트만으로는 정확하거나 완전하게 포착할 수 없는 풍부한 정보를 담고 있다. 하지만 현재의 사전 학습 접근법은 문서 및 웹 페이지와 같은 시각적으로 풍부한 자료를 언어 지능 학습을 위해 일반 텍스트로 변환함으로써 이러한 시각적 단서를 무시한다. 본 논문은 언어 모델이 반드시 텍스트 전용 표현으로 학습되어야 한다는 기본 가정에 도전하며, 시각적 사전 학습이 기반 모델 지능을 위한 확장 가능한 학습자임을 보여준다. 이를 위해, 우리는 텍스트 추출 없이 시각적 문서를 직접 활용하는 비지도 시각적 사전 학습 패러다임에 대한 체계적 연구를 수행한다. 여러 백본과 벤치마크에 걸쳐, 동일한 기반 말뭉치에 대한 시각적 사전 학습은 텍스트 전용 사전 학습보다 일관되게 우수한 성능을 보이며, 확장 가능한 언어 지능을 위한 효율적인 경로를 제공한다.
English
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.