ChatPaper.aiChatPaper

픽셀 텍스트 표현 학습의 설계 기본 원리에 대하여

On the Design Fundamentals of Pixel Text Representation Learning

September 1, 2026
저자: Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang, Yu Rong, Hong Cheng, Hou Pong Chan, Chenghao Xiao
cs.AI

초록

텍스트가 풍부한 시각 입력은 픽셀 공간에서 언어를 직접 읽고, 검색하고, 압축할 수 있는 모델을 요구하지만, 기존의 픽셀-텍스트 인코더는 고정 해상도 사전학습, 시각적 지름길 학습, 약한 시각적 접지, 다국어 시각 텍스트 이해에서 어려움을 겪는다. 본 연구에서는 강건한 시각 텍스트 표현 학습에 필요한 근본적인 설계 원칙을 조사한다. 체계적인 통제 절제 실험을 통해 우리는 네 가지 핵심 요소를 식별한다: 다양한 이미지 해상도와 렌더링된 글꼴 크기는 고해상도 문서 일반화를 위한 공간적 프록시를 제공하며, 자연 이미지-텍스트 쌍은 접지에 필수적이고 텍스트 전용 붕괴를 방지하며, 레이아웃을 인지하는 렌더링은 픽셀 수준의 지름길 학습을 방지하는 데 도움이 되고, 2단계 다국어 커리큘럼은 효과적인 교차 언어 정렬을 가능하게 한다. 이러한 원칙들을 확장 가능한 학습 레시피에 통합함으로써, 우리는 실시간 렌더링, 통합 대조 접지, 그리고 2억 8천만 개의 학습 예제에 대한 다국어 커리큘럼으로 학습된 네이티브 해상도 비전 인코더인 Pixel Linguist II를 학습시킨다. Pixel Linguist II는 영어, 교차 언어, 다국어 Visual STS 및 ViDoRe에서 새로운 최첨단 결과를 달성하며, 더 나은 MLLM 다운스트림 평가 또한 가능하게 한다. 특히, Pixel Linguist II는 80%의 시각 토큰 압축에서도 강건함을 유지하여 광학 맥락 압축에 큰 가능성을 보여준다. 우리의 코드와 리소스는 https://github.com/Pixel-Linguist/Pixel-Linguist-II에서 확인할 수 있다.
English
Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretraining, visual shortcut learning, weak visual grounding, and multilingual visual text understanding. In this work, we investigate the fundamental design principles required for robust visual text representation learning. Through systematic controlled ablations, we identify four critical components: variable image resolutions and rendered font sizes provide spatial proxies for high-resolution document generalization; natural image-text pairs are indispensable for grounding and prevent text-only collapse; layout-aware rendering helps prevent pixel-level shortcuts; and a two-stage multilingual curriculum enables effective cross-lingual alignment. By integrating these principles into a scalable training recipe, we train Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples. Pixel Linguist II sets new state-of-the-art results on English, cross-lingual, and multilingual Visual STS and ViDoRe, while also enabling better MLLM downstream evaluation. Notably, Pixel Linguist II remains robust under 80\% visual token compression, showing great promise for optical context compression. Our code and resources are available at https://github.com/Pixel-Linguist/Pixel-Linguist-II.