ChatPaper.aiChatPaper

UniSpace: 통합 시각 표현 및 확장 가능한 멀티모달 모델링

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling

August 9, 2026
저자: Jinbo Yan, Limeng Qiao, Jie Qin, Junyan He, Feize Wu, Guanglu Wan
cs.AI

초록

의미론적 비전 인코더는 이미지 생성에서 다중 모달 이해와 의미론적 조건화를 위한 핵심 시각 인터페이스가 되었다. 그러나 이들의 최종 토큰은 세밀한 시각적 세부 사항을 버려 픽셀 재구성 성능이 낮아지고, 이미지 생성 및 편집과 같은 재구성에 민감한 작업에서의 사용이 제한된다. 본 연구에서는 사전 훈련된 의미론적 ViT로 구축된 단일 시각 표현 공간에서 이해, 생성, 편집을 모델링할 수 있는지 살펴본다. 우리는 의미론적 ViT의 동결된 트랜스포머 블록이 본질적으로 시각적 세부 사항을 보존하지 못하는 것이 아님을 보여준다. 오히려 원래의 패치 파라미터화가 표현을 의미론적 추상화로 이끌어, 최종 토큰에서 세밀한 정보를 복구하기 어렵게 만든다. 이 관찰에 기반하여, 우리는 원래의 의미론적 경로를 보존하면서도 동일한 동결 ViT 블록에 세밀한 시각 정보를 제공하는 재구성 인지 패치 임베딩을 추가하는 패치 재파라미터화(Patch Reparameterization)를 도입한다. 결과적인 통합 표현은 다중 모달 이해를 보존하면서 고충실도 이미지 재구성과 유리한 재구성-생성 트레이드오프를 가능하게 한다. 우리는 이 표현을 별도의 VAE 경로 없이 동일한 시각 공간에서 이해, 생성, 편집을 수행하는 80억(8B) 규모의 Mixture-of-Transformer-Experts 모델인 UniSpace로 확장한다. 시스템 수준 평가는 실용적인 텍스트-이미지 생성과 지시 기반 이미지 편집을 입증하며, 재파라미터화된 사전 훈련 ViT가 확장 가능한 다중 모달 모델링을 위한 통합 시각 인터페이스로 작동할 수 있음을 보여준다.
English
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce Patch Reparameterization, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off. We further scale this representation into UniSpace, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.