ChatPaper.aiChatPaper

통합 다중 모달 생성으로서의 비전

Vision as Unified Multimodal Generation

July 7, 2026
저자: Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, Siyi Xie, Xin You, Jinsheng Quan, Zhongang Cai, Haiwen Diao, Ziwei Liu, Lei Yang, Dahua Lin, Quan Wang
cs.AI

초록

우리는 컴퓨터 비전을 통합 다중 모드 생성(unified multimodal generation)으로 정식화하며, 이질적인 시각 작업들이 특정 작업에 특화된 아키텍처 없이 통합 다중 모드 모델의 고유한 텍스트 및 이미지 생성 공간에서 표현된다. 이러한 정식화 하에서 SenseNova-Vision은 자연어 명령과 선택적 시각 프롬프트를 사용하여 작업, 대상 영역 또는 뷰, 디코딩 규칙을 지정하고, 응답을 기호 출력을 위한 텍스트, 조밀한 공간 예측을 위한 이미지, 또는 구성적 작업을 위한 혼합 텍스트-이미지 출력으로 생성한다. 대규모 훈련을 지원하기 위해, 우리는 다양한 컴퓨터 비전 주석을 이러한 생성 공간과 호환되는 명령-응답 예제로 변환하여, 텍스트, 이미지 및 혼합 대상을 포괄하는 컴퓨터 비전 명령-응답 코퍼스인 SenseNova-Vision Corpus를 구축했다. 기성 사전 훈련된 통합 다중 모드 모델로부터 시작하여, SenseNova-Vision은 주로 이 코퍼스에서 훈련되며, 보조 다중 모드 데이터는 능력 보존 혼합물로 사용되며, 작업별 예측 헤드나 아키텍처 수정이 필요하지 않다. 결과 모델은 탐지, OCR, 키포인트 추정, 세분화, 깊이 추정, 표면 법선 예측, 포인트 맵, 카메라 포즈 추정을 포함한 광범위한 비전 작업을 다루며, 카테고리, 색상, 영역 및 기타 시각적 단서를 결합하는 언어 정의 변형을 지원한다. 실험 결과, 단일 통합 모델이 구조적 시각 이해, 조밀한 기하학적 예측, 세분화, 다중 뷰 시각 기하학에서 최고 수준의 작업 특화 시스템과 성능이 일치할 수 있음을 보여준다. 이러한 결과는 통합 다중 모드 생성이 컴퓨터 비전 기능을 범용 기반 모델에 통합하기 위한 확장 가능한 경로임을 시사한다. 모델과 코퍼스는 공개적으로 이용 가능하다.
English
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.