ChatPaper.aiChatPaper

統合マルチモーダル生成としてのビジョン

Vision as Unified Multimodal Generation

July 7, 2026
著者: Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, Siyi Xie, Xin You, Jinsheng Quan, Zhongang Cai, Haiwen Diao, Ziwei Liu, Lei Yang, Dahua Lin, Quan Wang
cs.AI

要旨

我々はコンピュータビジョンを統一的なマルチモーダル生成として定式化する。この枠組みでは、タスク固有のアーキテクチャを用いることなく、異種の視覚タスクを統一的なマルチモーダルモデル本来のテキストおよび画像生成空間の中で表現する。この定式化に基づき、SenseNova-Visionは自然言語指示とオプションの視覚的プロンプトを用いてタスク、対象領域や視点、復号化規則を指定し、応答として、シンボル出力にはテキスト、密な空間予測には画像、構成タスクにはテキストと画像の混合出力を生成する。大規模学習を支援するため、多様なコンピュータビジョンアノテーションをこれらの生成空間と互換性のある指示応答例に変換し、テキスト、画像、および混合ターゲットにわたるコンピュータビジョン指示応答コーパスであるSenseNova-Vision Corpusを構築した。SenseNova-Visionは、既存の事前学習済み統一マルチモーダルモデルを出発点とし、主にこのコーパスで訓練され、補助的なマルチモーダルデータは能力維持のための混合として使用される。そのため、タスク固有の予測ヘッドやアーキテクチャの変更は不要である。得られたモデルは、検出、OCR、キーポイント推定、セグメンテーション、深度推定、表面法線推定、ポイントマップ、カメラ姿勢推定など広範な視覚タスクをカバーし、さらにカテゴリ、色、領域、その他の視覚的手がかりを組み合わせた言語定義のバリアントにも対応する。実験により、単一の統一モデルが、構造化された視覚理解、密な幾何学的予測、セグメンテーション、多視点視覚幾何学において、最先端のタスク特化システムに匹敵することが示された。これらの結果は、統一的なマルチモーダル生成が、コンピュータビジョン機能を汎用基盤モデルに統合するためのスケーラブルな経路であることを示唆している。モデルとコーパスは公開されている。
English
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.