MODUS: デコーダのみによる多様なモダリティのAny-to-Anyモデリング
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
July 28, 2026
著者: Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, François Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin Dehghan, Jesse Allardice, David Mizrahi, Oğuzhan Fatih Kar, Roman Bachmann, Amir Zamir
cs.AI
要旨
任意間モデルは、単一のネットワーク内で任意の組み合わせのモダリティから任意のモダリティを予測するものであり、マルチモーダル視覚および視覚言語モデルで用いられる定式化であり、さらに生態学や天文学などの科学分野でもその利用が広がっている。既存の任意間モデルは通常、エンコーダ・デコーダまたは拡散アーキテクチャを用いてゼロから学習され、その性能に影響を与え、強力な事前学習済みデコーダのみのモデルを事前知識として利用することを妨げている。本研究では、すべてのモダリティを対称的に扱い、モダリティ固有のヘッド、損失関数、タスクパイプラインを必要とせずに任意のモダリティを入力および出力としてサポートする、デコーダのみの任意間マルチモーダルモデリングを調査する。すべてのモダリティが同一モデルの入力かつ出力であるため、結果として得られるモデル「Modus」は、中間モダリティを介した連鎖生成や、別の生成モダリティでモデル自身の出力をスコアリングするクロスモーダル自己検証など、様々なアプリケーションをサポートできる。Modusは、初期状態での優れた性能を示し、単一モデルで様々なベンチマークにおいて専門家およびマルチタスクのベースラインと競合する。すべての資料は https://modus-multimodal.epfl.ch/ でオープンソース公開されている。
English
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.