ChatPaper.aiChatPaper

MODUS: 다양한 모달리티에 대한 디코더 전용 Any-to-Any 모델링

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

July 28, 2026
저자: Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, François Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin Dehghan, Jesse Allardice, David Mizrahi, Oğuzhan Fatih Kar, Roman Bachmann, Amir Zamir
cs.AI

초록

Any-to-any 모델은 단일 네트워크 내에서 다른 모든 조합으로부터 모든 모달리티를 예측하는 방식으로, 멀티모달 비전 및 비전-언어 모델에서 사용되며 생태학, 천문학과 같은 과학 분야에서도 점점 더 활용되고 있다. 기존의 Any-to-any 모델은 일반적으로 인코더-디코더 또는 확산 아키텍처를 사용하여 처음부터 학습되며, 이는 성능에 영향을 미치고 강력한 사전 학습된 디코더 전용 모델을 사전 지식으로 활용하지 못하게 한다. 본 연구에서는 모든 모달리티를 대칭적으로 처리하고, 모달리티별 헤드, 손실 함수 또는 작업 파이프라인 없이 임의의 모달리티를 입력 및 출력으로 지원하는 디코더 전용 Any-to-any 멀티모델링을 조사한다. 모든 모달리티가 동일 모델의 입력이자 출력이기 때문에, 결과 모델인 Modus는 중간 모달리티를 통한 연쇄 생성이나 다른 생성된 모달리티로 자체 출력을 평가하는 교차 모달 자체 검증 등 다양한 응용을 지원할 수 있다. Modus는 바로 사용 가능한 강력한 성능을 보여주며, 단일 모델로 다양한 벤치마크에서 전문가 및 멀티태스크 기준선과 경쟁력을 갖춘다. 모든 자료는 https://modus-multimodal.epfl.ch/에서 오픈소스로 공개된다.
English
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.