ChatPaper.aiChatPaper

S1-Omni: 과학적 이해, 예측 및 생성을 위한 통합 다중 모달 추론 모델

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

July 17, 2026
저자: Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
cs.AI

초록

S1-Omni는 과학적 이해, 예측 및 생성을 위한 통합 멀티모달 추론 모델입니다. 과학을 위한 AI(AI4S)는 도메인 특화 모델, 도구 보강 LLM, 과학 언어 모델을 통해 크게 발전해 왔습니다. 그러나 모델의 기능은 여전히 매우 파편화되어 있어 이질적 데이터, 과학 법칙 및 전문가 지식의 공동 모델링에 한계가 있습니다. S1-Omni는 이러한 격차를 해소하기 위해 이러한 기능들을 하나의 일관된 과학 추론 모델로 통합합니다. S1-Omni의 아키텍처는 과학 데이터의 통합 표현, 자연 세계 지식 정렬, 도메인 특화 태스크를 위한 디코딩이라는 세 가지 핵심 구성 요소를 기반으로 구축되었습니다. 첫째, S1-Omni는 자연어 명령과 CIF, SMILES, 단백질 서열, 스펙트럼, 과학 이미지 등의 과학 객체를 공유 표현 공간으로 매핑합니다. 둘째, 데이터 구축 및 학습에 과학 법칙과 전문가 지식을 통합하여 모델이 과학적 증거를 바탕으로 추론할 수 있도록 합니다. 셋째, 물성 예측, 스펙트럼-분자 생성, 단백질 부위 및 구조 예측, 과학 이미지 생성 및 편집 등 광범위한 응용을 지원하기 위해 태스크별 디코딩을 수행합니다. S1-Omni는 200개의 과학 태스크를 포괄하고 수백만 개의 추론 샘플을 포함하는 S1-Omni-코퍼스로 학습되었으며, 60개 이상의 과학 벤치마크에서 평가되었습니다. 대부분의 벤치마크에서 GPT-5.5 및 Gemini-3.1-Pro를 능가하며, 여러 벤치마크에서는 도메인 특화 모델과 동등하거나 더 나은 성능을 보입니다. 전반적으로 S1-Omni는 통합 과학 모델링을 위한 실용적인 경로를 제공합니다.
English
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.