S1-Omni: 科学的理解、予測、生成のための統合型マルチモーダル推論モデル
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
July 17, 2026
著者: Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
cs.AI
要旨
本稿では、科学理解・予測・生成のための統一的なマルチモーダル推論モデル「S1-Omni」を提案する。AI for Science(AI4S)は、ドメイン特化型モデルやツール拡張型大規模言語モデル、科学用言語モデルの発展により大きく進歩してきた。しかし、各モデルの能力は依然として高度に断片化されており、異種データや科学法則、専門家知識の統合的モデリングに限界がある。S1-Omniは、これらの能力を単一の一貫した科学推論モデルに統合することで、この課題に取り組む。S1-Omniのアーキテクチャは、科学データの統一表現、自然世界の知識アライメント、ドメイン固有タスクのためのデコードという3つの中核要素から成る。第一に、S1-Omniは自然言語による指示と、CIF、SMILES、タンパク質配列、スペクトル、科学画像を含む科学オブジェクトを、共有表現空間にマッピングする。第二に、科学法則と専門家知識をデータ構築とトレーニングに組み込み、モデルが科学的根拠に基づいて推論できるようにする。第三に、タスク固有のデコードを実行し、物性予測、スペクトルから分子への生成、タンパク質部位・構造予測、科学画像の生成・編集など、幅広いアプリケーションをサポートする。S1-Omniは、200の科学タスクをカバーし数百万の推論サンプルを含むS1-Omni-Corpusでトレーニングされ、60以上の科学ベンチマークで評価された。その結果、ほとんどのベンチマークでGPT-5.5やGemini-3.1-Proを上回り、いくつかのベンチマークではドメイン特化型モデルと同等かそれを上回る性能を達成した。全体として、S1-Omniは統一的な科学モデリングへの実用的な道筋を提供する。
English
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.