S1-Omni:面向科学理解、预测与生成的统一多模态推理模型
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
July 17, 2026
作者: Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
cs.AI
摘要
我们提出了S1-Omni,一个面向科学理解、预测与生成的统一多模态推理模型。AI for Science (AI4S) 已通过领域专用模型、工具增强的大语言模型及科学语言模型取得了显著进展。然而,模型能力仍高度碎片化,限制了异构数据、科学规律与专家知识的联合建模。S1-Omni通过将这些能力整合为一个统一的科学推理模型来弥补这一不足。S1-Omni的架构基于三个核心组件:科学数据的统一表示、自然世界知识对齐以及面向特定领域任务的解码。首先,S1-Omni将自然语言指令与科学对象(包括CIF、SMILES、蛋白质序列、光谱及科学图像)映射至共享表示空间。其次,它将科学规律与专家知识融入数据构建与训练过程,使模型能够基于科学证据进行推理。第三,它执行任务特定的解码,以支持广泛的应用场景,包括性质预测、光谱到分子生成、蛋白质位点与结构预测,以及科学图像生成与编辑。S1-Omni基于包含200个科学任务、数百万条推理样本的S1-Omni-Corpus语料库进行训练,并在超过60项科学基准上进行评估。在大多数基准上,它优于GPT-5.5与Gemini-3.1-Pro,在多项基准上达到或超越领域专用模型水平。总体而言,S1-Omni为统一的科学建模提供了一条可行路径。
English
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.