S1-Omni:面向科學理解、預測與生成的統一多模態推理模型
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
July 17, 2026
作者: Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
cs.AI
摘要
我們提出 S1-Omni,一個用於科學理解、預測與生成的統一多模態推理模型。AI for Science (AI4S) 透過領域專用模型、工具增強的 LLM 以及科學語言模型取得了顯著進展。然而,模型能力仍然高度分散,限制了對異質數據、科學定律與專家知識的聯合建模。S1-Omni 透過將這些能力整合到一個單一且連貫的科學推理模型中,填補了這一空白。S1-Omni 的架構建立在三個核心組件之上:科學數據的統一表示、自然世界知識對齊,以及領域特定任務的解碼。首先,S1-Omni 將自然語言指令與科學物件(包括 CIF、SMILES、蛋白質序列、光譜以及科學圖像)映射到一個共享的表示空間。其次,它將科學定律和專家知識融入數據構建與訓練過程中,使模型能夠依據科學證據進行推理。第三,它執行任務特定的解碼,以支援廣泛的應用,包括屬性預測、光譜到分子生成、蛋白質位點與結構預測,以及科學圖像生成與編輯。S1-Omni 在 S1-Omni-Corpus 上進行訓練,該語料庫涵蓋 200 項科學任務並包含數百萬個推理樣本,並在超過 60 個科學基準測試上進行評估。它在大多數基準測試上優於 GPT-5.5 和 Gemini-3.1-Pro,並在多項基準測試上達到或超越領域專用模型的表現。總體而言,S1-Omni 為實現統一的科學建模提供了一條實用路徑。
English
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.