OmniScientist:一种全模态、全学科的AI科学家
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
August 13, 2026
作者: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
cs.AI
摘要
基础模型的最新进展使AI科学家能够自动化日益完整的研究工作流程,从假设生成、代码执行到论文撰写。然而,仅覆盖工作流程并不能提供科学发现所依赖的完整证据。现有系统通常基于文本、代码、标签或预计算摘要进行推理,使得具有科学决定性的空间、时间、跨通道和程序性关系无法被智能体获取。我们引入OmniScientist,一个端到端的全模态AI科学家系统,可直接基于异质原始证据开展多学科研究。一个感知层和三个自主智能体(分别负责构思、实验和撰写)在确定性流水线中协同运作,使观察能够在整个研究生命周期中塑造研究问题、实验决策和最终结论。通过在代码中执行想法检查、严谨性检查和主张检查,系统实现了新颖性筛查、统计有效性、执行溯源和数值可追溯性。我们在涵盖5个学科大类、4类科学证据以及图像、信号、音频、视频、三维结构、轨迹、表格、公式和图等多种模态的36个真实数据案例上评估了OmniScientist。该系统在所有36个案例中完成了从原始数据到编译稿全文的完整路径,并在参考推理主干下取得了6.3分的平均论文综合评分。在与仅接收预计算标量特征的盲变体的配对比较中,直接感知在所有7个评估维度上均有所提升,并在85%的成对比较判断中胜出。这些结果表明,贯穿研究生命周期的感知能力对于基于证据的科学发现至关重要,并为构建广泛能力的AI科学家提供了切实可行的路径。
English
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.