OmniScientist:全模態全領域AI科學家
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
August 13, 2026
作者: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
cs.AI
摘要
近期基礎模型的進展使人工智慧科學家能夠自動化日益完整的研究流程,從假說生成、程式碼執行到稿件準備。然而,僅有流程涵蓋範圍並無法讓代理取得科學發現所依賴的完整證據。現有系統通常僅基於文字、程式碼、標籤或預計算摘要進行推理,使得具有科學決定性的空間、時間、跨通道及程序關係無法被代理取得。我們提出 OmniScientist,一個端對端、全模態的人工智慧科學家,可直接從異質原始證據進行跨學科研究。感知層與三個分別負責構思、實驗與撰寫的自主代理在確定性流程中運作,使觀察能在整個研究生命週期中塑造研究問題、實驗決策與最終主張。透過以程式碼執行構思、嚴謹性與主張檢查,系統強制進行新穎性篩查、統計有效性、執行溯源與數值可追溯性。我們在涵蓋5個學科類別、4類科學證據以及包括影像、訊號、音訊、影片、3D結構、軌跡、表格、公式與圖形等模態的36個真實資料案例上評估 OmniScientist。該系統在所有36個案例中完成從原始資料到編譯稿件的完整路徑,並在參考推理主幹下獲得平均論文總分6.3。在與僅接收預計算標量特徵的盲測變體進行配對比較時,直接感知在所有7個評估維度上均有提升,並在85%的兩兩比較判斷中勝出。這些結果顯示,全生命週期感知對於基於證據的科學發現至關重要,並為實現廣泛能力的人工智慧科學家提供了切實可行的途徑。
English
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.