ChatPaper.aiChatPaper

OmniScientist: 전(全) 모달·전(全) 학문 분야 AI 과학자

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

August 13, 2026
저자: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
cs.AI

초록

최근 기초 모델의 발전으로 AI 과학자들은 가설 생성, 코드 실행, 원고 작성에 이르기까지 점점 더 완전한 연구 워크플로우를 자동화할 수 있게 되었다. 그러나 워크플로우의 포괄성만으로는 과학적 발견의 기반이 되는 전체 증거에 대한 접근이 보장되지 않는다. 기존 시스템은 일반적으로 텍스트, 코드, 레이블 또는 사전 계산된 요약 위에서 추론하며, 과학적으로 결정적인 공간적, 시간적, 교차 채널적, 절차적 관계를 에이전트가 활용하지 못하게 한다. 우리는 OmniScientist를 소개한다. 이는 이질적인 원시 증거로부터 직접 다분야 연구를 수행하는 종단 간(end-to-end) 옴니-모달(omni-modal) AI 과학자이다. 지각 계층(perception layer)과 발상(ideation), 실험(experiment), 집필(writeup)을 담당하는 3개의 자율 에이전트가 결정론적 파이프라인 내에서 작동하여, 관찰이 연구 수명주기 전반에 걸쳐 연구 질문, 실험 결정, 최종 주장을 형성할 수 있게 한다. 아이디어 검증, 엄격성 검증, 주장 검증을 코드로 실행함으로써, 이 시스템은 참신성 심사, 통계적 타당성, 실행 추적성(execution provenance), 수치 추적 가능성을 강제한다. 우리는 5개 학문 계열, 4가지 과학적 증거 유형, 그리고 이미지, 신호, 오디오, 비디오, 3차원 구조, 궤적, 표, 수식, 그래프를 포함하는 양식에 걸친 36개의 실제 데이터 사례에서 OmniScientist를 평가한다. 이 시스템은 36개 전체 사례에서 원시 데이터부터 편집된 원고까지의 전체 경로를 완료하며, 참조 추론 백본(reference reasoning backbone)을 사용하여 평균 종합 논문 점수 6.3을 달성한다. 사전 계산된 스칼라 특징만을 입력으로 받는 맹검(blind) 변형과의 쌍별 비교에서, 직접 지각은 7개 평가 차원 모두에서 우수했으며 일대일 판단의 85%에서 승리하였다. 이러한 결과는 수명주기 전반에 걸친 지각이 증거 기반 과학적 발견에 필수적이며, 광범위하게 유능한 AI 과학자를 향한 실용적 경로를 제공함을 보여준다.
English
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.