AutoResearch: 통찰이 들어가면, 환각이 나온다
AutoResearch: Insight In, Hallucination Out
August 23, 2026
저자: Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
cs.AI
초록
자율 연구 시스템은 점점 더 긴 연구 워크플로우를 실행할 수 있게 되었지만, 자동화만으로는 결과적인 과정이 과학적으로 근거를 유지한다는 것을 보장하지 않는다. 우리는 아이디어 생성(Idea Generation)과 아이디어 실행(Idea Execution)을 연결하는 2단계 시스템인 AutoResearch를 제안한다. 이 시스템은 연구 아이디어가 형성되는 방식과 실험을 통해 그 아이디어가 신뢰할 수 있게 확립되는 방식을 모두 다룬다. 아이디어 생성 단계에서 AutoResearch는 새로운 연구 신호를 축적된 도메인 지식과 지속적으로 통합하고, 전이 가능한 메커니즘적 통찰을 식별하며, 다중 모델 생성과 교차 검토를 활용하여 근거가 있고 검증 가능한 연구 계획을 수립한다. 아이디어 실행 단계에서는 조정된 에이전트들이 이러한 계획을 실험으로 분해하고, 반복적으로 구현 및 진단하며, 연구 결론을 수용하기 전에 독립적인 증거 기반 검토를 수행한다. 교차 모달 검색, 시스템 최적화, 벤치마크 기반 머신러닝의 대표적 설정에서 AutoResearch는 생성된 아이디어를 측정 가능한 진전으로 전환하고, 신뢰할 수 없는 실험 결과를 탐지 및 교정하며, 연구 방향을 계속하거나 수정하거나 종료하는 증거 조건부 결정을 내린다. 예를 들어, RSICD 벤치마크에서 AutoResearch가 생성한 아이디어는 평균 재현율(Recall)을 32.84에서 34.69로 향상시켰으며, 다른 자율 연구 시스템의 11~27건과 비교하여 감사로 확인된 문제 이벤트는 단 5건만 기록하였다. 이러한 결과는 의미 있는 통찰이 실험 전에 근거를 갖추고, 결론이 수용 전에 근거를 갖추는 연구 과정을 입증한다: 통찰이 들어가고, 환각이 나온다(Insight In, Hallucination Out).
English
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.