自動研究:洞見輸入,幻覺輸出
AutoResearch: Insight In, Hallucination Out
August 23, 2026
作者: Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
cs.AI
摘要
自主研究系統日益具備執行長期研究工作流程的能力,然而僅靠自動化並不能確保整個過程維持科學嚴謹性。我們提出AutoResearch,一個兩階段系統,將想法生成與想法執行相銜接,以同時解決研究想法如何形成,以及如何透過實驗可靠地確立這些想法。在想法生成階段,AutoResearch持續整合新興研究訊號與累積的領域知識,識別可遷移的機制性洞見,並運用多模型生成與交叉審查,產出具嚴謹基礎且可驗證的研究計畫。在想法執行階段,協調式代理將這些計畫拆解為實驗,迭代地實作與診斷,並在採納研究結論前,進行獨立的證據本位審查。在跨模態檢索、系統最佳化及基準驅動的機器學習等代表性情境中,AutoResearch能將生成的想法轉化為可衡量的進展,偵測並修正不可靠的實驗結果,並做出以證據為條件的決策——繼續、修訂或終止研究方向。例如,在RSICD基準中,一個由AutoResearch生成的想法將平均Recall從32.84提升至34.69,且僅記錄了5次經稽核確認的問題事件,而其他自主研究系統則為11至27次。這些結果證明了一套研究流程:有意義的洞見在實驗之前即已確立基礎,結論在接受之前亦已確立基礎——洞見輸入,幻覺輸出。
English
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.