自动研究:洞见入,幻觉出
AutoResearch: Insight In, Hallucination Out
August 23, 2026
作者: Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
cs.AI
摘要
自主研究系统日益能够执行长周期研究工作流,然而自动化本身并不能确保研究过程始终保持科学根基。我们提出AutoResearch,一个连接想法生成与想法执行的两阶段系统,以同时解决研究想法如何形成以及如何通过实验可靠地确立这两个问题。在想法生成阶段,AutoResearch持续整合新兴研究信号与积累的领域知识,识别可迁移的机制性洞见,并利用多模型生成与交叉评审来产生有根基、可检验的研究方案。在想法执行阶段,协同智能体将这些方案分解为实验,迭代实施并诊断这些实验,并在接受研究结论前进行独立的基于证据的审查。在跨模态检索、系统优化和基准驱动的机器学习等代表性场景中,AutoResearch将生成的想法转化为可衡量的进展,检测并纠正不可靠的实验结果,并做出基于证据条件的决策,以继续、修订或终止研究方向。例如,在RSICD基准上,AutoResearch生成的想法将平均召回率从32.84提升至34.69,同时仅记录5次经审计确认的问题事件,而其他自主研究系统为11至27次。这些结果表明了一种研究过程:有意义的洞见在实验之前就已奠定根基,结论在接纳之前就已获得依据:洞见输入,幻觉输出。
English
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.