ChatPaper.aiChatPaper

AutoResearch:洞察を入れ、幻覚を出す

AutoResearch: Insight In, Hallucination Out

August 23, 2026
著者: Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
cs.AI

要旨

自律的研究システムは、長大な研究ワークフローを実行する能力をますます高めているが、自動化だけでは、そのプロセスが科学的に裏付けられたものとして維持されることは保証されない。本稿では、アイデア生成とアイデア実行を結びつける二段階システムであるAutoResearchを紹介する。これは、研究アイデアがどのように形成されるか、そして実験を通じてどのように信頼性をもって確立されるかの双方に対処するものである。アイデア生成段階では、AutoResearchは、新たに生じる研究シグナルを蓄積されたドメイン知識と継続的に統合し、転移可能なメカニズム的洞察を特定し、マルチモデル生成と相互レビューを用いて、根拠に基づいた検証可能な研究計画を生成する。アイデア実行段階では、協調するエージェント群がこれらの計画を実験に分解し、反復的に実装と診断を行い、研究結論を受け入れる前に独立したエビデンスに基づくレビューを採用する。クロスモーダル検索、システム最適化、ベンチマーク駆動型機械学習における代表的な設定を通じて、AutoResearchは生成されたアイデアを測定可能な進歩に変え、信頼できない実験結果を検出・修正し、エビデンスに基づく条件付き判断によって研究の方向性を継続・修正・終了する。例えば、RSICDベンチマークでは、AutoResearchが生成したアイデアにより平均リコールが32.84から34.69に向上し、監査で確認された問題イベントはわずか5件であり、他の自律的研究システムの11〜27件と比較される。これらの結果は、意味のある洞察が実験前に根拠付けられ、結論が受け入れ前に根拠付けられる研究プロセス、すなわち「洞察を入れ、幻覚を出さない」(Insight In, Hallucination Out) ことを実証している。
English
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.