AREX: 深層研究のための再帰的自己改善エージェントに向けて
AREX: Towards a Recursively Self-Improving Agent for Deep Research
July 23, 2026
著者: Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
cs.AI
要旨
深層リサーチでは、複数の制約を同時に満たす解をエージェントが見つける必要がある。そのような解の発見にはコストがかかる一方、候補解の検証は多くの場合、制約ごとの扱いやすいチェックに分解できる。この発見-検証の非対称性は、リサーチエージェントが単に検索時間を延ばす以上に、中間結果を検証し、部分的に検証済みの状態を利用してその後の改良を導くことで、現在の解を再帰的に改善すべきであることを示唆している。本稿では、再帰的自己改善(RSI)を行う深層リサーチエージェントのファミリーであるAREXを紹介する。AREXは、証拠を収集して暫定的な解答を構築する内部リサーチループと、解答を制約ごとに監査し、未解決の主張を特定し、対象を絞ったフォローアップリサーチを開始する外部自己改善ループを交互に実行する。長期間にわたるRSIを持続するため、AREXは自律的なコンテキスト更新ツールを学習する。このツールは、増大する相互作用履歴を、検証済みの証拠と未解決の制約を保持するコンパクトな改善状態に圧縮し、外部モデルに依存しない。我々はAREXを、検証済みの合成タスクと高品質な軌跡を用いて、エージェント的ミッドトレーニングと長期強化学習により訓練する。長期学習における最終報酬の疎らさを緩和するため、決定的な証拠が獲得されるか、誤った研究方向が修正される重要なステップを強調する。我々は高密度の4Bモデルと122B-A10Bの混合エキスパートモデルを実装した。BrowseComp、WideSearch、DeepSearchQA、Humanity's Last Exam(HLE)、その他の推論およびツール使用ベンチマークにおいて、AREXは同等規模のベースラインを大幅に上回り、より多くの活性化パラメータを使用するモデルと競合する性能を示す。
English
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.