ChatPaper.aiChatPaper

AREX: 재귀적으로 자기 개선하는 심층 연구 에이전트를 향하여

AREX: Towards a Recursively Self-Improving Agent for Deep Research

July 23, 2026
저자: Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, Hongwang Xiao, Zheng Liu, Lei Xiong, Jiahao Wang, Sen Wang, Xiyan Jiang, Wanli Li, Yuyang Hu, Hongjin Qian, Bingyu Yan, Ziyi Xia, Yingxia Shao, Kang Liu, Zhicheng Dou, Di He, Chaozhuo Li, Qiwei Ye, Zhongyuan Wang, Zheng Liu
cs.AI

초록

심층 연구 에이전트는 여러 제약 조건을 동시에 충족하는 답변을 찾아야 한다. 이러한 답변을 발견하는 데는 많은 비용이 드는 반면, 후보 답변을 검증하는 것은 개별 제약 조건별로 분해하여 다루기 쉬운 확인 작업으로 수행할 수 있는 경우가 많다. 이러한 발견-검증 비대칭성은 연구 에이전트가 단순히 더 오래 검색하는 것을 넘어, 중간 결과를 검증하고 부분적으로 검증된 상태를 활용하여 후속 개선을 안내함으로써 현재 답변을 재귀적으로 개선해야 함을 시사한다. 우리는 재귀적 자기 개선(RSI) 심층 연구 에이전트 제품군인 AREX를 소개한다. AREX는 증거를 수집하고 잠정적 답변을 구성하는 내부 연구 루프와, 답변을 제약 조건별로 감사하고 확인되지 않은 주장을 식별하며 표적 후속 연구를 수행하는 외부 자기 개선 루프를 번갈아 실행한다. 장기간에 걸친 RSI를 유지하기 위해 AREX는 외부 모델에 의존하지 않고, 증가하는 상호작용 이력을 검증된 증거와 해결되지 않은 제약 조건을 보존하는 간결한 개선 상태로 압축하는 자율적 컨텍스트 업데이트 도구를 학습한다. 우리는 에이전트 중간 훈련과 장기 지평 강화 학습을 통해 검증된 합성 과제와 고품질 궤적을 사용하여 AREX를 훈련한다. 장기 학습 중 드문 최종 보상 문제를 완화하기 위해 결정적 증거가 확보되거나 잘못된 연구 방향이 수정되는 핵심 단계를 강조한다. 우리는 조밀한 4B 모델과 122B-A10B 전문가 혼합 모델을 구현한다. BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam(HLE) 및 기타 추론 및 도구 사용 벤치마크에서 AREX는 유사한 규모의 기준 모델을 상당히 능가하며, 실질적으로 더 많은 활성화 파라미터를 사용하는 모델과도 경쟁력을 유지한다.
English
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.