ChatPaper.aiChatPaper

더 이상의 토큰은 가치가 없다: 효율적인 딥 리서치 에이전트를 위한 한계 가치 추정

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

August 9, 2026
저자: Harshitha Kolukuluru, Reshma Ashok, Kirat Arora, Evan William Ciccarelli, Nischal Ashok Kumar, Lunyiu Nie, Franck Dernoncourt, Samyadeep Basu, Ryan A. Rossi, Nedim Lipka
cs.AI

초록

장기 지평 연구 에이전트는 반복적 검색, 집계, 종합을 통해 개방형 작업을 해결하지만, 컨텍스트는 빠르게 증가하는 반면 추가 증거의 한계 가치는 종종 감소한다. 이는 불필요한 토큰 비용, 더 높은 지연 시간, 최종 보고서 생성을 위한 입력의 잡음 증가를 초래한다. 본 연구는 심층 연구 에이전트에서 컨텍스트 관리를 위한 한계 가치 추정을 다루며, 파이프라인 전반에 걸친 가지치기 전략의 첫 번째 체계적인 단계별 비교를 제시한다. 우리는 검색 전, 검색 후, 종합 전 단계에서 경량 휴리스틱 기준과 학습된 가치 모델을 평가한다. 결과에 따르면 가지치기 효과는 특정 점수 규칙보다 가지치기가 적용되는 위치에 더 크게 의존한다. 초기 가지치기는 가장 큰 종단 간 절감을 가져오는 반면, 후기 가지치기는 주로 최종 종합 컨텍스트를 정제한다. 경량 휴리스틱은 품질 저하가 거의 없이 토큰 사용량을 최대 73%까지 줄이며, 학습된 가지치기는 선택된 절충 측면에서 경쟁력을 유지한다. 품질, 효율성, 충실성 전반에 걸쳐 어떤 단일 방법도 지배적이지 않다. 이러한 발견은 효율적인 장기 지평 에이전트 시스템을 설계하기 위한 실용적 지침을 제공한다.
English
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation. We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline. We evaluate lightweight heuristic criteria and a learned value model at pre-retrieval, post-retrieval, and pre-synthesis stages. Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness. These findings provide practical guidance for designing efficient long-horizon agentic systems.