不值再費一個Token:高效深度研究智能體的邊際價值估計
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
August 9, 2026
作者: Harshitha Kolukuluru, Reshma Ashok, Kirat Arora, Evan William Ciccarelli, Nischal Ashok Kumar, Lunyiu Nie, Franck Dernoncourt, Samyadeep Basu, Ryan A. Rossi, Nedim Lipka
cs.AI
摘要
長程研究代理透過迭代檢索、彙整與綜合來解決開放式任務,但隨著脈絡快速增長,額外證據的邊際價值往往遞減。這導致不必要的代幣成本、更高的延遲,以及最終報告生成時更嘈雜的輸入。我們研究深度研究代理中脈絡管理的邊際價值估計,並提出首個系統性的跨流程階段感知修剪策略比較。我們在檢索前、檢索後與綜合前三個階段,評估輕量級啟發式標準與學習式價值模型。結果顯示,修剪效果更多取決於應用的階段而非具體評分規則:早期修剪能帶來最大的端到端節省,而後期修剪主要優化最終綜合脈絡。輕量級啟發式方法最多可減少73%的代幣使用量,且品質下降極小;學習式修剪在特定取捨下仍具競爭力,但沒有任何單一方法能在品質、效率與忠實度上全面勝出。這些發現為設計高效的長程代理系統提供了實務指引。
English
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation. We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline. We evaluate lightweight heuristic criteria and a learned value model at pre-retrieval, post-retrieval, and pre-synthesis stages. Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness. These findings provide practical guidance for designing efficient long-horizon agentic systems.