FinanceHarness:自主金融深度研究框架
FinanceHarness: Autonomous Financial Deep Research Framework
July 30, 2026
作者: Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
cs.AI
摘要
得益于LLM和自主智能体的进步,深度研究已成为被最广泛采用的智能体产品之一。然而,大多数深度研究系统生成通用报告,难以满足金融深度研究的需求。金融研究需要专业知识来分析历史模式并预测未来事件。因此,自动化金融深度研究既需要分层式的驱动框架来引导研究智能体,也需要一个可验证的时点基准,以防止未来信息泄露。我们提出FinanceHarness,这是一个运行面向金融的工具和从业者引导工作流的框架,能够端到端地自动化金融深度研究:包括环境和数据构建、智能体执行循环以及奖励建模。我们进一步提出FinanceGym,其中包含以投资论点驱动的研究问题,以及结合信息截止日前与截止日后标准的评分细则。专业专家验证通过率达82%。即使是领先的LLM和智能体,其在评分细则上的得分也低于40%,这表明FinanceGym具有挑战性,且仍有巨大的提升空间。在相同的开放权重骨干模型下,FinanceHarness将整体评分从25.3%提升至32.4%。FinanceHarness可在 https://github.com/Yijia-Xiao/FinanceHarness 获取。
English
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at https://github.com/Yijia-Xiao/FinanceHarness.