FinanceComplexQA:基於工業級金融文件的代理推理基準測試
FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
July 21, 2026
作者: Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, Jian Yang, Xinyuan Wang, Ming Zhou, Weixiao Zhou, Xiangyuan Guan, Xiang Li, Zhenhe Wu, Ziyi Ni, Zhoujun Li, Bingjing Xu
cs.AI
摘要
代理推理因其整合大规模信息并生成可靠准确内容的能力,已成为金融分析领域的一场变革性力量。然而,在处理复杂的现实问题时,不同代理仍表现出显著的性能差异。本文设计了Finance-LaTeX SKILL——一种基于专家知识合成复杂版式金融文档的技能。利用基于该技能构建的代理工作流,我们生成了2,000份专业金融文档及6,000个高质量问答对。为全面评估代理的综合能力,我们引入FinanceComplexQA——一个贴近真实场景的金融文档综合性开放生成基准。该基准包含2,026项深度研究任务,覆盖1,009份金融文档。FinanceComplexQA具有八大核心特征:双语支持;覆盖六类主流场景及七类任务;具备专家级文档推理问题;支持复杂版式深度研究;提供相对稳定且可永久参考的标准答案;通过"代理即裁判"机制结合多项评估指标实现精准评价。利用FinanceComplexQA,我们对领先的RAG系统及代理推理工具在金融文档问答中的表现进行了全面评估。通过识别与分析失败案例,我们深入研究了这些工具在数值计算、多跳推理、内容摘要及行业分析等方面的能力。
English
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.