FinanceComplexQA: 産業用金融文書におけるエージェント的推論のベンチマーキング
FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
July 21, 2026
著者: Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, Jian Yang, Xinyuan Wang, Ming Zhou, Weixiao Zhou, Xiangyuan Guan, Xiang Li, Zhenhe Wu, Ziyi Ni, Zhoujun Li, Bingjing Xu
cs.AI
要旨
エージェンティック推論(Agentic Reasoning)は、大規模情報を統合し、信頼性と正確性の高いコンテンツを生成できる能力により、金融分析において変革をもたらす力となっている。しかし、複雑な実世界の問題を扱う際には、エージェント間で依然として顕著な性能差が生じる。本研究では、専門家の知識に基づき複雑なレイアウトの金融文書を合成するスキル「Finance-LaTeX SKILL」を設計する。このスキルを基盤としたエージェントワークフローを用いて、2,000件の専門的金融文書と、6,000件の高品質な質問応答ペアを生成する。エージェントの総合的な能力を評価するため、実世界のシナリオに極めて近い金融文書向けの包括的なオープンエンド生成ベンチマーク「FinanceComplexQA」を導入する。これは1,009件の金融文書を対象とする2,026件の深層調査タスクを含む。FinanceComplexQAは8つの主要な特徴を持つ:バイリンガル対応、6つの主要シナリオと7つのタスクを網羅、専門家レベルの文書推論問題、複雑なレイアウトに対する深層調査、比較的安定かつ永続的な参照回答、そして複数の評価指標によるAgent-as-a-Judge方式の精密な評価である。FinanceComplexQAを用いて、金融文書QAにおける主要なRAGシステムとエージェンティック推論ツールの包括的な評価を実施する。失敗事例の特定と分析を通じて、数値計算、マルチホップ推論、内容要約、業界分析におけるそれらの能力を詳細に検討する。
English
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.