ChatPaper.aiChatPaper

FinanceComplexQA: 산업용 금융 문서에서의 에이전트적 추론 벤치마킹

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

July 21, 2026
저자: Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, Jian Yang, Xinyuan Wang, Ming Zhou, Weixiao Zhou, Xiangyuan Guan, Xiang Li, Zhenhe Wu, Ziyi Ni, Zhoujun Li, Bingjing Xu
cs.AI

초록

에이전트 추론(Agentic Reasoning)은 대규모 정보를 통합하고 신뢰할 수 있으며 정확한 콘텐츠를 생성하는 능력으로 인해 재무 분석 분야에서 변혁적인 힘으로 자리 잡았다. 그러나 복잡한 실제 문제를 처리할 때 서로 다른 에이전트 간에 여전히 상당한 성능 차이가 나타난다. 본 연구에서는 전문가 지식을 기반으로 복잡한 레이아웃의 금융 문서를 합성하는 스킬인 Finance-LaTeX SKILL을 설계한다. 이 스킬을 기반으로 구축된 에이전트 워크플로우를 활용하여 2,000개의 전문 금융 문서와 6,000개의 고품질 질문-답변 쌍을 생성한다. 에이전트의 전반적 능력을 평가하기 위해 실제 시나리오와 매우 유사한 금융 문서 대상의 포괄적인 개방형 생성 벤치마크인 FinanceComplexQA를 도입한다. 이는 1,009개의 금융 문서를 대상으로 하는 2,026개의 심층 연구 과제를 포함한다. FinanceComplexQA는 8가지 주요 특징을 갖는다: 이중 언어 지원; 6가지 주요 시나리오와 7가지 과제를 포괄함; 전문가 수준의 문서 추론 질문; 복잡한 레이아웃에 대한 심층 연구; 비교적 안정적이고 영구적인 참조 답변; 다중 평가 지표를 이용한 Agent-as-a-Judge 기반의 정밀 평가. FinanceComplexQA를 활용하여 금융 문서 QA를 위한 주요 RAG 시스템과 에이전트 추론 도구에 대한 포괄적 평가를 수행한다. 실패 사례를 식별하고 분석함으로써 수치 계산, 다중 단계 추론, 콘텐츠 요약 및 산업 분석 측면에서 이들의 능력에 대한 심층 연구를 제공한다.
English
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.