FinMCP-Bench: Het benchmarken van LLM-agenten voor het gebruik van financiële tools in de praktijk volgens het Model Context Protocol
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
March 26, 2026
Auteurs: Jie Zhu, Yimin Tian, Boyang Li, Kehao Wu, Zhongzhi Liang, Junhui Li, Xianyin Zhang, Lifan Guo, Feng Chen, Yong Liu, Chi Zhang
cs.AI
Samenvatting
Dit artikel introduceert FinMCP-Bench, een nieuwe benchmark voor het evalueren van grote taalmodellen (LLM's) bij het oplossen van praktische financiële problemen door middel van tool-aanroeping via financiële modelcontextprotocollen. FinMCP-Bench bevat 613 voorbeelden verdeeld over 10 hoofdsenario's en 33 subsenario's, met zowel echte als synthetische gebruikersvragen om diversiteit en authenticiteit te waarborgen. Het integreert 65 echte financiële MCP's en drie soorten voorbeelden: enkelvoudige tool, meervoudige tools en multi-turn, waardoor evaluatie van modellen op verschillende niveaus van taakcomplexiteit mogelijk is. Met behulp van deze benchmark evalueren we systematisch een reeks mainstream LLM's en introduceren we metrieken die tool-aanroepnauwkeurigheid en redeneervermogen expliciet meten. FinMCP-Bench biedt een gestandaardiseerde, praktische en uitdagende testomgeving voor het bevorderen van onderzoek naar financiële LLM-agenten.
English
This paper introduces FinMCP-Bench, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of financial model context protocols. FinMCP-Bench contains 613 samples spanning 10 main scenarios and 33 sub-scenarios, featuring both real and synthetic user queries to ensure diversity and authenticity. It incorporates 65 real financial MCPs and three types of samples, single tool, multi-tool, and multi-turn, allowing evaluation of models across different levels of task complexity. Using this benchmark, we systematically assess a range of mainstream LLMs and propose metrics that explicitly measure tool invocation accuracy and reasoning capabilities. FinMCP-Bench provides a standardized, practical, and challenging testbed for advancing research on financial LLM agents.