AdvancedMathBench:一個用於高等數學證明生成與驗證的基準套件
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
July 13, 2026
作者: Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen
cs.AI
摘要
大型語言模型(LLMs)在高中與奧林匹亞風格的數學問題上已展現卓越表現,但其在高等數學領域的能力仍未被充分理解。然而,現有基準在範疇與評估粒度上皆有不足:它們的學科覆蓋範圍有限,且往往僅依賴最終答案的正確性或粗略的判斷,導致推理過程的有效性難以獲得充分評估。為填補此缺口,我們提出AdvancedMathBench,一套專為評估高等數學推理能力而設計的基準套件。其核心的證明生成基準ProverBench包含296道題目,涵蓋大學部與博士資格考的水平。為提供可靠的證明評估,我們開發了一套專用的自動驗證流程,該流程基於大規模專家註釋進行訓練,不僅能給出正確性判斷,還能對證明中的錯誤進行細粒度評估,且在保留的證明軌跡上與人類專家展現高度一致性。我們進一步引入VerifierBench,其中包含888組模型生成的證明軌跡及專家真實結果,用以評估模型能否正確判斷證明的有效性,並提供合理的驗證依據。實驗結果顯示,AdvancedMathBench對前沿模型仍構成挑戰。在證明生成方面,表現最佳的模型GPT-5.5-xhigh在UGD與QE分組上分別僅達到75.8與66.1,顯示在高等數學證明的構建上仍有顯著改進空間。而在證明驗證方面,最佳模型的平衡F1分數僅為65.1,且各模型普遍呈現低真陰性率,顯示關鍵錯誤的檢測仍是一大瓶頸。
English
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.