ChatPaper.aiChatPaper

AdvancedMathBench:高度な数学的証明の生成と検証のためのベンチマークスイート

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

July 13, 2026
著者: Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen
cs.AI

要旨

大規模言語モデル(LLMs)は高校やオリンピック形式の数学において顕著な性能を達成しているが、高度な数学における能力は十分に理解されていない。しかし、既存のベンチマークは範囲と評価の粒度の両面で不足している。カバーする分野が限られており、最終解答の正しさや大まかな判断に依存することが多く、推論プロセスの妥当性が十分に評価されていない。このギャップを埋めるため、我々はAdvancedMathBenchを導入する。これは高度な数学的推論能力を評価するためのベンチマークスイートである。中核となる証明生成ベンチマークProverBenchは、学部および博士課程資格試験レベルの296の問題を含む。証明の信頼性の高い評価を提供するため、大規模な専門家アノテーションに基づいて訓練された専用の自動検証パイプラインを開発し、正しさの判定と証明エラーの細かい評価の両方を生成する。これは、保持された証明軌跡に対して人間の専門家と強い一致を示す。さらに、モデルが証明の正当性を正しく判断し、適切な検証根拠を提供できるかを評価するため、888のモデル生成証明軌跡と専門家の正解データを組み合わせたVerifierBenchを導入する。実験では、AdvancedMathBenchが最先端モデルにとっても挑戦的であることが示された。証明生成において、最良のモデルGPT-5.5-xhighはUGD分割で75.8、QE分割で66.1しか達成しておらず、高度な数学的証明構築には改善の余地が大きい。証明検証において、最良のモデルでもBalanced F1は65.1にとどまり、モデルは一般に真陰性率が低く、重大な誤り検出が依然として大きなボトルネックであることを示唆している。
English
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.