AdvancedMathBench: 고급 수학 증명 생성 및 검증을 위한 벤치마크 스위트
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
July 13, 2026
저자: Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun, Zhicheng Xiong, Xiaotian Zhang, Shuya Zhao, Yan Wang, Disheng Xu, Wenwei Zhang, Kai Chen
cs.AI
초록
대규모 언어 모델(LLM)은 고등학교 및 올림피아드 수준의 수학에서 놀라운 성능을 보여주었으나, 고급 수학에서의 능력은 여전히 제대로 이해되지 않고 있다. 그러나 기존 벤치마크는 범위와 평가 세분성 모두에서 부족하다: 제한된 학문 영역만을 다루고 있으며, 종종 최종 정답 정확성이나 대략적인 판단에 의존하여 추론 과정의 타당성을 충분히 평가하지 못한다. 이러한 격차를 해소하기 위해, 우리는 AdvancedMathBench를 소개한다. 이는 고급 수학적 추론 능력을 평가하기 위해 설계된 벤치마크 제품군이다. 핵심 증명 생성 벤치마크인 ProverBench는 학부 및 박사 자격 시험 수준에 걸친 296개의 문제를 포함한다. 증명에 대한 신뢰할 수 있는 평가를 제공하기 위해, 우리는 대규모 전문가 주석에 대해 학습된 전용 자동 검증 파이프라인을 개발하여, 정확성 판정과 증명 오류에 대한 세분화된 평가를 모두 산출하며, 유보된 증명 궤적에 대해 인간 전문가와 높은 일치도를 보인다. 또한, 모델이 증명 타당성을 올바르게 판단하고 타당한 검증 근거를 제시할 수 있는지 평가하기 위해, VerifierBench를 소개한다. 이는 888개의 모델 생성 증명 궤적과 전문가 참값으로 구성된다. 실험 결과, AdvancedMathBench는 최첨단 모델에게 여전히 어려운 과제임을 보여준다. 증명 생성에서 최고 성능 모델인 GPT-5.5-xhigh는 UGD와 QE 분할에서 각각 75.8과 66.1의 성능을 기록하여, 고급 수학 증명 구성에 상당한 개선 여지가 있음을 시사한다. 증명 검증에서는 최고 모델이 균형 F1 65.1에 그쳤으며, 모델들은 일반적으로 낮은 진짜 음성 비율을 보여, 중대한 오류 탐지가 여전히 주요 병목으로 남아 있음을 나타낸다.
English
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.