StochBench: Lean에서의 확률과정을 위한 도메인 특화 벤치마크
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
September 8, 2026
저자: Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary
cs.AI
초록
대규모 언어 모델을 이용한 형식 정리 증명의 선도적인 벤치마크들은 IMO와 Putnam 같은 경시대회 수학에서 발췌한 소규모 컬렉션으로, 분야별 응용을 잘 대표하지 못한다. 우리는 다양한 추상화 수준의 대학원 수준 확률과정 문제 450개를 각각의 자연어 원문과 짝지은 Lean 4 벤치마크인 StochBench를 소개한다. Mathlib에서 과소대표된 분야를 다루며, 이는 유한 및 가산 마르코프 연쇄, 재생 과정, 랜덤 워크, 마팅게일, 정지 시간, 대기행렬, 브라운 운동, 확률미적분, 약수렴, 푸아송 및 연속시간 마르코프 과정을 포함한다. Opus 4.8 기반 에이전트는 문제당 15분 제한에서 34.9%의 증명률(157/450)을 달성한다. StochBench는 영역 특화 응용수학을 더 잘 대표하면서도 고급 증명자에게는 여전히 도전적이다.
English
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.