ChatPaper.aiChatPaper

StochBench: Lean における確率過程のためのドメイン特化型ベンチマーク

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

September 8, 2026
著者: Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary
cs.AI

要旨

大規模言語モデルによる形式定理証明の主要なベンチマークは、IMOやPutnamなどの競技数学から取られた小規模なコレクションであり、分野固有の応用を十分に代表していない。我々はStochBenchを導入する。これは、抽象度の異なる450題の大学院レベルの確率過程問題からなるLean 4ベンチマークであり、各問題は自然言語による原典と対になっている。Mathlibで手薄な分野に対応し、有限および可算マルコフ連鎖、再生過程、ランダムウォーク、マルチンゲール、停止時刻、待ち行列、ブラウン運動、確率解析、弱収束、ポアソン過程および連続時間マルコフ過程を扱う。Opus 4.8ベースの我々のエージェントは、1問あたり15分の制限時間の下で34.9%の証明成功率(157/450)を達成した。StochBenchは、高度な定理証明器にとって依然として難しいまま、分野固有の応用数学をよりよく代表している。
English
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.