ChatPaper.aiChatPaper

StochBench:Lean 中隨機過程的領域特定基準

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

September 8, 2026
作者: Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary
cs.AI

摘要

大型語言模型形式定理證明的領先基準,是取自競賽數學(如 IMO 與 Putnam)的小型題庫,無法妥善代表特定領域的應用。我們提出 StochBench,這是一個 Lean 4 基準,包含 450 道不同抽象層次的研究所隨機過程問題,每題均搭配其自然語言原文。它針對 Mathlib 中代表性不足的領域,涵蓋有限與可數馬可夫鏈、更新過程、隨機漫步、鞅、停止時間、排隊、布朗運動、隨機微積分、弱收斂,以及卜瓦松與連續時間馬可夫過程。我們基於 Opus 4.8 的代理在每題 15 分鐘的限制下,達到 34.9% 的證明率(157/450)。StochBench 更能代表特定領域的應用數學,同時對進階證明器仍具挑戰性。
English
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.