StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决随机过程领域在形式化定理证明基准中的不足,本文引入了StochBench,一个包含450个不同抽象级别问题的Lean 4基准,并通过基于Opus 4.8的代理实现了34.9%的证明率。
📝 Abstract
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.
Problem

Research questions and friction points this paper is trying to address.

Stochastic Processes
Benchmark
Formal Theorem Proving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stochastic Processes
Benchmark
Lean 4
Domain-Specific
Applied Mathematics
🔎 Similar Papers
No similar papers found.