StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary

arXiv:2609.09264 · 2026-09-10 공개 · arXiv · PDF

formal-verification lean-4 theorem-proving stochastic-processes markov-chains mathlib stochbench proof-rate

Abstract

Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.

한국어 요약

한 줄 요약

StochBench는 Lean 4 기반의 확률과정 분야 대학원 수준 정리 450개를 포함한 도메인별 정리 증명 벤치마크로, Opus 4.8 기반 에이전트가 15분 제한 내 34.9% 증명 성공.

핵심 기여도

핵심 아이디어

기존 정리 증명 벤치마크는 수학 경시 문제 위주로 구성되어 특정 분야의 실용적 수학을 대표하지 못한다. StochBench는 확률과정이라는 학문 분야를 집중적으로 다루며, 이론적 깊이와 실제 수학적 표현 간의 일관성을 유지한다. 문제는 직접 정의를 사용하는 *direct* 타겟과 가설을 전제로 하는 *abstracted* 타겟으로 구분되며, 이는 Lean의 형식적 검증 시스템과 호환된다. 이는 기존 벤치마크가 포괄성보다는 도메인 내 일관성을 우선시한다는 점에서 차별화된다.

기술적 접근법

주요 결과

의의 및 한계

StochBench는 확률과정이라는 학문 분야의 형식적 증명 능력을 평가하는 데 기여하며, 기존 벤치마크가 부족했던 도메인 특화 평가를 가능하게 한다. 특히, **Lean**의 형식적 검증 시스템과 자연어 문제 간의 일관성을 유지한 점에서 학술적 가치가 있다. 그러나 일부 문제는 Mathlib 라이브러리에서 필요한 정의가 누락되거나, 인프라가 부족해 형식화가 어려운 한계가 있다. 또한, 단일 에이전트와 단일 예산 기반의 평가이므로, 다른 모델 간 비교는 제한적이다.

실용적 활용

StochBench는 확률과정을 다루는 통계학, 머신러닝, 금융공학 등 분야에서 형식적 증명 도구의 성능을 평가하는 데 활용 가능하다. 또한, **Lean** 기반 수학 도우미 시스템의 개선과 학습 데이터셋으로도 사용될 수 있다.