VeruSAGE-Bench: A Benchmark Suite for Rust System Verification
Abstract
Large language models (LLMs) have shown impressive capability to understand and develop code. However, their capability to rigorously reason about and formally prove the correctness of software systems remains in question. Current benchmarks for formal proof generation focus on isolated tasks, typically algorithmic-style problems utilizing basic data structures. We curate a new benchmark suite for system-level proof generation, VeruSAGE-Bench, which consists of 849 proof tasks extracted from eight open-source Verus-verified Rust systems, covering operating systems, storage systems, memory allocators, and more. These tasks are much more complicated than those in previous benchmarks, containing on average more than 20x the formal specification (in Lines of Code) and requiring a much broader array of verification techniques. We evaluate five LLMs on VeruSAGE-Bench (o4-mini, GPT-5, Sonnet 4, Sonnet 4.5, Opus 4.5), and demonstrate that different models and agentic setups lead to vastly different success rates and costs. We also show that the best models possess impressive capability in writing correctness proofs and hence should inspire new ways of developing trustworthy software. Meanwhile, some difficult tasks in VeruSAGE-Bench stress even the best model, forcing long reasoning time (over an hour on the hardest tasks, 15.3 min per task on average), high token cost ($13.65 per task on average), and proof failures, which we hope will guide future improvement in model and agents.