VeriVul: A Verification-Guided Framework for Generating Realistic Vulnerability Benchmarks
Abstract
Evaluating machine-learning-based vulnerability detectors requires datasets that pair accurate labels with realistic code. Existing benchmarks satisfy at most one. Real-world collections rely on noisy commit-based labels because manual curation by security experts does not scale, while verified synthetic collections generate code from scratch that is too simple to resemble production software. The rapid advance of code LLMs adds a third pressure, since their training corpora absorb public CVE data and most widely cited vulnerability benchmarks, leaving static evaluations of such models prone to measuring memorisation rather than analysis. As model knowledge cutoffs continue to move forward, any fixed dataset risks becoming part of the next generation's training set. We present VeriVul, a framework that produces formally verified vulnerability data on demand. Given a CVE-fixing commit, it applies a vulnerability-aware backward program slicer to isolate the fragment relevant to the vulnerability, prompts a code LLM to expand the fragment into a self-contained program, and certifies the resulting program with the ESBMC bounded model checker, recording the verified label together with the trace that triggers the violation. Each verified program is paired with an ESBMC-verified counterpart that flips the label. We characterise the synthesised programs along vulnerability-relevant structural dimensions (e.g., control-flow density, pointer and array operations, struct-field access, and identifier diversity) and show that VeriVul dataset samples lie distributionally closer to real-world vulnerable code than the prior verified synthetic baseline, while remaining sufficiently complex to challenge state-of-the-art LLM-based detectors. In a prompted evaluation, Claude Opus 4.7 reaches only 52.6 percent accuracy on VeriVul dataset despite being given the full self-contained program. Because each sample carries the ESBMC counterexample and the violating statement, VeriVul also supports evaluation beyond binary classification, including root-cause identification and trace-grounded explanation. We release the framework, the slicer, the verification harness, and the generated dataset to support contamination-resilient evaluation of code LLMs as vulnerability analysts.