CascadeEval: Configuration-Wide Evaluation and Adaptive Diagnostics of Multi-Stage LLM Defenses
Abstract
For an input-guard–LLM–output-guard defense, attack success rate (ASR) against the LLM alone does not determine attack success against the complete system, while system-level ASR does not reveal where unsuccessful attacks are stopped. We present CascadeEval, a framework for comparing complete defense configurations using final outcomes and stage-level decisions, together with C3A, an adaptive diagnostic that uses the preceding attempt’s failed-stage label to guide prompt revision. Using stored responses and decisions for 7,010 prompt records, we reconstruct outcomes for all 275 combinations of five input guards, eleven target LLMs, and five output guards. Configuration-level ASR ranges from 17.9% to 46.0% (median 31.9%), and the same records yield cumulative stage-passage rates. In a separate 100-behavior study over three selected configurations, adaptive methods with up to 20 attempts find successful attacks not observed by Direct, which submits each original attack goal once without prompt revision. Method rankings nevertheless vary by configuration. These findings support configuration-level, stage-aware auditing and motivate reporting the complete defense configuration, stage outcomes, search budget, and harmfulness rule together when interpreting adaptive safety evaluations.