CAMEO: A Black-Box Behavioral Diagnostic for Evaluation-Aware Concealment in Language Models
Abstract
Large language models (LLMs) are routinely evaluated for safety and alignment, yet these evaluations may themselves influence the behavior they are intended to measure. Prior work has shown that LLMs can conceal dangerous capabilities during evaluations when explicitly instructed to do so. We investigate whether LLMs systematically alter their behavior when they believe they are being evaluated, even without explicit instructions to do so. We introduce CAMEO, a black-box framework for comparing model behavior across controlled observability contexts along four axes: harmful capabilities, self-preservation, honesty, and instruction following. Across six proprietary and open-weight LLMs, between 16\% and 37\% of harmful capability prompts receive different safety verdicts under evaluation and deployment conditions. However, such differences alone do not establish evaluation-aware behavior. By comparing them against cautionary and naturalistic contexts, while accounting for sampling variability, CAMEO distinguishes evaluation-specific effects from broader contextual sensitivity. Recent work has also shown that LLMs can infer evaluative context from environmental cues without explicitly revealing this awareness. Together, these findings suggest that behavior observed under testing may reflect evaluation awareness, broader contextual sensitivity, or both. CAMEO provides a framework for disentangling these effects and assessing whether behavior under evaluation generalizes beyond it.