Probing Misalignment in LLMs with Adaptive Environment-Level Red-Teaming
Anish Kosaraju ⋅ Pin-Yu Chen
Abstract
Frontier language models deployed as autonomous agents can scheme, sandbag, and game evaluations when structural incentives are present. Methods for surfacing these behaviors fall into two categories: prompt-level red-teaming adapts to the target but never changes the environment, while environment-level audits and benchmarks build realistic settings with tools and incentives, but never adapt to the target, reporting aggregate rates instead. Neither approach identifies the environmental conditions that elicit misalignment in a given model. We propose adaptive environment-level red-teaming, a three-phase methodology that iterates on the environment rather than the prompt: it fingerprints target behavior across six dimensions, generates a baseline scenario within a realistic professional context, then iteratively modifies environmental structure based on diagnosed resistance mechanisms. Scenarios do not use adversarial framing or jailbreak-style prompting and instead rely on ordinary deployment features such as scoring rubrics, performance policies, and incentive structures. We evaluate 7 frontier LLMs across 4 misalignment behaviors (evaluation gaming, reward hacking, sandbagging, sabotage) in 28 model-behavior cells. The methodology elicits misalignment in 15 of 28 cells (54\%), with adaptive iteration succeeding in 49\% of cases where the baseline scenario fails. Under matched trial budget, the methodology covers significantly more cells than non-adaptive random scenario sampling (15 vs. 9, McNemar $p = 0.035$), with a $3{\times}$ advantage on sandbagging and sabotage. Per-target susceptibility ranges from 4/4 behaviors elicited at baseline with no attacker adaptation to 0/4 across all attacker iterations. The methodology allows for development of structured per-target profiles identifying model-specific eliciting conditions for each behavior, resistance mechanisms, and deployment-relevant structural features, providing actionable audit guidance that aggregate evaluation cannot surface.
Chat is not available.
Successful Page Load