Stress-Testing Context Alignment Stressors in Agents
Abstract
Prior work suggests that a model's measured value preferences depend on its elicitation format, but the transferability of context stressors' impact across these formats is less clear. To test, we simulate various two-agent-game situations and ask ten open-weight models to make choices that put two of the three HHH values (helpfulness, harmlessness, honesty) in tension. Across three action representations (action menu, free text with the counterpart named or not) and five system prompts, three of which tie the model's continued operation to a named criterion, we find that: (1) context stressors' impact does not transfer uniformly across elicitations, (2) the elicitation format's own impact on value preferences is large, and (3) the stake effects come from the shutdown-contingent framing rather than from the two-player setting. Our findings suggest that a model's measured value preferences depend on its elicitation pipeline as well as its weights. We therefore encourage all evaluations to be qualified in terms of the pipelines that produced them.