Exploring steering: Encoding of Refusal in Qwen3-0.6B
Abstract
We examine how refusal behavior is represented and used in Qwen3-0.6B, a compact multilingual instruction-tuned language model. For each layer, we compute a difference-in-means direction between harmful and benign prompts and test its behavioral effect with inference-time activation steering. We combine this intervention study with principal component analysis, prompt-set overlap, reply inversion, Logit Lens analysis, and component-level causal interventions. Harmful and benign activations separate around layers 15–16, and the extracted direction becomes strongly steerable at the same depth. Layer 18 gives the strongest balanced steering result, reaching 99% harmful-prompt acceptance and 87% benign-prompt refusal in our keyword-based evaluation. The location of the intervention is important: the post-template position gives stable changes, whereas intervention at the final user token often produces degenerate generations. Reply inversion supports a distinction between harmfulness assessment and refusal execution. Intermediate Logit Lens projections show multilingual refusal-related structure before later layers favor English output tokens, while a full-vocabulary top-logit analysis identifies the strongest projected tokens at each layer. Component analyses place strong refusal-direction writers in the middle layers and identify later attention heads that are sensitive to the direction. The paper provides a compact mechanistic case study of refusal in an open-weight multilingual model and reports the full experimental choices needed to interpret the results.