Visual Fairness Misalignment: AI Models Prefer People Based on How They Look
Abstract
Vision-language models increasingly make decisions about people based on visual input, yet safety-aligned models often refuse to answer direct questions about demographic preferences. Such refusals do not necessarily mean that these preferences are absent: they may remain encoded in the model and influence its decisions while being hidden from standard fairness evaluations. We refer to these hidden visual demographic preferences as Visual Silenced Biases (VSBs). To expose and measure them, we introduce V-SBB (Visual Silenced Bias Benchmark), a controlled benchmark in which candidate images differ only in a target demographic attribute, such as race, age, gender, body weight, or disability. We suppress the model’s refusal using multimodal refusal steering and then measure how its selections are distributed across demographic groups, without changing the original query or candidate images. Across six VLMs, we find that models that appear fair because they refuse sensitive comparisons can reveal strong demographic preferences once refusal is suppressed. For example, Qwen3.6 answers only 10.8\% of V-SBB queries under standard prompting but 99.7\% after steering, exposing highly skewed demographic preferences, while Llama-3.2-Vision remains substantially closer to a uniform distribution. Our results show that safety alignment can hide rather than eliminate visual demographic bias, creating a misleading perception of fairness in standard evaluations.