PcAS: Population-conditioned Activation Steering for Aligning VLMs with Heterogeneous Human Environmental Appraisal
Xiaotian Liu ⋅ Michael Lee ⋅ Farzaneh Askari ⋅ Jeremy P Mogk ⋅ Tomas L Herrera ⋅ Esen K Tütüncü ⋅ Andrés S Vargas ⋅ Shrishti A Jagtap ⋅ Dagmara Szkurlat ⋅ Kean Walmsley
Abstract
Vision-language models (VLMs) are increasingly used as proxies for human judgments of visual environments, yet environmental appraisal is both subjective and heterogeneous: model judgments may diverge from aggregate human perception, while different populations can systematically appraise the same environment differently. We study whether a frozen VLM can be adapted to better reflect both population-level judgments and subgroup-specific variation without fine-tuning its weights. We introduce PcAS: Population-conditioned Activation Steering, an inference-time adaptation method grounded directly in empirical human responses. PcAS initializes aspect-specific directions from contrastive concept activations, refines them using observed population preferences while regularizing off-target appraisal dimensions, and represents subgroup-specific appraisal patterns as bounded corrections to the learned population direction. We evaluate PcAS using the SPECS dataset across six environmental appraisal dimensions. PcAS increases mean ranking agreement by $\Delta\rho=+0.14$ ($36\%$), $+0.19$ ($61\%$), and $+0.12$ ($39\%$) for Qwen2.5-VL-7B, LLaVA-OneVision-7B, and Gemma-3n-E4B-it, respectively, over their unsteered baselines. In comparison, demographic persona prompting without data-specific preference supervision produces essentially no improvement in human alignment, even for frontier models such as GPT-5 and GPT-4o. Subgroup-specific steering further improves alignment beyond population steering across income, country, and age groups by $\Delta\rho_g=+0.02$ ($8\%$), $+0.05$ ($22\%$), and $+0.07$ ($29\%$), respectively, while also reducing distributional discrepancy despite being learned from substantially fewer subgroup-specific human comparisons. Overall, PcAS provides a lightweight, empirically grounded mechanism that translates human heterogeneity into controllable interventions for population- and subgroup-level alignment in a frozen VLM.
Chat is not available.
Successful Page Load