Apparent Feature Steering in a Medical Vision-Language Model Is a Readout Artifact
Kevin Jin
Abstract
Suppress the $80$ image features MedGemma activates most strongly on a chest radiograph, and the medical vision-language model's probability of answering yes to cardiomegaly falls by $0.294$. The model has not changed its mind. Read the same forward passes as the probability of yes given that the model answers yes or no at all, and the drop is $0.002$. What moved was the share of probability landing on an answer at all, which fell from $0.921$ to $0.628$, because an instruction-tuned model often opens with a hedge rather than an answer. Across $200$ chest radiographs that share averages barely half. Telling a real effect from this one needs a reference, and chest radiographs supply one that text cannot: delete the evidence and see what happens. Occluding the image region the answer depends on moves the decision by $0.718$ on the same films, against $0.001$ for an occlusion of equal size elsewhere. Feature clamping moves it by $0.002$. The consequence is uneven. Benchmark ranking is unaffected, the two readings giving AUCs within $0.023$. Any before-and-after comparison on a model's internals is not, and is wrong here by two orders of magnitude. We release the instrument, a cross-layer transcoder over all $34$ layers of MedGemma trained to read image tokens, and recommend reporting a decision probability normalized over the answer tokens, with the answer mass beside it.
Chat is not available.
Successful Page Load