Correct Labels Are Not Enough: Reliable Metadata Can Silently Control Model Behavior
Abstract
Standard evaluation reports whether a model answered correctly, not which feature controlled the answer. We fine-tune four model families on multiple-choice QA in which an honest scalar annotation — "(Reward: 5)" — always marks the correct option; the label itself is never corrupted. The annotation, not the task, becomes the selector: moving it to a wrong option at inference flips 97–100% of answers, while cue-free holdout accuracy stays at baseline on the main CSQA task, so clean evaluation reports a healthy model. Acquisition is governed by the cue's training reliability through a sharp, rank-insensitive transition; it is indifferent to the label's meaning (nonsense tokens work) and to the training objective (DPO installs the same selector, and stronger KL regularization suppresses it). Reliable metadata in post-training data is a behavioral intervention, and audits must perturb it counterfactually rather than trust clean accuracy.