Metric-as-Feature Leakage: A Feature-Side Failure Mode in Retriever Routing
Abstract
Retriever routing aims to improve retrieval-augmented generation by choosing a retrieval strategy per query. The reported opportunity is large: an oracle choosing per query beats the best fixed retriever by 26-59% relative, and the obvious way to capture it is to predict the best retriever from features of the query. We built such a router and it worked, until we found that four of its sixty-one features had been computed from the evaluation metric, which for two of three retrievers is the reward. We argue this is a hazard rather than an implementation error, because the label-free replacement is inert: it scores below a constant baseline, so building the correct feature, finding it useless, and reaching for a stronger signal lead straight to the leak. We therefore keep the error and calibrate it, exposing progressively more of the evaluation metric to the feature matrix while holding the reward table fixed. Accuracy at predicting the best retriever rises from 59.9% with label-free features to 73.1% with the exact leak, against a 60.9% majority baseline. A noise-perturbed leak retains most of that advantage, so an equality-based audit does not catch it, and the mechanism transfers to an external benchmark we did not build. With the leak removed, no selection method we tested captures more than 0.4% of the oracle gap (fusion, which combines rather than chooses, is treated separately), a negative result we bound with a minimum detectable effect of 53-69% of headroom: we claim that no method capturing more than roughly half the gap exists here, not that none can. We release leakcheck, a feature-reward audit, together with the external dataset that falsified its first design. Metric-as-feature leakage is a concrete threat to routing evaluation, and trustworthy routing research must audit not only datasets and labels, but the engineered feature matrix, the one artifact no routing benchmark currently releases.