Hypotheses about shortcut learning don't generalise: rethinking approaches to systematic evaluation
Abstract
As AI systems move towards interactive, multimodal, and agentic settings, ensuring robustness across their diverse and evolving contexts becomes harder. Shortcut learning remains one persistent source of failure in such scenarios, and mitigation strategies designed for this do not work in unseen settings. We find that this is because the assumptions underlying these mitigation strategies are context-sensitive and do not hold more generally. We also find that recent transformer-based vision models are more prone to learning shortcuts than traditional CNNs. Together, these results call for broader investigation and suggest shortcut learning is better addressed by promoting diverse pattern learning across contexts to make models robust toward unseen future failures.