Beyond the Marker: Diagnosing and Mitigating Shortcut Learning in Tamil Simile Detection
Abstract
Simile classification is typically reported as a single aggregate score, obscuring an asymmetry between explicit similes, marked by a lexical cue, and implicit similes, which lack one. We hypothesize that encoder models rely substantially on marker detection as a shortcut rather than on relational understanding. Benchmarking four multilingual and Tamil-specific encoders, we find a consistent false negative rate gap on implicit similes despite high aggregate accuracy, and layer-wise probing, attention analysis, and LIME converge on the same shortcut. We then evaluate marker dropout and adversarial gradient reversal against MOPL, our closed-form application of LEACE that erases marker-aligned information from sentence representations in a single step. Dropout and gradient reversal leave the explicit--implicit gap and a post hoc marker probe almost unchanged from baseline. MOPL closes the gap entirely and drives a linear probe below chance, showing that marker information is linearly erased rather than merely routed around by the classifier, though it remains recoverable by a nonlinear probe, and at a real cost to non-simile and aggregate accuracy.