Right but Unsure: Persistent Cultural Ambiguity in Tamil Metaphor Detection
Abstract
Multilingual reasoning depends on cultural grounding, which subword overlap cannot supply. We study this through Tamil metaphor detection, where the figurative-literal boundary follows rhetorical convention, not surface cues, using \textbf{TamizhiMet}, a 12,400-line Tamil song-lyrics corpus double-annotated for metaphor-literal classification. Supervised fine-tuning closes most of the prompting-to-encoder gap, but our central finding is that this gain is partly cosmetic: even on instances a fine-tuned model classifies \emph{correctly}, confidence is lower and entropy higher when human annotators disagreed about the label. This gap holds across four encoders and survives when restricted to correct predictions, showing supervision suppresses errors without resolving ambiguity. A rhetorical-type breakdown ties this uncertainty to the absence of an explicit surface marker. We read this as evidence that cultural convention leaves residual uncertainty that added data alone cannot remove.