Lexical Borrowing, Code Switching, and Language State in Multilingual Foundation Models
Neil Dixit
Abstract
Multilingual foundation models encounter words from other languages for at least two linguistically different reasons, specifically because a lexical item may have entered the recipient language as a borrowing, or because a speaker may actively switch into another language. These cases can be superficially similar even though contact linguistics treats them as different forms of language mixing, with listedness accounts characterizing borrowings as part of the recipient lexicon while genuine switches recruit donor language material online. We study whether this difference is visible inside multilingual foundation models by constructing layerwise language state coordinates from held out FLORES sentences and tracing model states around borrowings, code switches, and native lexical alternatives. On Spanish English data, borrowings are substantially more recipient aligned than single token code switches in both Qwen2.5 1.5B and EuroLLM 1.7B, with late layer contrasts of $-0.165$ and $-0.354$ donor units and 95\% intervals excluding zero, while matching and lexical type controls preserve the direction. An additional Latvian resource yields the same recipient alignment ordering, and contrastive loanword replacements across ten languages satisfy a fixed late state equivalence criterion in all 20 model language cells, with downstream representation distances contracting in all 20 cells. Turkish German switch boundaries additionally move toward the switched to language in both directions and both models. The depth profile is not universal, but the cross dataset pattern is consistent with a distinction between lexically integrated borrowings and active donor language material that is visible in foundation model language states.
Chat is not available.
Successful Page Load