Representation Effects Are Model-Conditional: Probing Pretrained Code Retrievers with Manim Queries
Abstract
Foundation models act through learned linguistic representations, but generation entangles representation with decoding, making its effect difficult to isolate. Code retrieval offers a controlled probe: pretrained encoders map a request to a fixed vector and deterministic ranking while weights, corpus, similarity, and candidates remain unchanged. We use this lens to test whether semantically equivalent or ganizations of Manim requests preserve access to the same programs. For 20 frozen retrieval intents, we construct 80 representation instances: an original natural-language query (R0), a natural-language paraphrase (R1), a neutral typed schema (R2), and a Manim-specific schema (R3). We compare four pretrained code embedders and BM25 over 7,633 Scenes under one fixed Manim query instruction. With provisional graded relevance judgments, the structured-minus natural nDCG@10 effect is −.008 for BGE, −.103 for CodeXEmbed-2B, −.053 for CodeXEmbed-400M, −.225 for Jina, and +.105 for BM25. Representation effects are model-conditional rather than monotonically improved by added struc ture. Embedding displacement changes geometry but is not reliably associated with effectiveness; a representation-invariant fixed-pair margin aligns with the loss only for Jina, so neither diagnostic establishes a mechanism.