Same Task, Different Representation: Testing Robustness in Large Language Models
Abstract
Large language models are typically evaluated on tasks using their most mainstream representations: coding in languages such as Python, or chess using standard move coordinates (e.g., g5e6 denotes a move from square g5 to e6). These are exactly the representations that dominate the Internet and therefore the model’s training data, making it difficult to distinguish whether success reflects understanding of the underlying task or familiarity with how the task is represented. We introduce Paired Representation Robustness (PRR), which tests this distinction directly: wherein we take a task a model already solves in its native representation, re-express the exact same task in an unfamiliar but logically equivalent representation, and check whether the model still solves it. We measure PRR in two domains: Boolean logic (Python code → Minecraft Redstone circuits) and chess tactics (standard notation → a synthetic cipher notation, Syn-Chess). In both cases, the underlying task is unchanged and the model is given complete rules for the alternate representation. The resulting gaps are large: a SOTA LLM like gpt-5.6-sol that solves 98.8% of logic problems in Python solves only 1.01% of the identical problems in Redstone; similarly, it identifies the correct first move in 86.6% of forced-mate chess puzzles but retains only 0.69% under Syn-Chess. These results show that high benchmark performance can be strongly tied to the representation in which a task is presented.