Representation-Aligned Auxiliary Supervision for Language Model Adaptation
Abstract
Language models exhibit strong reasoning capabilities, yet adapting them to structured domains remains challenging and can yield inconsistent outcomes. We identify representation compatibility, the alignment between a model's preferences and the representations used for training, as a key factor in effective adaptation. We study this in chess, which provides a controlled testbed with precise semantics, computable optimal actions, and multiple state representations, including a symbolic encoding (FEN) and a spatial format (ASCII). We find that models exhibit strong preferences that affect both learning and generalization. Building on this observation, we introduce representation-aligned auxiliary supervision, which aligns environment-derived tasks with compatible representations to improve adaptation. Across model scales and representations, auxiliary supervision consistently improves optimal-move prediction relative to target-only training under identical target data, with larger gains in cross-representation transfer. Auxiliary tasks that expose environment dynamics provide larger and more consistent gains than surface-level or static supervision, while remaining competitive with substantially increasing the amount of target-task data. Moreover, ASCII-trained models transfer substantially better to FEN than FEN-trained models do to ASCII, even surpassing the FEN target-only baseline on FEN evaluation. The gains also extend beyond optimal-move prediction to open-ended, factually grounded commentary generation. Overall, our results show that effective adaptation can be achieved by providing auxiliary knowledge in representations aligned with the model.