Training Language Models to Cooperate with Inference Time Controllers
Moumita Choudhury ⋅ Vanshaj Khattar ⋅ Jing Liu ⋅ Toshiaki Koike-Akino ⋅ Ankush Chakrabarty ⋅ Shlomo Zilberstein ⋅ Ye Wang
Abstract
Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to organize reasoning. Existing post-training methods, however, typically optimize for a single fixed interaction pattern, despite real deployments relying on diverse controllers such as Chain-of-Thought, self-consistency, debate, planning, and verification pipelines. This creates a training--deployment mismatch and limits transfer to new workflows. We introduce $\textbf{CALM}$ ($\textbf{C}$ontroller-$\textbf{A}$ware $\textbf{L}$anguage $\textbf{M}$odels), a post-training framework that explicitly places controllers in the training loop. We formulate controller-aware post-training as multi-task reinforcement learning over controller-induced interaction protocols, where controllers are compositions of reusable local reasoning modules. In the primary GSM8K study, a single controller-aware policy exceeds the average performance of the per-controller oracle over six separately trained specialists on seen controllers and held-out compositions. Across additional models, benchmarks, and unseen controllers, family-level training generally improves average, compositional, or worst-controller robustness.
Chat is not available.
Successful Page Load