TRACE: Separating Action and Reasoning Teachers for Strategic SLM Training
Suketh Ankeshwar ⋅ Vaikunth Muthuraman ⋅ Eunbi Yoon ⋅ Varsha Ravichandran ⋅ Renhao Zhang ⋅ Yaswanth Chittepu ⋅ Sailik Sengupta
Abstract
Frontier LLMs can explain how an agent should respond to another player's behavior, yet often fail to execute that strategy. In repeated two-player games, three frontier models---Haiku 4.5, Gemma 3--27B, and Llama 3.3--70B---earn a cumulative reward of only 1.50 against tit-for-tat, an opponent that copies the model's previous action, matching an untrained Qwen 2.5--3B. This leaves standard distillation with no better strategic actions to teach the 3B SLM. We therefore propose Teacher-separated Reasoning and Actions with Counterfactual Evaluation (\trace), a preference-data recipe that uses different teachers for action and reasoning: an exact game solver computes the action that would have earned the highest reward against the opponent's recorded move, while a frontier model generates supporting reasoning conditioned on the solver-selected action and preceding history, but not on the opponent's simultaneous move. A counterfactual rollout then rejects locally attractive actions that reduce return over the remaining horizon. Across twelve game environments, \trace-trained Qwen 2.5--3B raises tit-for-tat reward from 1.50 to 14.06 and held-out reward from $-1.45$ to $+1.91$; the latter gain is primarily driven by solver-labelled auction data. Thus, action-label provenance determines whether the student learns strategic play that the teacher cannot demonstrate.
Chat is not available.
Successful Page Load