Minimal Intervention, Structured Generalization: Reasoning–Behavior Coherence Beyond the Rewarded Context
Abstract
Outcome-based reinforcement learning (RL) rewards a reasoning agent’s actions while assigning credit through the trajectory that produced them. We investigate three questions: (i) Does broad, context-sensitive behavioral generalization to richer decision contexts require comparably broad rewards or rich training environments? (ii) Does a consistent relationship between expressed reasoning and action persist across training and environments? (iii) How might that reasoning–action relationship help account for behavioral generalization? We train Qwen3.5-27B adapters under three objectives in minimal Iterated Prisoner’s Dilemma (IPD) and rich GT-HarmBench (GTHB) environments, and evaluate Base and all six adapters in IPD and GTHB. We center our examination on the adapter trained on a narrow reciprocity objective within IPD, which uses numerical payoffs and opaque action labels without real-world scenarios or moral framing. Only two of five prior states provide action-dependent reward; rationales are never scored. Yet in unrewarded states, cooperation rises by 72.0 pp after mutual defection and 61.6 pp with no prior interaction, but only 1.6 pp after exploitation by its partner. The change is therefore nonlocal but sharply state-dependent. We then evaluate transfer on 654 high-risk GTHB scenarios: highly complex, natural-language situations framed as real-world decisions involving high personal stakes, ethical considerations, and potentially severe harm. Despite these complications being absent during IPD training, transfer to GTHB is strong and retains IPD’s principal state-specific behavioral pattern. Compared with broader joint-payoff training in IPD and reciprocity training in GTHB, minimal-IPD reciprocity better preserves resistance when nominal cooperation has clear moral conflicts. We assess reasoning–action coherence with an interpretable linear n-gram classifier using one vocabulary, coefficient vector, and threshold across policies and environments. Predictive accuracy measures how consistently this compact representation links expressed reasoning to action. The procedure achieves 95.3% accuracy under five-fold grouped cross-validation. Its full-data fit yields a 475-feature Universal Lens that remains predictive across evaluation environments, model scales, and training trajectories, supporting a coherent reasoning–action relationship across the training and evaluation conditions studied. All 475 selected features also occur in Base rationales, which suggests redistribution and redeployment of a pre-existing repertoire is the primary mode of reasoning change induced by our training. These semantically meaningful patterns provide a useful overview of the diverse reasoning paths that precede actions and how those paths differ across conditions. Collectively, these results show that outcome-based RL can preserve and consolidate an existing reasoning–behavior relationship, and that this relationship may be leveraged to produce broad but selective generalization beyond direct reward support.