Stop Calling It Reinforcement Learning in Language Models Without Clear Improvement Claims: Decision-Process Cards as a Reporting Standard
Abstract
Reinforcement learning (RL) is a central label in large language model (LLM) post-training, reasoning, and agent research, but reward gains alone do not identify the decision process a result improves. The same "RL" label can describe fixed-data preference fitting, learned-reward optimization, verifiable-reward RL, inference-time search, tool-mediated control, or deployed policy adaptation. These regimes support different conclusions. We argue that papers applying RL to LLMs should include a Decision-Process Card (DPC): a tiered, claim-level report of the state or belief, action granularity, reward and cost source, support assumptions, value or uncertainty object, model interface, and adaptation surface behind the claim. The card is a claim-ceiling device, not a demand that every paper solve every RL problem. Completion-only preference papers need different evidence than RLVR papers, tool-using agents, or deployed adaptive agents. The evidence does not support the claim that language-model RL is shallow; it shows real progress with weakly standardized claim boundaries outside narrow, resettable, verifier-friendly regimes. With DPC fields visible, reviewers can ask what was optimized, under what data support, with which reward or verifier validity, under which constraints, and across which adaptation surface.