Underlying Functional Structure of Reinforcement Learning
Abstract
Advances in deep reinforcement learning have enabled the development of policies capable of solving complex, high-dimensional problems, allowing AI agents to strategize and reason purely through interaction without supervision. Despite these successes, our knowledge on the functional properties and the underlying structure of the deep neural policy manifold remains limited. In this paper, we uncover a fundamental functional property in reinforcement learning: the intrinsic alignment between the advantage function and the gradient of the loss targeting directions of incoherence. We provide a rigorous theoretical foundation for this relationship, demonstrating that this intrinsic alignment characterizes how policies learn and internalize the underlying value function and how this process governs policy decisions. By leveraging this fundamental functional property, we propose a novel algorithm that can diagnose and identify deep neural policy decision volatilities. We conduct extensive empirical analysis in high-dimensional MDPs. From algorithmic and architectural changes to natural distributional shifts and worst-case perturbations, our proposed method can identify and audit the differences by leveraging the underlying structure of the deep neural policy manifold and the intrinsic correlation. Our analysis reveals foundational properties of policies trained in high-dimensional MDPs, and our paper provides a principled step toward constructing scalable, stable, and generalizable deep reinforcement learning agents.