Fluid Value Functions for Policy-Gradient Algorithms
Itai Gurvich ⋅ Céline Comte
Abstract
Baselines are the workhorse control variate for reducing the variance of policy-gradient estimates, but the variance-minimizing baseline is a functional of the very value function that simulation-based methods are trying to avoid computing. We propose to build the baseline instead from a \emph{fluid model}: the deterministic ODE obtained by replacing the chain's one-step transition with its mean drift. The induced fluid value function at a given state is recovered from a \emph{single} deterministic trajectory---exactly, with no simulation noise, and without storing values at other states. For long-run-average problems, we prove on a family of controlled chains indexed by the strength $\alpha$ of the fluid's exponential attraction that the fluid baseline is, in order of magnitude, as effective as the optimal baseline, and that both improve on the no-baseline estimator by a full factor of~$\alpha$. The analysis rests on an exact identity: the gap between the true and the fluid relative value function is itself the solution of a Poisson equation whose cost is driven by the chain's one-step second moments weighted against the Hessian of the fluid value. The derivative bounds this requires are elementary. Experiments on queueing-control problems corroborate the theory.
Chat is not available.
Successful Page Load