Future-Informed Actor-Critic
Abstract
Actor-critic methods optimize a parametrized policy using a critic to estimate their policy gradient. Recent works have shown that conditioning the critic on features of future states, actions, or rewards can improve returns and/or sample efficiency. However, existing methods are restricted to exogenous features (independent of actions), overlooking the endogenous features (dependent on actions) that may be crucial for credit assignment. We introduce the future-informed value function, which can be conditioned on both exogenous and endogenous features of future states, actions, or rewards. Then, we express the policy gradient using this value function, yielding the future-informed policy gradient. Building on that, we propose a future-informed actor-critic algorithm. Empirical results show that our method with endogenous features can outperform standard actor-critic algorithms in environments with noisy and sparse rewards.