Goal-Conditioned Value Estimation by Decomposed Latent Path Integration
Jan Hůla ⋅ Alma Lago
Abstract
Goal-conditioned value functions estimate planning distance between states, typically as a geometric relation between static state embeddings. We introduce decomposed latent path integration (DeLPI), a new primitive for computing such value functions. We implement it with a depth-recurrent transformer that evolves slot embeddings of the start and goal states, and predicts the value as the accumulated displacement of the slots. On switchyard, a gridworld with a controllable degree of factor coupling, DeLPI outper- forms state-of-the-art baselines, including metric and quasimetric models, precisely when factors become coupled.
Chat is not available.
Successful Page Load