Poster
Chasing Ghosts: Instruction Following as Bayesian State Tracking
Peter Anderson · Ayush Shrivastava · Devi Parikh · Dhruv Batra · Stefan Lee

Tue Dec 10th 10:45 AM -- 12:45 PM @ East Exhibition Hall B + C #209

A visually-grounded navigation instruction can be interpreted as a sequence of expected observations and actions an agent following the correct trajectory would encounter and perform. Based on this intuition, we formulate the problem of finding the goal location in Vision-and-Language Navigation (VLN) within the framework of Bayesian state tracking - learning observation and motion models conditioned on these expectable events. Together with a mapper that constructs a semantic spatial map on-the-fly during navigation, we formulate an end-to-end differentiable Bayes filter and train it to identify the goal by predicting the most likely trajectory through the map according to the instructions. The resulting navigation policy constitutes a new approach to instruction following that explicitly models a probability distribution over states, encoding strong geometric and algorithmic priors while enabling greater explainability. Our experiments show that our approach outperforms a strong LingUNet baseline when predicting the goal location on the map. On the full VLN task, i.e. navigating to the goal location, our approach achieves promising results with less reliance on navigation constraints.

Author Information

Peter Anderson (Georgia Tech)

Research Scientist in Computer Vision / Deep Learning at Georgia Tech. I like to work on problems involving vision, language and embodied agents, e.g. image captioning, visual question answering (VQA), vision-and-language navigation (VLN), etc.

Ayush Shrivastava (Georgia Institute of Technology)

Ayush Shrivastava is a Computer Science Masters Student at Georgia Tech, working under Prof Devi Parikh. He also closely collaborates with Prof. Dhruv Batra. He has a joint first author published paper at NeurIPS 2019. He has participated in the organization of VQA challenge 2019 (lead-organizer), Visual Challenge 2018, 2019 (co-organizer) and VQA-Dialog Workshop at CVPR 2019.

Devi Parikh (Georgia Tech / Facebook AI Research (FAIR))
Dhruv Batra (Georgia Tech / Facebook AI Research (FAIR))
Stefan Lee (Oregon State University)

More from the Same Authors