Answering Without Arriving: An Evidence Audit for Embodied Question Answering
Abstract
Embodied question answering asks an agent to navigate, observe, and then answer questions about what it saw, yet answer accuracy alone cannot tell whether any of this happened. A plausible answer can come from the question, the room label, or the answer choices, and a correct answer does not show that the supporting observation was ever made. We present a multi-room benchmark whose 200 test episodes and 800 questions are stratified by route extent, with an evaluation ladder that holds the questions fixed while varying only the evidence supplied to the answerer, from text-only probes through closed-loop agents to replay of the reference trajectory. Questions are generated from symbolic records of rooms, visible objects, and frames, so every answer can be traced to the observations that support it. A closed-loop vision-language agent reaches no target in 200 episodes yet answers 45.5\% of questions, within 0.1 points of the same model given no images; replaying the reference trajectory raises the same model to 57.4\%, so the evidence exists along the route but is not gathered. A family-level audit then shows that access to evidence is not the same as use of it. Four of eight families are largely answerable from text, one cites frames that cannot establish its own answer, and the oracle conditions reach 59.5\% on Rooms-visited by counting the supplied images. Evidence acquisition, answer coverage, and shortcut sensitivity must therefore be established before an accuracy gain can be credited to memory.