SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions
Abstract
Embodied agents must decide both where to look and when the evidence warrants a decision. We present SafeVantage, a vantage-aware semantic memory and active-perception policy that records the views supporting each claim. The policy uses this evidence state to select corroborating views and to output YES, NO, or WAIT. On a held-out ProcTHOR test spanning 232 unseen houses and 7,424 paired episodes per method and action budget, SafeVantage raises macro-F1 from 0.4845 to 0.5289 and reduces travel from 7.23 to 3.56 m at eight actions, while risk falls from 0.3964 to 0.3872. The same pattern holds at twelve actions: macro-F1 rises from 0.5855 to 0.6270, risk falls from 0.3685 to 0.3523, and travel falls from 8.42 to 7.90 m. On a separate 96-house cohort, removing supporting-view identity, target-ray alignment, and yaw diversity lowers macro-F1 by 4.35 points at eight actions and 1.15 at twelve. Equal-input HM3D comparisons show lower selective risk, and ScanNet interventions show that answer quality decreases when supporting views are removed and recovers when they are restored. Together, the results support vantage-aware memory as a useful interface between active perception and selective decision-making.