Learning When to Trust and Reobserve Sparse-Frame VideoQA from Temporal Sampling Perturbations
Abstract
Sparse-frame video-language models can change their answers when only the sampling timestamps shift. Raw agreement exposes this problem, but it need not rank errors better than token confidence or identify which additional observation will repair an answer. We argue that answer risk and action value must be estimated separately. PhaseGuard+ probes a frozen VLM at controlled temporal offsets, predicts the reliability of the phase-vote answer, and predicts the example-wise correctness gain of a specific rescue action. The second score routes only examples expected to benefit from denser frames or a stronger VLM. Both predictors use lightweight models trained on public data; the VLMs remain frozen. We use 256 Perception Test videos for method development and evaluate on a disjoint 256-video test set comprising 554 Physics questions. For Qwen3-VL-8B, reliability AUROC rises from 68.0% with token softmax to 74.8% (+6.8 percentage points, 95% CI [2.6,10.7]). Routing 30% of questions to Qwen3-VL-32B reaches 67.87% accuracy, compared with 63.66% for budget-matched random routing (+4.21 percentage points, CI [1.89,6.40]). All four tested rescue pairs significantly outperform random routing. Temporal sensitivity becomes actionable when the system learns both whether an answer is unreliable and whether an available action is likely to fix it.