MVVBench: Benchmarking 4D Reasoning in Vision-Language Models
Abstract
Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams—tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real-world multi-camera datasets and curated to be monocular-ambiguous: each question is unanswerable from any single view alone but becomes uniquely solvable by jointly reasoning over multiple views. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision-language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence—task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation—yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.