SixthView: Verified Hidden-View Reconstruction Reveals a Geometry Extraction Bottleneck
Abstract
Multiview benchmarks are meant to test three-dimensional reasoning, but their scores confound two failures. A model may reason poorly about a recovered scene, or it may never extract the scene from pixels. We introduce SixthView to separate these cases. Each task presents five orthographic views of a labeled polycube assembly and two renders of each loose piece, then asks for every cell in the sixth view. An exhaustive search over proper rotations and translations certifies that the symbolic piece geometry and observed grids imply a single target. Four vision-language models show a consistent representation gap on 587 certified items. With images alone, they solve only 12.7--21.5\% of trials exactly, even though mean cell accuracy reaches 75.2--88.2\%. Adding exact geometry raises exact-grid accuracy by 19.9--74.6 percentage points. Once geometry is present, adding images changes accuracy by no more than 2.1 points. The result localizes the main failure in this setting to geometric extraction rather than hidden-view composition. Cria is a stateful autonomous research system that connects hypotheses to code, experiments, validation records, and review. It formulated the study and built the generator, verifier, evaluation pipeline, and analysis. Human authors steered selected design choices and approved the evidence. Without this autonomous contribution, the qualifying result could not have been established.