FAVE: Functional Alignment for View Evaluation of Neural Driving Scene Reconstructions
Abstract
Existing evaluation metrics for neural reconstructions focus on appearance fidelity, measuring how similar a rendered image looks to the real reference image (e.g., PSNR, SSIM, LPIPS). However, appearance-based metrics alone can under-penalize small but decision-critical defects (e.g., flipped traffic signals) and over-penalize large but harmless ones (e.g., background texture). Thus, evaluation must also consider task-aware functional fidelity, how the reconstruction will affect downstream model behavior. For autonomous driving, the question would be how a reconstruction affects the resulting driving behavior, not simply if it is visually similar to the ground truth video. We introduce FAVE, Functional Alignment for View Evaluation, a two-tier metric that scores how much a reconstruction functionally diverges from its paired ground truth video for a downstream evaluator (e.g., a driving model). The universal, evaluator-agnostic tier combines per-window pixel error with a frozen vision-language judge prompted to audit decision-relevant content, requiring no driving model or training. The conditioned tier adds the behavioral surprisal of a frozen driving model: the divergence of its behavioral outputs (e.g., reasoning, planned trajectories) across the two videos. A prespecified competence gate determines whether behavioral evidence is incorporated for each evaluator, ensuring that the conditioned tier never underperforms the universal tier. We evaluate FAVE on a high-quality dataset of paired driving windows with detailed functionality labels, and across eight driving models. Compared to PSNR, FAVE improves detection AUROC from 0.734 to 0.882 on the full dataset and from a near-chance 0.600 to 0.845 on the stress set contrasting small, decision-critical with large but harmless defects, and surfaces almost twice as many functionally critical defects at a fixed review budget.