The Parallax Principle: Position-Marginalized Decoding for Order-Robust Language Models
Abstract
Autoregressive language models can produce different answers from the same evidence when its order changes. Motivated by this phenomenon, we articulate a parallax principle: when input order is incidental, predictions should depend on the evidence rather than on any single serialization. We introduce Position-Marginalized Decoding(PMD), a training-free method that aggregates next-token probabilities across four structured input views while maintaining one shared autoregressive trajectory. Unlike output ensembling, PMD combines views before their generations diverge and requires only next-token logits. Across seven backbones on controlled Natural Questions evaluations, PMD improves mean best-subspan exact match by 0.8-4.4 points, worst-placement accuracy by 1.9-9.4 points, and reduces position-induced variance by 85.6-98.8%. Across 15 additional model--dataset settings, it improves mean accuracy in 12, ties in one, and improves or preserves worst-placement accuracy in all 15. PMD requires no training or model-internal modifications, but uses four model streams, approximately fourfold computation, and four KV caches. These results show that synchronized multi-view decoding can improve robustness to incidental input order.