MineOcclude: Separating Visual Evidence, History, and Motion Predictability
Abstract
Decoding a hidden object's position can reflect information in past images, predictable motion, or the way a probe is fitted. MineOcclude separates these factors through paired Minecraft captures with shared motion and continuous position targets. A fixed probe of frozen MineWorld with verified RGB inputs achieves 0.083-0.107 blocks RMSE in air and glass and 1.600 under stone on 25 central trajectories. We then evaluate 64 training and 64 test families of varied paired motions across two layouts. In a retrospective readout analysis using the same representations and family split, training directly for position differences reduces stone contrast error at ages 1-10 from 1.176 to 0.280 for MineWorld, from 1.125 to 0.183 for history pixels, and from 1.019 to 0.179 for VideoMAE. History pixels, VideoMAE, and an image-history positive control outperform the zero-difference reference (0.265 blocks RMSE). Source-selected scaling alternatives reduce some extreme transfer errors but do not restore reliable transfer. These results identify probe objectives and preprocessing as important parts of an occlusion evaluation, while remaining specific to finite input windows rather than persistent object memory.