Your world model may be worse than copying the last frame
Abstract
In one of our runs a world model explained 97.9% of the variance in the representations it was predicting. Repeating the previous frame explained 99.6% of the same targets; in squared error the learned model was five times worse than doing nothing. Eighteen measurements in our logs behave the same way, clearing an absolute 0.5 threshold while predicting worse than copying. On a stationary target with lag-1 autocorrelation ρ, copying scores 2ρ − 1. World models are evaluated at short horizons and dense sampling, where ρ is close to 1, so the trivial baseline is close to 1 as well. It also changes with the read-out: across 197 measurements the floor runs from 0.28 to 0.996 on one and from −0.20 to 0.76 on another. Between two configurations of a single model, the one with the higher score is the one predicting worse than copying the previous frame. Of 20 recent world-model papers, 12 report a prediction-quality number and 9 of those report nothing to read it against. We report lift, (v − floor)/(ceiling − floor), with the floor measured on every run. For a coefficient of determination lift is the mean-square-error skill score, which forecasters have computed against a persistence reference for decades.