Competence without Performance: Representation Without Control on ARC-AGI-3
Hasaan Ahmad ⋅ Yuxuan Sun ⋅ Ali Zahin ⋅ Vaibhav Mehra ⋅ Soumya Banerjee
Abstract
DreamerV3 masters more than 150 control tasks under a single configuration, from Atari to collecting diamonds in Minecraft from scratch, and beats PPO on multi-step ARC-AGI-1. A 7M-parameter Tiny Recursive Model reports 87\% on Sudoku-Extreme, 85\% on Maze-Hard and 45\% on ARC-AGI-1, above language models orders of magnitude larger. We put both through one Gymnasium- and DreamerV3-compatible substrate for ARC-AGI-3, a benchmark of interactive games whose rules and goals must be inferred by playing them, and both collapse: $0.005$ and $0.018$ against a human-parity $1.0$, over six games and two seeds, with cross-game pretraining giving no consistent benefit. We ask why, and separate the two answers a terminal score cannot tell apart. The models do represent these environments, but not uniformly. DreamerV3's reconstruction converges on every game and level identity is linearly decodable from its frozen latent above raw-pixel and elapsed-time controls on two games it never scores on, while its imagined rollouts fall below copying the last frame at every horizon on every game; the recursive model, trained on the same replays under a loss whose optimum is not the static scene, clones human play into non-zero scores and predicts transitions above that same baseline in 11 of 12 cells. What fails is downstream of all of it. On 21 of 25 games the reward is identically zero for the entire 500k-step run, imagined returns are constant, and the actor sits at its uniform maximum entropy of $\ln(4102)$ nats; score appears in exactly the four games where it departs. The two interventions that move anything, a factored action interface and a cloned human prior, act on that step rather than on the model: the factored interface ignites reward discovery on \texttt{cd82} in 6 of 6 seeds against 0 of 6 while making the world model's fit fifteen times worse. We read this as competence without performance, and identify what ARC-AGI-3 withholds that both models' prior successes supplied.
Chat is not available.
Successful Page Load