Interpreting Ego-Motion in a Driving World Model
Abstract
Driving world models can generate realistic future video continuations from past context frames, but whether their internal representations track physical kinematics or motion-correlated visual patterns remains an open question. We examine residual-stream activations in the flow-matching model Orbis across three continuous nuPlan odometry channels: speed, signed yaw rate, and longitudinal acceleration. Speed is decoded accurately even from early layers, whereas yaw rate and acceleration emerge deeper in an intermediate layer band before decaying toward the output. Iterative probe erasure indicates that yaw rate is concentrated within a compact subspace of 44 directions, whereas speed remains broadly distributed across more than 400 directions. Finally, we evaluate activation interventions during rollout using probe subspaces and sparse dictionaries. While specific intervention methods, such as rank-20 probe writes, can induce coherent heading changes in selected contexts, internal interventions do not reliably execute intended route turns across diverse scenes, often introducing visual artifacts such as structural distortion and color shifts.