Equivariance as an interface property: state-carried positional codes in neural cellular automata
Abstract
Coordinate-conditioned decoders, used in many implicit representations such as neural radiance fields, sinusoidal representation networks, and neural cellular automata, evaluate a function at coordinates supplied by a sampling grid. This directly couples the outputs of these decoders to the grid. This coupling can introduce two side effects: broken equivariance and visible sampling artifacts in the output. While these artifacts are typically addressed by increasing the encoding's capacity and consequently its expressivity, we explore a different approach with the realization that equivariance is fundamentally tied to the decoder interface. If all query-dependent inputs to a decoder shift alongside the content, exact translation equivariance is guaranteed by design. We demonstrate this by comparing grid-anchored encodings to a state-carried encoding within a neural cellular automaton. Our state-carried code achieves this exact equivariance at lattice-aligned shifts, whereas every grid-anchored baseline we train fails. We also perform ablations to separate the two side effects, finding that the lattice artifacts are downstream of what the coordinate encodes, while the decoder equivariance follows from how that coordinate transforms. For coordinate-conditioned models, we argue both concerns are ultimately a matter of interface design.