Conformal Safety Certificates in Learned Representations: What Survives and What Must Be Rebuilt
Abstract
Real environments are often too complex to model by hand, and robots therefore learn a world model instead. The world model compresses raw sensor data into a compact summary known as a latent state and predicts how that state evolves as the robot acts. Planning then happens inside this learned representation rather than in the physical world. Safety requirements such as clearance are nonetheless defined in the physical world where conformal prediction constructs guaranteed margins in metres, and no existing work examines what those metres mean inside a representation never trained to respect physical distance. This paper argues that such a margin carries three properties that current systems do not distinguish and that only the first survives a change of representation. The three are whether coverage still holds, whether it retains its original physically bounding property, and whether its shape remains useful for planning. A proof-of-concept study calibrates the same margin in physical space, inside the representation, and by a decoder that reads clearance from the latent state, and then repeats every comparison on frozen DINOv2 features as a second and more generic representation. Every margin kept its coverage while regions built on the DINOv2 features hid larger clearance errors and falsely excluded more of the usable space. Only the decoder-calibrated margin remained a bound in metres at no measurable cost in usable space, which suggests that a physically grounded safety guarantee must be built into the world-model interface rather than assumed to survive the change.