Causal Calibration of Symbolic State in Embodied Language Agents
Abstract
Neuro-symbolic agents often rely on symbolic facts such as whether a container is open, an object is being held, or a target object has been heated. Mechanistic read- outs can reveal such concepts inside a neural model, but the presence of a readable representation does not establish that the representation is actually used to select an action. We test whether the strength of symbolic predicates under the Jacobian lens (J-lens) predicts their causal influence on embodied decisions. We construct controlled decision points from ALFWorld household tasks and measure J-lens representations of five grounded predicates: OPEN, HOLDING, HOT, CLEAN, and COOL. At each decision point, we suppress the corresponding J-space com- ponent and measure the change in the model’s preference for the correct next ac- tion, comparing against strength-matched unrelated concepts, norm-matched ran- dom directions, and wrong-layer controls. Across decision-relevant states, J-lens strength reliably predicts intervention effect (Spearman ρ = 0.46, p < 0.001), whereas the relationship is decoupled when the same predicates are behaviorally irrelevant (ρ = 0.09, p = 0.12). Targeted ablations reduce correct-action margin substantially more than matched controls (+0.47 logits), scaling monotonically with intervention strength. These results indicate that J-lens magnitude provides a calibrated proxy for causal importance conditioned on decision context, distin- guishing active symbolic state from merely accessible representations.