From Embodied Contingency to Arbitrary Symbols? A Diagnose-Repair-Retest Benchmark
Yuhao Xu
Abstract
Embodied symbol grounding is often motivated developmentally, but competence in one sensorimotor relation need not transfer to later symbolic tasks. We test whether performance-matched mirror-contingency pretraining improves two nonlinguistic downstream families: arbitrary token binding across morphology and remapping, and source-mediated retrospective revision. Seven capacity-matched 22,190-parameter recurrent world models receive canonical, yoked, reversed, direct-identity, or no-mirror histories; their cores are then frozen and matched downstream heads are trained from ordinary task feedback. A single heldout campaign evaluates nine systems on 2,160 episodes and 4,320 decisions. Mirror competence is successfully matched, but the candidate fails three absolute token-family criteria and has no positive course effect over the no-mirror control in either family. The revision family reaches ceiling for candidate and generic controls. An audit further reveals exact functional equivalence between candidate and generic recurrent branches. We prospectively repair that defect with a nonisomorphic operator, then freeze it before a 32-seed retest. Candidate-minus-generic core accuracy is $-0.0020$, 95\% CI $[-0.0120,0.0081]$, and both proper scores are worse, although operator lesion lowers accuracy by $0.1195$, $[0.1061,0.1328]$. Thus reachability and local causal use are insufficient for transfer superiority; developmental claims require strong controls, pre-evaluation audits, and a separate retest.
Chat is not available.
Successful Page Load