Physical Understanding for Decision-Making: Bridging Foundation Models and Reliable Agents
Abstract
Foundation models are moving from interpreting observations toward supporting embodied interaction, making physical understanding an increasingly important capability. In this setting, tactile sensing provides physical agents with direct evidence of how objects respond to contact. Existing tactile language models typically align tactile features with an external vision and language space before passing them to a text LLM through a prefix or adapter. However, this endpoint design does not match the hierarchical perceptual interface now shipped by multimodal LLMs (MLLMs) built on DeepStack, including the Qwen3-VL family and Cosmos~3, which condition language generation jointly on final visual tokens and intermediate readouts. To address this gap, we introduce NativeTouch, which co-designs tactile alignment with the downstream interface of such an MLLM. It aligns a tactile encoder with the intermediate and final visual readouts consumed by the MLLM, then delivers tactile features through dedicated native tokens and DeepStack without another large projector. On a tactile understanding benchmark with examples from HCT and SSVTP, the NativeTouch-Nano and NativeTouch-Super versions achieve relative gains over TVL of 65.12\% and 62.15\% in CKA, 5.59\% and 6.23\% in tactile-vision Top 1 retrieval, and 9.21\% and 6.93\% in judge score.