PhysTacGen: Physics-Aware Visual-Tactile Sensor Image Generation
Abstract
Realistic physical interaction is a cornerstone of embodied intelligence, yet obtaining high-fidelity tactile data remains significantly more expensive than visual data. This scarcity necessitates Visual-to-Tactile synthesis to bridge the Sim-to-Real gap. However, existing data-driven approaches often treat this as a naive image-to-image translation task, neglecting fundamental contact mechanics. Moreover, the inherent spatial misalignment in real-world visual-tactile datasets frequently forces generative models to average out spatial uncertainties, resulting in blurry and textureless outputs. To address these limitations, we introduce PhysTacGen, a physics-aware generation framework that enforces consistency between visual appearance and tactile mechanics. Our system makes three technical contributions. Firstly, we propose Group Tactile Policy Optimization (GTPO), a reinforcement learning alignment strategy that refines Large Multimodal Models to infer latent physical properties by rewarding physics-consistent reasoning chains. Secondly, we introduce a Semantic-Driven Data Curation and Visual Prior pipeline, leveraging DINOv2 to filter spatially misaligned data and extracting pure depth maps to decouple macro-geometry from surface texture. Finally, we present a physics-driven generative architecture based on Stable Diffusion XL (SDXL). It synthesizes tactile images via a ControlNet conditioned on GTPO-generated physical descriptions and depth-augmented visual priors. Extensive experiments demonstrate that PhysTacGen achieves state-of-the-art perceptual quality and significantly enhances grasp force prediction for zero-shot Sim-to-Real transfer.