Grounded and Faithful Vision-Language Models for Real-World Deployment
Abstract
Vision-language(-action) models are rapidly crossing a threshold: from systems that describe the visual world to agents that must act within it. Yet a critical gap persists between what these models appear to understand and what they can reliably and faithfully ground in the underlying visual and physical world. Despite remarkable progress across robotics, autonomous systems, embodied agents, and interactive AI, current systems frequently exhibit failures in grounding reliability, hallucination mitigation, reasoning consistency, and robust behavior under dynamic and uncertain real-world conditions. This workshop brings together researchers and practitioners from academia and industry to advance methods, benchmarks, and systems for grounded and faithful multimodal intelligence. Rather than viewing grounding and faithfulness as downstream properties to optimize after model development, we emphasize them as fundamental principles for building AI systems whose predictions, reasoning, and actions remain aligned with visual and physical reality. By convening researchers across multimodal learning, robotics, embodied AI, autonomous systems, and world models, the workshop aims to establish concrete pathways toward AI systems capable of reliable perception, faithful reasoning, and robust interaction in open-ended real-world environments.