Neuro-Symbolic Validation for Multi-Robot Task Updates from Human Language
Abstract
Multi-robot systems operating in dynamic environments may encounter conditions that were not known when tasks were initially assigned. Human operators can provide this new information during execution, and large language models (LLMs) offer a flexible way to interpret it and suggest how robot tasks should be updated. However, correctly understanding the human input does not guarantee that an LLM will respect robot capabilities or environment constraints. We study a neuro-symbolic approach in which an LLM identifies the affected task context, while deterministic symbolic reasoning determines feasible robots and whether the task should be reassigned or the route replanned. We evaluate Qwen3 8B and Llama 3.1 8B on reports describing deep water, stairs, and blocked routes. Both models perform well on language grounding but frequently make incorrect capability or task-update decisions. Symbolic reasoning raises end-to-end exact-match accuracy from 0\% to 96.7\% for Qwen and from 16.7\% to 100\% for Llama. Qwen's remaining errors arise from an incorrectly grounded location, which the symbolic layer does not reinterpret. These results show that neuro-symbolic decomposition can substantially reduce downstream constraint errors when semantic grounding is correct, while exposing a clear boundary when it is not.