OntoPlan: An Ontology-Grounded Scene Representation and Agentic Framework for Scalable Robot Task Planning
Hyeongwoo Nam ⋅ Woongje Cho ⋅ Juwon Kim ⋅ Jongeun Choi
Abstract
Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 benchmark tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.83 average task success, compared with 0.30 for the strongest baseline, while using 17.2k total tokens per task on average, about 5.8$\times$ fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in both success and token cost. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://anonymous.4open.science/r/OntoPlan.
Chat is not available.
Successful Page Load