Relevance-Filtered Symbolic Context Remediates the Task-Planning Bottleneck in Embodied LLM Agents
Abstract
Large language models (LLMs) are increasingly used as task planners for virtual embodied agents, where success depends on long chains of procedural, domain-specific knowledge that is dense, version-specific, and unreliably encoded in pretraining data—the tacit domain knowledge a symbolic planner is given explicitly. We characterise the resulting procedural-context bottleneck. Across five models from four architecture families (8B to 671B where disclosed), success on the multi-hop (L2) tier of 60 held-out Minecraft tasks (44 crafting, 11 gathering, 5 combat) is only 10–30% without retrieval (median 15%), despite 45–95% (median 85%) on single-material (L1) tasks—a failure model scale alone does not fix. Using a training-free retrieval mechanism (SAER: State-Aware Example Retrieval) and a plan-quality metric, we decompose what remediates it. The existence of any same-category in-context example is the primary lever (random retrieval gains +31 percentage points, pp, on Qwen2.5-Max, +20 pp under a stricter material-balance scorer); relevance adds a smaller increment that clears Holm only at the pre-specified family size, never the Bonferroni bar—a bar its discordance count (b+c=7) put out of reach—and does not survive the stricter scorer at all (BM25 +38 pp, SAER +50 pp on Qwen2.5-Max, SAER over BM25 +12 pp, 95% CI [+3, +20]; under debiting +15 pp, +23 pp and p=0.125 for the increment); retrieved-template correctness also binds on one exploratory contrast (SAER fell to 1/13 at L2 against random's 6/13 until six defective demonstrations were repaired, on a superseded 35-task benchmark); and model capacity does not gate the remedy, though three of five sizes are undisclosed so we test no scaling trend (all five gain +20 to +50 pp; four clear pre-specified Bonferroni, five clear post hoc Holm). The benchmark's five combat tasks, on a separate demonstration pool, replicate the existence effect outside crafting but are too few to separate the variants. We thus identify which facet of a symbolic domain model a neural planner must be handed rather than assumed to hold.