Demonstration Conditioning Improves Jagged Frontier Model Performance on Diverse Agentic Tasks Across Digital and Physical Domains
Abstract
Frontier models perform unevenly across tasks that one benchmark average re- ports as a single number, and the cheapest training-free repair is a prior episode placed in the context window, which for a multimodal agent adds thousands of input tokens of frames and action history, 6.7k to 16.1k across domains. A score does not show whether the model reasoned over that context or retrieved an an- swer from it. Holding the prompt fixed, we vary only the demonstration block across five multimodal agentic domains, VisualWebArena, VLABench, Assem- bly101, IndEgo and Bench2Drive-VL, and record per item whether the demon- stration held the item’s answer. Competence is uneven inside every benchmark’s own categories. A demonstration raises every domain and all five models, +2.8 to +14.1pp on the web tasks, by uneven amounts and in a pattern that does not follow the competence. With the answer removed, the lift survives in two of the four con- structed domains, and the surviving share runs from 15% on driving to 62–72% on VLABench. Irrelevant context costs up to 5.1pp, and the same episode as frames rather than text roughly halves answer copying.