AgentVMEval: Calibrating When Feedback Earns Its Cost in Visual-Editing Agents
Abstract
When should a visual agent inspect intermediate results instead of planning once? Existing evaluations entangle instruction understanding, tool reliability, renderer quality, and judge error, obscuring the value of feedback. We introduce AgentVMEval, an exactly scored intervention that varies linguistic form and tran- sient tool failure while holding source scenes and hidden edit plans fixed. We report 6,720 executions over deterministic controls and three vision-language model fam- ilies. A public template parser scores 100% on canonical instructions and 0% on parser-resistant paraphrases. With clean tools, full learned planners achieve 97.5–100% exact success; iterative REACT provides no reliable gain while costing 2.99–5.23×more than PLANONCE. At 20% failure probability, PLANONCE falls to 50.8–51.7%, whereas REACT reaches 100% and recovers every injected fail- ure (p= 10−5). Across six Claude failure rates, PLANONCE tracks the analytic compound-risk curve within 3.1 points on average, while REACT remains at least 99.6%. A generic non-agentic retry also restores 100% success on canonical tasks, showing that feedback earns its premium only when recovery cannot be reduced to a fixed policy. AgentVMEval is a calibrated unit-test complement to naturalistic benchmarks: it separates semantic competence, recovery, and operational cost with exact ground truth.