House Rules: A Benchmark for Agents under Changing Requirements
Abstract
Agents carrying out long tasks must accept requirements that a user, an operator, or a safety policy revises mid-task while the goal stays the same. Most agent benchmarks instead fix the specification for an episode. We introduce House Rules, a benchmark for agents that must follow requirements revised during a task. Its constructor augments existing interactive tasks with an editable symbolic specification of rules, hard constraints, and soft preferences, and with timed natural-language updates. Each update is paired with ordered gold edit operations, so the evaluator knows what changed while the agent receives only language, and feasibility checks keep a successful route open after every edit. The benchmark comprises 450 instances across ALFWorld, CookingWorld, and ScienceWorld, and we evaluate direct, reactive, and planning agents on one frozen seven-billion-parameter language model. Task success, compliant success, and constraint violation come apart. The strongest agent completes 17.8% of tasks yet violates an active constraint in 15.1% of episodes, its violation rate rises from 4% to 25% across difficulty profiles, and a privileged symbolic shield removes every executed violation without changing task success.