Anchor: Preventing Artifact Drift for Reliable Agent Verification
Abstract
Reliable evaluation of long-horizon agents requires verifiers that match what the task actually asks. Verifiable rewards have driven progress in domains like math and coding, yet building similarly reliable environments for open-ended enterprise workflows repeatedly struggles to balance realism, verifiability, and scale. This creates a validity problem we call artifact drift: when instructions, environments, oracles, and verifiers are created separately, they frequently fail to agree on what a task requires, resulting in environments that are unsolvable, reward-hackable, or inconsistent. We introduce Anchor, a benchmark-construction pipeline that formalizes domain expert specifications of business workflows into constrained optimization programs. From a single parametric specification, the pipeline jointly produces a natural-language instruction, environment configuration, solver-certified ground-truth solution, and state-based verifier. With this pipeline, altering parameters yields new tasks with controlled difficulty and known optimal solutions, producing harness-agnostic environments whose rewards depend solely on end-state business correctness. We apply Anchor to produce ERP-Bench: a benchmark of 300 long-horizon tasks spanning procurement and manufacturing workflows in a production-grade ERP system. Our pipeline organizes tasks into diverse families with calibrated difficulty profiles, and allows generated environments to evaluate coding, browser, and computer-use agents using a shared verifier. Across these harnesses, we find that frontier models follow all rules and instructions in 26.1% of trials and reach fully optimal decisions in only 17.4% of trials. Overall, Anchor and ERP-Bench offer a concrete recipe for building auditable evaluation environments whose verifiers remain aligned with the economically valuable tasks they are intended to score.