When Reward Becomes the Wrong Objective: Auditing AI Benchmarks for Misaligned Outcomes
Abstract
A reward function is only useful insofar as optimising it, or ranking controllers by it, produces the outcome it was meant to proxy for; when that link breaks, a benchmark can reward exactly the wrong thing while still returning a clean, publishable number. We audit one applied AI-for-good benchmark, a reinforcement-learning environment for campus microgrid load shedding, and find the link broken through three classes of failure, spanning six specific findings being implementation-level bugs that silently favour some controllers over others, a structural failure in which a policy that does nothing at all outranks every evaluated controller, a result already visible, unaltered, in the benchmark's own published results, and, most consequentially, a regime in which the reward ranks a policy that abandons the critical-load cluster above one that protects it, precisely where that protection matters most. We re-implement and validate the benchmark to three decimal places before running any diagnostic, so every finding is a statement about the actual artifact. From this audit we distill a four-check protocol, null-policy, control-authority, reward decomposition, and reward-independent outcome reporting, each executable in under a minute without training a model, and argue it should be run routinely before an AI-for-good benchmark's headline claims are trusted. We report the evidential status of every claim explicitly, including which findings rest on a hand-constructed diagnostic rather than a trained optimiser's own behaviour, and which parts of our protocol remain demonstrated on one benchmark rather than validated broadly.