Widening Harm Evaluation Coverage with a Judge That Reads the World
Abstract
An agent acting in a stateful environment can leave evidence of harm in two places: the record of what it did and said, or the world it changed. Safety evaluation for discovering harm reads one or the other in limited ways: an LLM judge over the trajectory, or deterministic checks over the final state. We consider the trajectory and the environment state as two channels of evidence. LLM-based verifiers that go and inspect the environment exist, but all of them verify task completion; none has been used for the task of detecting and confirming harm in a safety setting. We fill this gap with an LLM that has read access to the environment state, supplied with a harm taxonomy. To evaluate the approach, we take three environments, two from DecodingTrust-Arena and one from OSWorld, and hold the judge model and the harm taxonomy fixed. We then vary only what the judge observes: the trajectory, read-only tools over the world, and a single judge given access to both. On the two DecodingTrust-Arena environments, a union of the trajectory and state-only judges recalls 94\% of confirmed harms against the trajectory judge's 53\%; on OS-Harm the union recalls 77\% against 67\%. We also show that a single judge given both channels does not perform better and gets more false negatives than the state-only judge. We also ledger disagreements and observe that the two channels fail on different episodes. Our results suggest a union of two channel-specific verdicts as a robust and flexible approach to detecting and verifying harm in agentic settings.