Evaluating LALM Judges on Freeform Audio Reasoning
Abstract
Audio reasoning benchmarks for Large Audio Language models(LALMs) are moving from multiple choice to free-form answers. That shift is useful because free-form answers are harder to game, but it creates a new problem: it now needs to be decided whether each answer is correct, and that job is increasingly handed to an automated judge. Whether those judges are up to the task is largely untested. This paper puts audio-language judges under scrutiny. We construct a human-labeled set of free-form MMAR answers that specifically allows for question ambiguity and semantic variance in answer, and we test for whether LALM judges can effectively replace a human annotator. We run seven judges across conditions that vary the input for each judge: the correct answer, the audio clip, or neither. We find that judges reach human level when given the correct answer, fall well below it when the answer is withheld, are not helped by the audio when ground truth is present, and grade most accurately on questions they could have answered themselves, which is evidence that they match responses against a reference rather than judge correctness from the signal. The practical consequence is that audio judges can be trusted where a reference exists but not in the reference-free setting where automated judging is needed most.