Measuring impact of evaluation format on LLM-as-a-Judge for Long-Form Audio Reasoning
Abstract
LLM-as-a-judge methods are increasingly used to evaluate open-ended model responses, but the evaluation format itself may influence measured performance. We isolate this effect by holding model outputs fixed and comparing two evaluation protocols on 396 long-form audio reasoning tasks: direct free-response grading and multiple-choice answer mapping. Gemini 3.6 Flash generated one open-ended response per task without seeing answer choices, and the same frozen responses were evaluated by Claude Opus 5, GPT-4o, and Gemini 3.5 Flash under both protocols. Free-response grading produced 83.59% measured correctness, compared with 88.89% under multiple-choice mapping, a difference of 5.30 percentage points (95% CI: +2.53 to +8.33). Among 35 protocol-discordant responses, 28 favored multiple-choice mapping and seven favored free-response grading (exact McNemar p = 0.00051). Qualitative analysis found that incomplete synthesis was the most common discordant pattern, consistent with answer choices providing semantic structure that helps evaluators map partial responses to the intended option. These results show that LLM-as-a-judge scores can vary systematically with evaluation format even when the underlying model output is unchanged.