Measuring impact of evaluation format on LLM-as-a-Judge for Long-Form Audio Reasoning
Abijah Sanjeev ⋅ AnaLuz Ferrer ⋅ Sarayu Badam ⋅ Aydin Feroz ⋅ Arjun Bahuguna
Abstract
LLM-as-a-judge methods are increasingly used to score open-ended model responses, but the evaluation format itself may affect the resulting judgment. We test this by holding the model response fixed and changing only how it is evaluated. Across 396 long-form audio reasoning tasks, Gemini 3.6 Flash generated one open-ended response per task without access to multiple-choice options. Each response was then frozen and evaluated by Claude Opus 5, GPT-4o, and Gemini 3.5 Flash under two protocols: direct free-response grading and multiple-choice answer mapping. Free-response grading produced 83.59% measured correctness, compared with 88.89% under answer mapping, a difference of 5.30 percentage points (95% CI: +2.53 to +8.33). Of the 35 responses whose classifications differed between protocols, 28 favored multiple-choice mapping and seven favored free-response grading (exact McNemar $p=0.00051$). Qualitative review found that many discordant responses contained relevant but incomplete information, a pattern consistent with answer choices helping evaluators map partial responses to the intended option. These results show that LLM-as-a-judge scores can change with evaluation format even when the underlying model output does not.
Chat is not available.
Successful Page Load