System-Prompt Invariance for LLM-as-Judge Evaluation: A Linguistic Robustness Study
Abstract
LLM-as-judge pipelines depend on system prompts to define the task a response-generation model is expected to perform, while a separate, fixed rubric governs how a judge scores the resulting output; the instruction hierarchy frames such task-defining system prompts as carrying privileged, application-level authority. Whether that authority translates into behavioral stability under character-level perturbations and task-outcome-preserving lexical rewrites of the prompt itself has not been systematically tested. We extend nine transformation families from BiGGen Bench’s grounding capability into three matched conditions each: an unperturbed base, a single character-level typo, and a task-outcome-preserving lexical substitution verified against a rule stricter than plain WordNet synonymy. Across four response-generation models’ 108 responses and three automated LLM graders, we collect 324 judgments and find observed but heterogeneous score drift: three of four models show drift under perturbation while one shows none, with no evidence of self-enhancement in the available same-family grader–response-model comparison. Grader compliance with a requested structured output format also varies substantially, from perfect to majority-fallback across the three graders. An independent human validation study on a 54-item DeepSeek-only subset, conducted by three blinded raters, achieves strong inter-rater reliability (Krippendorff’s alpha = 0.9264) and close agreement with automated median scores (r = 0.9946, 98.1% exact agreement), supporting the automated pipeline as a reasonably faithful proxy for human judgment on this task family. These results suggest model-dependent sensitivity to system-prompt perturbation rather than establishing invariance as a general property of instruction-following models, and identify structured-output compliance as an independent, underexamined, and empirically significant axis of LLM-as-judge reliability.