Faithful to the Teacher, Far from the Reference: Auditing Distilled LLM Judges for Ordinal Matching
Jasmine Qi ⋅ Elaine Xie ⋅ Danylo Dantsev ⋅ Jian Shi
Abstract
Distilling an LLM judge lowers inference cost, and student-teacher agreement, or fidelity, is commonly treated as evidence of success. In ordinal matching, however, a student can copy its teacher while further suppressing the strongest candidate-job matches. We audit this failure on resume-jd. Writing $S$, $T$, and $R$ for student predictions, teacher labels, and the released reference labels, we measure agreement with linear-weighted Cohen's $\kappa$, which penalises distant ordinal errors more heavily. Our identifying experiment is a three-seed matched ablation that holds the student, data, optimisation, and evaluation rows fixed while changing only the supervision source: teacher verdict-token distributions matched via KL (soft), their argmax labels (hard), or reference labels. Teacher supervision yields high fidelity, $\kappa_{\mathrm{lin}}(S,T)=0.649 \pm 0.020$ (soft) and $0.630 \pm 0.005$ (hard), but only $\kappa_{\mathrm{lin}}(S,R)=0.061 \pm 0.008$ and $0.061 \pm 0.007$. Reference supervision reverses the ranking: $\kappa_{\mathrm{lin}}(S,R)$ rises to $0.456 \pm 0.005$, and F1 on the highest, decision-relevant Good-Fit class rises from at most $0.048 \pm 0.001$ to $0.397 \pm 0.010$. Every within-seed paired contrast has the same sign. Four diagnostics localise the result. A1 finds the soft verdict distribution nearly one-hot. A2 finds high same-day, same-route teacher repeatability ($\kappa_{\mathrm{lin}}=0.908$). A3 finds the triangle-inequality bound slack in every run. A4 finds the prespecified 200-row aggregate gate underpowered. In contrast, comparing teacher and reference Good-Fit rates would flag the discrepancy in 47 rows. The teacher assigns Good Fit to 5.3% of pairs versus 26.0% in the reference, and teacher-supervised students use it still less. Two completed Gemini seeds reproduce the fidelity-reference gap. We conclude that fidelity and reference agreement are distinct evidence: screen consequential class marginals before distillation, then run a powered agreement audit. The protocol separates target imitation, agreement with a released reference, and construct validation.
Chat is not available.
Successful Page Load