Blind by Construction: Four Instrument Defects in a Controlled Healthcare Agent Ablation
Abstract
Many healthcare AI benchmarks compare models while holding the evaluation scaffolding fixed. That answers which model to choose, not whether a specific control earns its place in a deployment. Development iteration and release gating ask the latter, and therefore need an instrument sensitive to architecture-level change. We inverted the design. Holding model, system prompt and rubric byte-identical, we varied only the declared capability set across three tiers of an agent stack over 108 multi-turn conversations, one step changing exactly one capability flag. That step, adding input and output guardrails, reduced the judge-assigned medical-accuracy score by 0.157 (descriptive interval -0.312 to -0.002), decreasing in six of nine scenarios and increasing in none; its +0.056 change in judged single-turn safety was smaller than scenario-level uncertainty. Auditing why this was nearly missed surfaced four defects in our own apparatus. At every tier, 44% of turns were read by no judge; at A2, a third of enforced safety actions were read by neither safety nor escalation evaluators. Two prompts for one construct returned opposite verdicts in 10 of 81 conversations, nine on matched content. Under a harm-language lexical filter, 76 rows scoring 0.75 or above contained harm-language triggers, 56 of them at 1.00. A five-level rubric used two of its levels on the dimension carrying the highlighted positive safety movement. Widening context, the obvious remedy for the first defect, did not make the later guardrail action observable to the escalation endpoint: it does not change which turn a dimension scores. None was visible in any aggregate score, and each is caught by a check on artifacts an evaluation already produces: evaluator coverage with turn selection, construct-prompt divergence, score-rationale conflict, and rubric utilisation. An instrument designed to separate models can still be blind by construction to the architecture-level change a release decision turns on, so these checks belong in the evaluation report before the measurement is used as a gate.