MedConsultBench: A Construct-Based, Traceable Measurement Framework for Clinical Consultation Agents
Abstract
Final-answer accuracy cannot distinguish consultation agents that reach the same diagnosis through clinically different evidence, treatment, and revision paths. We introduce MedConsultBench, a multi-construct measurement framework that maps six consultation capabilities to eligibility predicates, trace events, estimands, and validation gates. Unsupported case--construct cells remain unavailable rather than becoming zero, while hard safety events are non-compensatory. The released bank contains 9,832 scenarios; the seven-model evaluation uses an independent 400-case evidence panel, a 400-case common panel, and a 56-case safety stress panel; a same-protocol 9,196-case Qwen run supports full-bank calibration. Within their declared scopes, the evidence release matcher attains 0.989 precision and 0.978 recall, the diagnosis--evidence linker attains 0.987 precision, and the deterministic safety checker exactly reproduces 99 assessable adjudicated labels under its frozen rule scope. Under the three-turn protocol, seven models yield evidence coverage of .556--.640, Clinical Ordering compliance .015--.073, actively elicited Diagnostic Grounding .519--.617, strict failure .411--.929 under adversarial embedded constraints, and Adaptive Revision .434--.533; one configuration triggers the hard-safety gate. No configuration dominates all reported construct cells: the model with the lowest strict safety failure also has the lowest elicited grounding, and typed perturbations selectively change their target estimands across six trace-complete profiles. Panel--bank calibration further quantifies transport behavior and detects a directional Adaptive Revision shift. The resulting profiles expose capability reversals that a scalar leaderboard cannot represent.