Deployment-Centered Evaluation: A Case Study on User Feedback with EHR-Integrated Language Models
Abstract
Large language models (LLMs) are increasingly integrated into clinical care, making it essential to both evaluate and improve the utility of these systems. However, static benchmarks fall short: these evaluations tend to measure correctness rather than user acceptance, aggregate performance across queries, and require densely annotated datasets. In this work, we conduct a deployment-centered evaluation where we perform a query-level analysis of user feedback that is dynamically collected by a deployed EHR-integrated LLM system in an academic medical center. First, we find that deployment-specific context (e.g., provider types, departments, model type) significantly affects whether users reject or accept the system output. Then, we operationalize this observation to develop a query-level intervention. Using deployment-specific context together with query content as features, we then train a pre-response model to predict whether a specific user will reject a specific query. We conduct a prospective analysis of our model over 4.5 months of user feedback, finding that our prediction model achieves an AUROC of 0.719. Altogether, our empirical case study demonstrates the feasibility of predicting user rejection using deployment-specific context, opening the door to targeted guardrails.