Evaluating Generative AI for Personalized Longitudinal Diabetes Decision Support
Abstract
We present DM-Bench, the first large‑scale benchmark for evaluating large language models (LLMs) on patient‑facing diabetes management tasks grounded in longitudinal, individual‑level wearable and behavioral data. Existing health benchmarks focus on generic medical knowledge, clinician‑facing workflows, or objective tasks that fail to capture the continuous, highly personalized, and context-dependent reasoning required in daily diabetes management. To address this gap, DM-Bench introduces a dataset-agnostic evaluation framework tailored to the unique challenges of prototyping patient-facing AI solutions in diabetes, glucose management, and metabolic health domains. The benchmark spans 7 distinct task categories, reflecting the breadth of diabetes decisions, from basic glucose interpretation to long-term planning over extended time horizons. We instantiate DM-Bench using the open‑source OhioT1DM dataset and DM‑large, a dataset comprising one month of continuous glucose monitoring and behavioral time‑series data from 15,000 individuals across type 1, type 2, and prediabetes/health & wellness populations. Using these datasets, we generate 360,988 personalized, contextual questions across the 7 tasks. DM-Bench defines a multi-dimensional evaluation protocol covering 5 metrics: accuracy, groundedness, safety, clarity, actionability. We evaluate recent LLMs and observe substantial performance variability across tasks and metrics, with no model consistently outperforming the others. DM-Bench exposes domain-specific performance gaps not captured by existing benchmarks, advancing the reliability and practical utility of AI solutions in diabetes care.