Zdravko AI: A Paired Pilot Evaluation of Prompt-Level Safety Guardrails for Conversational Nutrition Advice
Abstract
General-purpose conversational assistants are widely used for everyday nutrition questions, where harm can arise both from insufficient caution (enabling unsafe restriction, disordered eating, or unsafe supplement use) and from excessive caution (refusal, over-referral and loss of useful guidance). We describe Zdravko AI, an instruction-level safety layer for conversational nutrition advice explicitly designed to optimise both objectives, and we report a paired pilot evaluation against an unmodified assistant on a 50-case safety benchmark spanning eight domains, including deliberately low-risk cases intended to detect over-caution. Responses were scored by an automated evaluator on five 0–2 safety dimensions and a holistic 0–4 Safety Score. The safeguarded condition received a higher overall score in 13/50 cases, the same score in 34/50, and a lower score in 3/50 (analysis-plan-defined exact two-sided sign test, nominal p = 0.0213, not to be read confirmatorily). Cases were administered sequentially within one conversation per condition, so observations are not independent replicates; discordant outcomes clustered in contiguous runs and in recurring context-dependent patterns rather than as independent events. All three deteriorations were consecutive supplement questions in which the evaluator judged the safeguarded system over-cautious rather than unsafe, consistent with a categorical prohibition in the instruction text. We report this as a hypothesis-generating signal about safeguard behaviour, not evidence of clinical safety, and draw methodological lessons for evaluating context-dependent health conversations.