Confabulation in a Caption-then-Generate Pipeline: Evaluating Vision-Language Generation of Adolescent Health Information
Abigail Naa Amankwaa Abeo
Abstract
Many of the health habits people carry into adulthood are formed during adolescence, yet adolescent health literacy is often low, and the technologies built to deliver health information are rarely designed with young people in mind. Vision-language models make it possible to turn a photograph into health information (output), but text an adolescent can act on is a stricter target than text that is merely fluent. An output can be well-formed and factually correct yet fail its audience: too generic, pitched for an adult, or detached from what the photograph shows. We ask a narrow version of a broad question: can an off-the-shelf vision-language pipeline take a photograph an adolescent captured and produce a short, usable piece of health information from it? We designed and built a two-stage, caption-then-generate pipeline. In the first stage, an image is passed to a vision-language captioning model (Florence-$2$-Large) which produces a caption. In the second stage, that caption, together with a fixed system instruction and a few-shot exemplar, is passed to an instruction-tuned language model (Llama $3$ $8$B Instruct), which generates a short health message. We use both models as released, without fine-tuning; our contribution is the system's design, its prompting, and the evaluation and analysis framework around it. One property of this design governs everything downstream: the language model never sees the image; it works only from the caption, so anything the caption omits is gone before generation begins. We applied the pipeline to images from a school Photovoice session (Wang and Burris, 1997), a participatory method in which people photograph their surroundings; in our study, the participants were students aged $12$-$16$. Six health literacy experts independently rated every output for Clarity, Relevance, and Helpfulness, each on a three-point scale (Very, Somewhat, Not at all). To locate why the weaker outputs failed, we introduce a four-level confabulation severity scheme. We use confabulation rather than hallucination deliberately (Smith et al., 2023); the model is not misperceiving an image it cannot see, but filling an underspecified caption with plausible, invented content. The scheme codes how far each output departs from its caption, annotated blind to the expert ratings. Outputs were rated highly for Clarity far more often than for Relevance or Helpfulness (rated "Very" roughly $69\%$ of the time for Clarity, versus $59\%$ for Relevance and $57\%$ for Helpfulness): fluency was not the bottleneck. Confabulation severity, coded independently of the ratings, tracked the expert judgements closely. As severity rose, Helpfulness "Very" ratings fell from $75\%$ for outputs that stayed within the caption, to $14\%$ for those built substantially on invented content, while "Not at all" ratings rose from $3\%$ to $48\%$ (Kendall's $\tau_b = -0.39$ helpfulness, $-0.51$ relevance, $-0.32$ clarity; all $p<.001$). The outputs that departed most from their captions were, independently, the ones experts found least helpful. Beyond the ratings, experts' open-ended comments surfaced three recurring gaps, each matching a recognised generation challenge: misinterpretation of visual elements (error propagation), missed contextual focus (content selection), and insufficient actionability (pragmatic generation). Prior work has generated health content and studied adolescent health literacy separately, but generating health information for adolescents from their own photographs remains, to our knowledge, unexplored. Our contribution is a purpose-built pipeline and evaluation framework for this task, together with a confabulation severity scheme for caption-to-output generation. Our central evidence is that fluent output is not the same as usable output, and that domain-expert evaluation, rather than fluency or automatic metrics alone, is needed to tell them apart (Li et al., 2025).
Chat is not available.
Successful Page Load