MAWA: A Clinician Co-Designed Multi-Agent Framework for Weight-Management Advice
Jiahe Liu ⋅ Siteng Zhang ⋅ Li Lu ⋅ Qi Tang ⋅ Zongyuan Ge
Abstract
Overweight and obesity are leading modifiable drivers of metabolic disease, yet individualized weight-management follow-up remains labor-intensive. Large language models can draft lifestyle advice at scale, but clinical use requires trustworthy targets, safety accountability, and fit to care pathways. We introduce MAWA, a clinician co-designed Multi-Agent Weight-management Assistant combining guideline-grounded, rule-based clinical computation with personalized generation. Ten clinicians define the specification; a rule-based engine computes management modes and clinical targets; specialized agents render section-level advice; an orchestrator integrates and compresses the report; and critic-guided revision with programmatic validation enforces coherence and safety before physician review. On 267 longitudinal cycles from 209 users, we compare MAWA against a content-matched single-prompt baseline receiving the same user data, clinical specification, and report requirements across GPT-5.5, Claude Opus 4.8, and Qwen3-235B. Two cross-family LLM judges (Gemini 3.1 Flash-Lite and DeepSeek-V4-Pro) score six quality dimensions, and six blinded physicians evaluate 12 paired Qwen cases. MAWA improves the LLM-judge-rated six-dimension mean by 0.17, largest on Qwen, driven by internal consistency ($+0.98$; $p<10^{-31}$; Cohen's $d_z=1.52$) and correctness ($+0.17$); actionability decreases slightly ($-0.10$), while personalization, safety, and clarity remain stable. Across backbones, verified arithmetic contradictions fall from $29.9\%$ under the bare prompt to $0\%$ under MAWA. In the Qwen ablation, injection alone reduces report quality; adding agent decomposition raises the mean from 4.00 to 4.31 while shortening reports from 792 to 653 words. Physicians rate both arms as clinically sound and safe, with a descriptive preference for MAWA. MAWA thus delivers more consistent, auditable advice across models under clinician supervision.
Chat is not available.
Successful Page Load