Replacing Prompt-Engineering with Agentic Reflective Meta-training for Medical LLM Deployments
Abstract
Large Language Models (LLMs) are highly sensitive to prompt formulation, making expert prompt engineering a bottleneck for reliable deployment in medicine. Although prior work automates prompt optimization, it remains unclear whether such methods can replace expert engineering, improve already optimized prompts, and remain practical under limited data and compute. We introduce Agentic Reflective Meta-training (ARM), a simple Actor–Validator–Critic framework that iteratively rewrites prompts from task errors and optimization history while rolling back catastrophic updates. Across four clinical information-extraction tasks from orthopedic operative notes, ARM significantly improves over expert-tuned prompts that plateau at 86–93% accuracy. Prompts learned from a minimal instruction perform indistinguishably from expert-initialized prompts, indicating that expert initialization can be bypassed. We further quantify the accuracy–runtime trade-off across training-set and Actor/Critic sizes: 40 labeled examples perform comparably to the full training set, while larger training models increase runtime without improving accuracy. The 8B/8B configuration achieves the highest mean accuracy at the lowest cost, and its prompts transfer effectively to a 70B inference model. On some tasks, ARM reaches the reported 96–100% accuracy of human abstractors. Together, ARM replaces manual prompt engineering while providing a favorable, measurable performance–cost trade-off.