MeditronFO: An Auditable Pipeline for Clinical LLMs
Abstract
A limitation of current medical LLMs is that none are fully-open, which limits auditability and reproducibility. Most recent medical LLMs release weights, some release data, but none also release the training code and data pipelines and none of them use fully-open model bases. In this work, we ask whether clinically meaningful specialization is possible under strict full-openness constraints, where the data, synthetic generation, preprocessing, training and evaluation pipelines are all inspectable and reproducible. We introduce MeditronFO (Meditron Fully-Open), the first fully-open pipeline for clinical LLM finetuning. The pipeline creates three new synthetic datasets from clinician-authored adversarial vignettes, clinical practice guidelines and public medical QA, using strong open-weight LLMs for rejection-sampling distillation, with generation prompts co-authored by a panel of four clinicians. We apply the pipeline to three fully-open base models (Apertus, OLMo2, EuroLLM) and one open-weight control (Gemma-3), and release them as public artifacts. Our results show that the pipeline halves the gap between open-weight and fully-open medical LLMs, with Apertus-70B-MeditronFO reaching 53.77\% accuracy on medical benchmarks versus 47.18\% for Apertus-70B-Instruct and 60.67\% for MedGemma-27B. We also introduce AutoMOOVE, an open-ended clinician preference benchmark, which is compared on real clinicians preferences from blinded comparisons. Both fully-open Apertus-70B-MeditronFO, and open-weight control Gemma-3-27B-MeditronFO are competitive with MedGemma-27B on this new benchmark.