Responsible Communication of Machine Learning Research in Biomedicine
Abstract
A persistent gap has emerged between what machine learning (ML) systems can deliver in biomedical contexts and how their capabilities are communicated to those who use or regulate them. In early discovery, high-profile systems such as AlphaFold and agentic research platforms such as AI Scientists have increasingly accelerated the pace at which unchecked capability claims enter public and policy discourse. In clinical deployment, decision-making tools are expanding rapidly but frameworks for communicating their limitations remain underdeveloped, illustrated by cases from IBM Watson for Oncology's widely documented capability overstatements to more recent findings that LLMs achieving near-perfect medical benchmark scores fail to improve clinical decision-making with real patients. Across the pipeline, this gap drives hype, misuse, misinterpretation and poorly informed governance, with direct consequences for user trust, funding priorities and the effective adoption of advances in ML. Yet in early discovery, frameworks to address this challenge are largely absent; in clinical settings, reporting standards such as TRIPOD+AI and TRIPOD-LLM represent important steps, but adherence remains low. A deeper contributing challenge is that the technical conventions and vocabulary that make findings legible within ML do not translate cleanly across the diverse stakeholders involved: researchers, clinicians, policymakers and science communicators frequently lack a common language and are left uncertain about what ML systems can and cannot do and unable to evaluate their claims. Because the challenges arise in the translation between communities, reporting standards alone or solutions developed by a single community in isolation cannot feasibly close the gap that unintended miscommunication creates. In response, this workshop treats structured, interdisciplinary dialogue as the method. Initiated from within the ML research community, it brings those who produce ML findings into direct exchange with the clinicians, policymakers and science communicators who must interpret and act on them. Previous NeurIPS, ICML and ICLR workshops have advanced related themes primarily from the ML perspective, including explainability, interpretability and responsible AI (RAI). This workshop builds on that foundation by shifting focus from how findings are communicated within ML to how they translate across the biomedical landscape and those who shape it. Grounded in real-world case studies, the workshop is designed to surface opportunities for evidence-based communication approaches when moving from problem to practice, with interdisciplinary exchange at its core.