ELLIS Workshop on the Foundations of LLM Post-Training in Changing Environments
Abstract
Large language models (LLMs) are routinely adapted to downstream applications through post-training methods, such as instruction tuning and domain adaption. Yet in real-world deployment, downstream tasks rarely remain fixed: objectives shift, data distributions drift, feedback signals evolve, and evaluation standards change over time. Post-training therefore becomes a process of repeated adaptation in non-stationary environments. Despite its central role in modern foundation models, the theoretical foundations of this adaptive post-training paradigm remain limited. Current practices are largely heuristic, with incomplete understanding of statistical identifiability, optimization dynamics, robustness to misspecification, and trade-offs between adaptation and capability preservation. These gaps are particularly consequential in safety-critical settings, where unintended regressions or feedback loops may arise under evolving conditions. This workshop aims to develop principled foundations for LLM post-training under task evolution. We bring together researchers from machine learning theory, reinforcement learning, and AI safety to develop principled foundations for this adaptive post-training paradigm.