Skip to yearly menu bar Skip to main content


Workshop

The Third Workshop on Agents in the Wild: Safety, Security, and Beyond

Chenguang Wang ⋅ Wei Wang ⋅ Yuan Xue ⋅ Nicholas Crispino ⋅ Tianneng Shi ⋅ Vincent Siu ⋅ Zhe Ye
Dec 11, 8:00 AM - 5:00 PM Hall 5
AI agents are increasingly deployed in the real world, acting autonomously with access to external tools and persistent memory. However, this same autonomy makes them harder to keep safe and secure, exposing them to new failure modes and adversarial threats that existing safeguards were not built to handle. These risks are no longer hypothetical, making it urgent to address them as agents move from prototypes into systems acting at scale. The Third Workshop on Agents in the Wild brings together researchers across AI, security, systems, policy, and related disciplines to confront the safety and security challenges of agents operating in the real world. We focus on both foundational and emerging challenges through a mix of invited talks, contributed work, and open discussion.
Show more
View full details
Workshop

World Models in Physical AI

Jenny Schmalfuss ⋅ German Ros ⋅ Despoina Paschalidou ⋅ Roberto Martín-Martín ⋅ Jose M. Alvarez ⋅ Grossanchez
Dec 11, 8:00 AM - 5:00 PM MR C4.1
We propose World Models for Physical AI, a one-day NeurIPS 2026 workshop on learned models of physical-world dynamics and their use as the computational substrate for embodied, real-world AI systems. World models have rapidly moved from a research curiosity to a central organizing idea for robotics, autonomous driving, and embodied agents: action-conditioned video generators, latent dynamics models, and generative simulators are now used to train policies, evaluate agents, and even act directly in the physical world. Yet the communities of generative modeling, reinforcement learning, robotics, and computer vision, which are driving this progress, remain fragmented across venues. This workshop brings them together to crystallize the shared problems of evaluation, controllability, long-horizon consistency, sim-to-real, and safety, to chart a research agenda for world models that are actionable, physically grounded, and deployable.
Show more
View full details
Workshop

Machine Learning for Simulations in Biology and Chemistry - The 2nd SIMBIOCHEM Workshop

Bruno Trentini ⋅ Emine Kucukbenli ⋅ Jigyasa Nigam ⋅ Maxim Secor ⋅ Ole Winther ⋅ Runzhong Wang
Dec 11, 8:00 AM - 5:00 PM MR C4.6 & C4.7
Machine learning has achieved landmark progress in structure prediction, generative molecular design, learned potentials, and automated scientific workflows. Yet many models still treat molecular systems as static objects, or rely on physical representations that do not capture conformational ensembles, kinetics, allostery, rare events, thermodynamics, and experimental context. Physics-based simulation provides the mechanistic grounding needed to study these phenomena, but remains too costly for routine use at discovery scale, and a persistent gap remains between computational predictions and experimental reality. The 2nd SIMBIOCHEM Workshop addresses a pressing research challenge: how simulation, generative models, and agentic scientific tooling can be combined into reliable molecular AI systems. Its distinctive emphasis is on molecular simulation as a training signal, mechanistic constraint, and validation layer for models that reason over dynamics, uncertainty, and physical evidence. The workshop will bring together machine learning, computational chemistry, biophysics, life science, and materials science communities to discuss foundation models for molecular systems, simulation-derived post-training, learned force fields, differentiable and enhanced molecular dynamics, uncertainty-aware prediction, and agents that plan simulations, call MD or QM tools, evaluate outputs, and guide the next computational or experimental step.
Show more
View full details
Workshop

NeurIPS 2026 Workshop on Tackling Climate Change with Machine Learning

Utkarsha Agwan ⋅ John Duncan ⋅ Kim Bente ⋅ Rajanie Prabha ⋅ Jonathan Richetti ⋅ Chen Chen ⋅ Yuanyuan Shi ⋅ David Rolnick ⋅ Utkarsha Agwan
Dec 11, 8:00 AM - 5:00 PM Cockle Bay Room 2
Climate change is a complex, multifaceted, and far-reaching challenge with increasingly severe consequences for humanity, as natural disasters multiply, sea levels rise, and ecosystems falter. Climate action takes many different forms and includes both climate change mitigation, e.g. designing smart electric grids or tracking greenhouse gas emissions through satellite imagery, and climate change adaptation, e.g. building flood resilience. Machine learning is a valuable tool in mitigating and adapting to climate change, and at the same time, climate change has been noted as a valuable area for inspiring the development of cutting-edge machine learning algorithms. However, using machine learning to address climate change requires close interdisciplinary collaboration among various fields with diverse practitioners. This workshop is one of a series of workshops intended to form connections and foster cross-pollination between researchers in machine learning and experts in complementary climate-relevant fields, in addition to providing a forum for those in the machine learning community who wish to tackle climate change.
Show more
View full details
Workshop

NEmo: Neuro-Symbolic Embodied Intelligence

Elena Umili ⋅ Emanuele Musumeci ⋅ Vincenzo Suriani ⋅ Daniil Dobriy ⋅ Anna Sofia Lippolis
Dec 11, 8:00 AM - 5:00 PM MR E5.4 & 5
Recent advances in embodied AI increasingly rely on the integration of neural perception, symbolic reasoning, and large language models (LLMs) to support robust decision-making in complex environments. Yet, significant challenges remain in enabling agents to operate reliably over long horizons, revise their internal knowledge structures from experience, and maintain interpretable and transferable representations of the world. This workshop explores emerging neuro-symbolic approaches for embodied intelligence, with a focus on long-horizon planning, self-evolving agents, trustworthy integration of LLMs, and reliable knowledge management. Particular attention is devoted to the development of interoperable semantic representations that connect perception, affordances, planning, and execution across robotic systems. By bringing together researchers from neuro-symbolic AI, robotics, machine learning, knowledge representation, and embodied reasoning, the workshop aims to identify open challenges, share methodologies, and foster a research agenda for scalable and trustworthy embodied neuro-symbolic systems.
Show more
View full details
Workshop

AI for Stochastic Dynamics: From Theoretical Foundations to Scientific Applications

Dai Shi ⋅ Andi Han ⋅ Bingxin Zhou ⋅ Luke Thompson ⋅ Valentin De Bortoli ⋅ Junbin Gao ⋅ José Miguel Hernández-Lobato
Dec 11, 8:00 AM - 5:00 PM MR C3.2
Stochastic dynamics provides a mathematical language for modeling systems driven by intrinsic randomness, with applications across fluids, climate, finance, biology, materials, and molecular science. It is also becoming central to modern AI through diffusion models, neural SDEs, stochastic neural operators, probabilistic forecasting, and learning-based control. Despite this progress, there remains a gap between the theoretical foundations of stochastic dynamics and their use in scientific machine learning applications. Key challenges include defining learning objectives aligned with stochastic quantities of interest, incorporating stochastic structure into models and solvers, and validating whether learned systems reproduce the behavior required by the application. This workshop aims to bridge theoretical foundations and scientific applications by bringing together researchers from stochastic analysis, applied probability, numerical simulation, machine learning, and AI-driven scientific domains. Through invited talks, contributed presentations, posters, and open discussion, the workshop will build a focused community around principled, reliable, and scientifically meaningful AI for stochastic dynamics.
Show more
View full details
Workshop

Grounded and Faithful Vision-Language Models for Real-World Deployment

Mozhgan Nasr Azadani ⋅ Yimu Wang ⋅ Milan Ganai ⋅ Jiayuan Mao ⋅ Weiming Zhi ⋅ Elahe Arani ⋅ Krzysztof Czarnecki ⋅ Marco Pavone
Dec 11, 8:00 AM - 5:00 PM MR C2.2 & C2.3
Vision-language(-action) models are rapidly crossing a threshold: from systems that describe the visual world to agents that must act within it. Yet a critical gap persists between what these models appear to understand and what they can reliably and faithfully ground in the underlying visual and physical world. Despite remarkable progress across robotics, autonomous systems, embodied agents, and interactive AI, current systems frequently exhibit failures in grounding reliability, hallucination mitigation, reasoning consistency, and robust behavior under dynamic and uncertain real-world conditions. This workshop brings together researchers and practitioners from academia and industry to advance methods, benchmarks, and systems for grounded and faithful multimodal intelligence. Rather than viewing grounding and faithfulness as downstream properties to optimize after model development, we emphasize them as fundamental principles for building AI systems whose predictions, reasoning, and actions remain aligned with visual and physical reality. By convening researchers across multimodal learning, robotics, embodied AI, autonomous systems, and world models, the workshop aims to establish concrete pathways toward AI systems capable of reliable perception, faithful reasoning, and robust interaction in open-ended real-world environments.
Show more
View full details
Workshop

Interpretability as a Science: Toward Rigorous Foundations for Understanding LLMs

Dhanya Sridhar ⋅ Navita Goyal ⋅ Shruti Joshi ⋅ Patrik Reizinger ⋅ Gemma Moran ⋅ David Klindt ⋅ Hal Daumé
Dec 11, 8:00 AM - 5:00 PM Hall 2
LLM interpretability has attracted increasing attention, driven by questions about what concepts LLMs encode, where they are represented, and how those representations give rise to behavior. These are the right questions, but the field has yet to converge on what valid answers look like or what kind of evidence can establish them. Thus, despite rapid progress, interpretability remains epistemologically fragile. This reflects the need to scientifically ground interpretability research. Interpretability needs what other sciences already have: a principled account of hypothesis formation and testing, formal criteria for what counts as a genuine explanation, rigorous notions of measurement and falsifiability, and experimental designs that can distinguish mechanisms from artifacts. This workshop brings together researchers from diverse disciplinary background spanning interpretability, causal representation learning, neuroscience, mathematics, statistics, and physics to discuss how to establish interpretability as a rigorous empirical science. It will foster structured dialogue with researchers from adjacent fields that have confronted analogous problems to identify frameworks and experimental practices that can be adapted to interpretability. Through keynote talks and facilitated breakout discussions, we will diagnose what interpretability gets right and what it gets wrong, and work toward shared criteria for how its claims can be formalized, tested, and refined, working toward a more scientifically rigorous foundation for interpretability.
Show more
View full details
Workshop

AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models

Neeraj Kumar ⋅ Ruchika Verma ⋅ Jia Wu ⋅ Hamid Tizhoosh ⋅ Joel Saltz ⋅ Katherine Hoadley ⋅ Gabriele Campanella ⋅ Chad Vanderbilt
Dec 11, 8:00 AM - 5:00 PM MR C4.11
AI for oncology is entering a new phase in which the central challenge is no longer simply training large models on digitized pathology slides, but building clinically reliable systems that learn from the full complexity of hospital-scale cancer data. Pathology remains the diagnostic cornerstone of oncology, yet modern cancer care increasingly depends on an interconnected multimodal ecosystem spanning whole-slide histology, immunohistochemistry, special stains, molecular assays, spatial and multiplexed imaging, genomics, clinical text, longitudinal records, and treatment outcomes. The AI at Scale for Clinical Impact (ASCI) workshop will bring together machine learning researchers, computer vision scientists, computational biologists, pathologists, oncologists, clinical informaticians, and industry leaders to define the next generation of scalable AI methods for cancer diagnosis, prognosis, therapy selection, and real-world clinical deployment. Building on the rapid emergence of pathology foundation models and the NeurIPS 2025 Self-supervised Learning for Cancer Pathology Foundation Models competition, this workshop expands the agenda beyond slide-centric representation learning toward multimodal, continually improving, clinically grounded AI systems. Core themes include pathology-aware and biology-aware learning, continual learning from growing institutional archives, data-efficient adaptation for rare cancers and emerging biomarkers, integration of gigapixel images with omics and clinical text, rigorous benchmarking and reporting standards, uncertainty estimation, interpretability, biological discovery, workflow integration, regulatory readiness, and prospective validation. By focusing on the full lifecycle from model development to measurable patient impact, ASCI aims to catalyze a cross-disciplinary research community around a central question: how can AI systems continuously learn from millions of patients and diverse cancer data streams while remaining trustworthy, generalizable, interpretable, and useful in real clinical practice?
Show more
View full details
Workshop

Agentic AI Benchmark and Application for Enterprise Tasks

Atsunori Moteki ⋅ Graham Neubig ⋅ Yonatan Bisk ⋅ Hideo Saito ⋅ Alexandre Drouin
Dec 11, 8:00 AM - 5:00 PM MR C3.4 & C3.5
This workshop focuses on agentic AI, a rapidly evolving and critical field, with a particular emphasis on its application and evaluation in enterprise-level operations. While the emergence of powerful multimodal large language models (LLMs) has made it possible to deploy agentic AI for a wide range of tasks, its implementation in complex enterprise environments remains largely unexplored and under-researched. The primary objective of this workshop is to promote discussion and collaboration aimed at building robust, efficient, and reliable agentic AI technologies for complex and dynamic enterprise operations. This workshop will be beneficial not only for researchers at institutions conducting research on agentic AI benchmarks and applications, but also for those working within enterprises to build and operate systems using agentic AI. We estimate approximately 250+ attendees, based on the success of the inaugural edition held at AAAI-26.
Show more
View full details
Workshop

Managing Agents that Manage Agents: Workshop on Responsible Use of Meta-Agents that Build, Optimize, and Supervise Other Agents

Simon Yu ⋅ Dilara Soylu ⋅ Apurva Gandhi ⋅ Jiuding Sun ⋅ Zichen Liu ⋅ Christopher D Manning ⋅ Weiyan Shi ⋅ We Shi ⋅ Ananjan Nandi ⋅ Derek Chong
Dec 11, 8:00 AM - 5:00 PM Hall 3
Agents now write their own harness and shape their own training. How do we advance these capabilities while ensuring responsible deployment? We define these higher-order systems that build, optimize, and supervise other agents as **meta-agents**. While meta-agents will likely play an increasingly central role in the future of AI, research remains dispersed across disparate communities. This workshop provides a dedicated venue to address both the technical and societal challenges of this emerging field. Technically, the transition to automated agent design requires new optimization methods, learning signals, and meta-level benchmarks to ensure these systems can safely generalize and improve. Societally, as meta-agents take on the manager roles (i.e. optimizing prompts and assigning tasks to downstream worker agents), they require strict governance. A misaligned objective can propagate to every sub-agent, and the potential for agents to manage human labor raises urgent questions about human autonomy. By spanning the full meta-agent lifecycle, from automated design and open-ended evolution to verifiable halt controls and human oversight, this workshop brings together *researchers in AI and organizational science** to ensure these systems are developed and deployed responsibly.
Show more
View full details
Workshop

Diffusion Language Models: Foundations, Efficiency, and Reasoning

Irina Belousova ⋅ Shansan Gong ⋅ Amin Karimi Monsefi ⋅ Pavlo Molchanov ⋅ Jinjie Ni ⋅ Michael Shieh ⋅ Yizhe Zhang ⋅ Yuchen Zhu ⋅ Amin Karimi Monsefi
Dec 11, 8:00 AM - 5:00 PM Parkside 1
Autoregressive (AR) generation has dominated language modeling for years, but a fundamentally different paradigm is gaining rapid traction: diffusion language models (DLMs). Rather than generating tokens left-to-right, diffusion language models corrupt sequences through forward noising over categorical spaces and learn to reverse this corruption, enabling parallel decoding, bidirectional context, and fine-grained controllable generation. In 2025–2026, this paradigm transitioned from theoretical curiosity to commercial reality: Inception Labs launched Mercury, the first commercial-scale diffusion LLM; academic labs released LLaDA and Dream; and academic work—SEDD, MDLM, FS-DFM, LaViDa—has shown diffusion language models can match or exceed AR baselines while achieving up to 10× inference speedups. This full-day workshop brings together researchers from academia and industry to consolidate theoretical understanding, benchmark competing approaches, and chart a roadmap for diffusion language models, catalyzing collaborations across the generative modeling, NLP, and systems communities.
Show more
View full details
Workshop

Workshop for Autonomous Machine Learning Research

Arjun Prakash ⋅ Aditya Iyer ⋅ Jack Liell-Cock ⋅ Hamish Ivison ⋅ Amy Greenwald ⋅ Nora Ayanian
Dec 11, 8:00 AM - 5:00 PM MR C3.6
The research process comprises the formulation of a hypothesis, the design of an experiment, and the judgment of a result. The implicit assumption that this process is a fundamentally human act is suddenly being challenged by autonomous research. We believe the machine learning community should proactively confront this structural change in the research process. We therefore propose the Workshop for Autonomous Machine Learning Research for research that was substantially carried out by autonomous AI agents. To ensure that conferences continue to exist for both science and scientists, our proposal anchors this new track in human judgment and participation. In keeping with current conventions, we assert that authorship remains exclusively human while recognizing that the role of author may shift to one of curation of autonomously generated research. Therefore, we propose a discussant-style format, where both a human author and a human reviewer present the accepted research. This design assigns visible credit to human participants for thoughtful evaluation and judgment, which is essential for both scientific excellence and maintaining a sense of community.
Show more
View full details
Workshop

NeurIPS 2026 Workshop on Dynamic Alignment in Human-AI Coupled Systems

Hua Shen ⋅ Divy Thakkar ⋅ Gordon Dai ⋅ Vivek Myers ⋅ Nick Haber ⋅ Joan Bruna ⋅ Dawn Song ⋅ Yoshua Bengio
Dec 11, 8:00 AM - 5:00 PM Cockle Bay Room 1
Large language models, multimodal models, and agentic AI systems are increasingly becoming long-term interactive actors in education, healthcare, scientific discovery, robotics, and other high-stakes domains. In these settings, AI systems shape human trust, dependence, values, and behavior, while human feedback and behavior in turn shape agents’ objectives, reward signals, and policy updates. This workshop introduces **Dynamic Alignment in Human–AI Coupled Systems** as a research agenda for studying alignment as an evolving property of coupled human–AI systems, rather than a static property of AI models alone. Bringing together researchers across machine learning, AI safety, human–AI interaction, cognitive science, social computing, governance, and philosophy, the workshop will examine how human states, agent policies, feedback signals, objectives, and evaluation criteria co-evolve over time. Through keynotes, panels, papers, posters, and structured discussions, the workshop aims to advance foundations, evaluations, and interventions for modeling, measuring, and controlling alignment dynamics, and to catalyze an interdisciplinary agenda for safer human–AI futures.
Show more
View full details
Workshop

Robot Learning with World Models: Capabilities, Frontiers, and Challenges

Kuang-Huei Lee ⋅ Hiroki Furuta ⋅ Homanga Bharadhwaj ⋅ Wenhao Yu ⋅ Joycelyn (Jing-Wen) Chen ⋅ Ruoshi Liu ⋅ Zeyi Liu ⋅ Yifu QIU ⋅ Kuang-Huei Lee
Dec 11, 8:00 AM - 5:00 PM Parkside 2
Despite early successes of world models on robot learning, significant challenges persist in physical accuracy, spatiotemporal consistency, and the lack of critical modalities beyond vision (e.g., tactile sensing, proprioception). This workshop aims to bring together researchers from generative modeling, computer vision, and robotics to exchange ideas on (1) the state-of-the art of robot learning with world models (where are we now?), and (2) the open challenges and gaps (where should we go?). Through invited talks, panel, oral presentations and poster sessions, we will discuss topics including world action models, improving physical accuracy, non-vision modalities, evaluation metics, and learning and planning in imagination. The workshop welcomes contributions ranging from short papers, full papers, demos, and networking group proposals. We hope to create an inclusive venue for active idea exchanging and community building among people with diverse backgrounds.
Show more
View full details
Workshop

Agentic AI for Biological Discovery: Toward Closed-Loop Life-Science Intelligence

Mengdi Wang ⋅ Le Song ⋅ Masatoshi Uehara ⋅ Xiner Li ⋅ Ehsan Hajiramezanali ⋅ Philipp S. L. Schäfer
Dec 11, 8:00 AM - 5:00 PM Grand Ballroom B1
The life sciences are entering an era of agentic AI, systems that go beyond static prediction to read literature, call specialized tools, plan multi-step analyses, propose experiments, and in some cases interact directly with laboratories and robotics. This shift is enabled both by frontier large language models and by a rapidly maturing stack of biology-specialized foundation models for protein structure, protein design, genomes, and single cells. Yet the field is strikingly young: there is little consensus on how to build an effective life-science agent (harness design, multi-agent orchestration, memory, tool ecosystems), how to deploy and evaluate one in high-stakes settings such as drug discovery and medicine, or when biology-specialized models are needed versus when general-purpose LLMs already suffice. This workshop brings together researchers from machine learning, computational biology, experimental biology, drug discovery, and lab automation around four open questions: (i) generalist vs.\ specialist agents; (ii) systems-design for building and deploying agents; (iii) lab-in-the-loop integration with automation, robotics, and human scientists; and (iv) evaluation, reliability, and safety of scientific agents. We solicit contributions across two tracks: Building Agentic Systems for Life Science (architectures, harness design, literature agents, alignment, RL) and Closed-Loop Discovery and Applications (autonomous labs, biological design, single-cell and multi-omics agents, benchmarks, biosecurity).
Show more
View full details
Workshop

Physical Understanding for Decision-Making: Bridging Foundation Models and Reliable Agents

Felix Juefei-Xu ⋅ TIANYU SHI ⋅ Mengyue Yang ⋅ Rong Zou ⋅ Yuke Zhu ⋅ Michael Rabbat ⋅ Steven Waslander
Dec 11, 8:00 AM - 5:00 PM Grand Ballroom B2 & B3
Classical control and planning approaches—such as model predictive control, trajectory optimization, and physics- based simulation—addressed this challenge by encoding physics and mathematical constraints explicitly, and achieved remarkable precision in structured environments. However, these methods do not scale gracefully to the diversity, partial observability, and contact-richness of real-world deployments, and they require painstaking manual modeling of every relevant physical relationship. The recent emergence of foundation models for decision-making—vision-language-action (VLA) models, robot foundation models, agentic systems, and interactive world models—offers a compelling alternative. These systems have sharply improved long-horizon video generation, embodied policy learning, and multimodal action grounding, building on a long tradition of world models for prediction and planning. Yet most of these models are still trained and evaluated primarily for visual realism or short-term prediction accuracy. For embodied agents, this gap is critical: a robot that cannot reason about the outcomes of its actions will fail on contact-rich manipulation, long-horizon planning, and safety-critical deployment. The field now stands at an inflection point where foundation and world models are powerful enough to serve as the backbone of physical agents, yet the scientific community lacks shared definitions, evaluation protocols, and benchmarks for physical understanding as a distinct and measurable capability. This workshop is designed to fill that gap before incompatible paradigms and vocabularies become entrenched.
Show more
View full details
Workshop

Collaborative, Open, and Decentralized Training of Foundation Models

Thalaiyasingam Ajanthan ⋅ Sameera Ramasinghe ⋅ Benjamin Thérien ⋅ Anastasia Koloskova ⋅ Kaja Gruntkowska ⋅ Eugene Belilovsky ⋅ Aakanksha Chowdhery ⋅ Nicholas Lane ⋅ Aj
Dec 11, 8:00 AM - 5:00 PM MR C4.4
Training large-scale foundation models today depends on massive, centralized GPU clusters that are inaccessible to most academic institutions, startups, and industries. This concentration of compute creates high barriers to entry, centralizes AI innovation, and limits broader progress in developing and studying frontier-scale foundation models. Open models are key to democratizing the know-how of frontier-model training. However, the cost of centralized training infrastructure largely excludes the broader research community from participating in open development at scale, leaving scaling efforts dependent on substantial centralized resources and offering the community limited opportunity to contribute to the training process itself. Decentralization and resource pooling provide a way to enable large-scale runs without this barrier by enabling model training across geographically distributed and heterogeneous devices, from coordinated inter-datacenter settings to consumer-devices connected over internet. This workshop will bring together researchers and practitioners to address the core technical challenges of this paradigm, including communication efficiency, asynchronous optimization, heterogeneous systems, fault tolerance, and security. Addressing these challenges enables large-scale foundation model training beyond the confines of a single datacenter, broadens participation in foundation model research, and provides complementary support for more scalable and collaborative open model development.
Show more
View full details
Workshop

AI for Drug Discovery: Bridging the Translation Gap

Ivor Tsang ⋅ Yu Xie ⋅ Sai Ho Ling ⋅ Yao Yinghua
Dec 11, 8:00 AM - 5:00 PM MR C3.3
A widening chasm has emerged between in-silico optimization and real-world therapeutic validation. Current deep learning models routinely claim state-of-the-art performance on static, curated academic benchmarks; however, these metrics are frequently artifacts of shortcut learning and selection biases inherent to public datasets. When deployed in production-grade drug discovery environments - characterized by severe data scarcity, unaligned modalities, and non-stationary biological drift—these models exhibit severe fragility. This workshop moves the community beyond superficial leaderboard chasing. We explicitly solicit papers that algorithmically diagnose benchmark vulnerabilities, develop provably robust and sample-efficient architectures, and design interactive AI Co-Scientists capable of safely steering generative models through data-starved biological regimes.
Show more
View full details
Workshop

I Can’t Believe It’s Not Better (ICBINB): Failure Modes of AI in Biology

Peter Koo ⋅ Maria Brbic ⋅ Su-In Lee ⋅ Bianca Dumitrascu ⋅ Siba Smarak Panigrahi ⋅ Masayuki Nagai ⋅ Soham Gadgil ⋅ Ozgur Beker
Dec 11, 8:00 AM - 5:00 PM MR C4.8
Artificial intelligence (AI) is rapidly transforming biology, enabling progress in gene regulation modelling, cellular response prediction, and protein structure determination. Despite strong performance on curated benchmarks, AI models often fail when deployed beyond controlled experimental settings. This gap arises from fundamental properties of biological systems, including inter-individual heterogeneity, substantial measurement noise, limited labelled data, weak or confounded ground truth, and intrinsic biological complexity, such that benchmark success does not reliably translate to real-world biological reliability. The workshop I Can’t Believe It’s Not Better (ICBINB): Failure Modes of AI in Biology will bring together the machine learning and life sciences communities to systematically document, analyze, and learn from negative results and real-world failure modes across genomics, transcriptomics, structural biology, and clinical prediction. In addition, we will emphasize evaluation beyond curated benchmarks toward deployment-relevant assessment, together with methodological advances that address these limitations, including robustness under distribution shift, interpretable and mechanistic modelling, causal learning, uncertainty quantification, and adaptive learning strategies that improve reliable generalization and trustworthy deployment. These challenges directly reflect core questions in modern machine learning concerning robustness, causality, interpretability, uncertainty, and evaluation in real-world settings. By centring real-world failure analysis in biologically consequential settings, the workshop aims to clarify the gap between benchmark performance and scientific reliability, establish principled evaluation perspectives for trustworthy discovery, and advance more reliable and scientifically meaningful AI systems in biology.
Show more
View full details
Workshop

Symmetry and Geometry in Neural Representations NeurIPS Workshops 2026

Francisco Acosta ⋅ David Klindt ⋅ Louisa Cornelis ⋅ Alison Pouplin ⋅ Nina Miolane ⋅ Fatih Dinc
Dec 11, 8:00 AM - 5:00 PM Hall 4
The fields of biological and artificial intelligence are increasingly converging on a shared principle: the mathematical structure of real-world tasks plays a central role in building efficient, robust, and interpretable representations. In neuroscience, mounting evidence suggests that neural circuits encode structure through low-dimensional manifolds, conserved symmetries, and structured transformations. In deep learning, principles such as sparsity, equivariance, and compositionality are guiding the development of more generalizable and interpretable models. The NeurReps workshop brings these threads together, fostering dialogue among machine learning researchers, neuroscientists, and mathematicians to uncover unifying geometric principles of neural representation. Following successful editions in 2022–2025 with over 160 submissions and 500 attendees last year, NeurReps 2026 expands into emerging frontiers including geometric representation steering and geometric mechanistic interpretability, and introduces a new Findings track to foster collaboration between experimentalists and theorists.
Show more
View full details
Workshop

Foundation Models for Temporal Systems: From Forecasting to World Modeling

Boris Oreshkin ⋅ Danielle Maddix Robinson ⋅ Omri Azencot ⋅ Ming Jin ⋅ Emadeldeen Eldele ⋅ Mayank Jauhari ⋅ Chenghao Liu ⋅ N. Benjamin Erichson ⋅ Mayank Jauhari Iitr
Dec 11, 8:00 AM - 5:00 PM MR C4.5
We propose a one-day in-person workshop at NeurIPS 2026 on foundation models for temporal systems, organized around temporal world modeling: how time-series models should be trained, evaluated, and deployed when real-world temporal data is multimodal, asynchronous, event-driven, and multi-scale, and when models must reason beyond extrapolation under distribution shift. Progress on these settings is fragmented across the forecasting, foundation-model, multimodal, and generative-modeling communities; the workshop unifies them around four directions: (i) forecasting and simulation tasks, (ii) temporal data and environments, (iii) temporal models, and (iv) evaluation and reliability, with applications in climate, healthcare, and industrial forecasting. The program features nine confirmed invited speakers spanning academia and industry, contributed papers and posters managed through OpenReview, a moderated panel, and an open community discussion, with approximately 200-300 in-person attendees anticipated. Distinct from djacent forecasting and foundation-model workshops, it emphasizes generative simulation, counterfactual rollouts, action-conditioned prediction, and long-horizon trajectory consistency: what temporal world modeling adds beyond forecasting.
Show more
View full details
Workshop

OPT 2026: Optimization for Machine Learning

Frederik Kunstner ⋅ Michael Crawshaw ⋅ Courtney Paquette ⋅ Fred Roosta ⋅ Sebastian Stich ⋅ Chulhee Yun
Dec 11, 8:00 AM - 5:00 PM Hall 1
This year marks the 18th edition of the long-running NeurIPS Workshop on Optimization for Machine Learning (OPT). OPT 2025 was a tremendous success, attracting over 160 paper submissions, a 30% increase over the previous year, and featuring 5 plenary speakers, 6 contributed talks, and 2 highly attended poster sessions with more than 400 participants. Widely regarded as one of the premier venues at the intersection of optimization and machine learning, the workshop has become a key forum for presenting cutting-edge research and fostering collaboration across the community. In particular, it serves as an important venue for early-career researchers, many of whom view OPT as the leading workshop for frontier research in optimization and machine learning. Over its 18-year history, OPT has played a central role in bringing together researchers from diverse backgrounds, catalyzing collaborations that have gone on to produce influential research at top-tier venues. Its impact on shaping the optimization and machine learning communities cannot be overstated.
Show more
View full details
Workshop

AgenticOS: Co-designing Systems and ML Foundations of an OS Layer for Agentic AI

Suparna Bhattacharya ⋅ Ian Foster ⋅ Jishen Zhao ⋅ Cong Wang ⋅ Tarun Kumar
Dec 11, 8:00 AM - 5:00 PM MR C2.5 & C2.6
Today’s agentic stacks are characterized by framework proliferation without common foundations. Each orchestration harness independently reimplements cross-cutting services (state management, memory, context engineering, resource budgeting, tool orchestration, and safety enforcement) making agents non-portable, difficult to audit, and brittle under composition. Simultaneously, self-evolving agents—systems that learn, adapt, and update their own behavior from operational experience—are proliferating in research and production. This introduces an entirely new class of challenge: how do we enable useful, constrained self-evolution while maintaining safety, reproducibility, and resource bounds? Current approaches leave this largely to individual frameworks, with no shared abstraction or principled mechanism for governing what changes, when, and under what conditions. The security implications of self-modifying systems operating autonomously at scale are underexplored and constitute a pressing open problem. Just as pre-OS computing was rescued not by better programs but by a shared execution substrate, the resolution to agentic fragmentation is not a better framework but an OS layer. Designing this OS layer requires genuine co-design between the ML and systems communities. The ML community must ask how models should be trained, structured, and exposed as system components—not merely as API endpoints. The systems community must ask what abstractions, memory hierarchies, scheduling policies, and execution substrates are required to make agentic behavior reliable and governable. These are not separable questions: the right abstraction boundary depends on what models can learn to do, and what models need to learn depends on what the system layer can support.
Show more
View full details
Workshop

ML for Systems

Dan Zhang ⋅ Xinlei XU ⋅ Mangpo Phothilimthana ⋅ Divya Mahajan ⋅ Haoran Qiu ⋅ Patrick Musau
Dec 12, 8:00 AM - 5:00 PM MR C4.5
Machine Learning (ML) for Systems applies machine learning techniques to computer systems. With the rise of large language models (LLMs) and generative AI agents, ML has the potential to revolutionize the entire hardware and software stack, replacing long-standing heuristics and even the process in which these systems are designed and implemented. This has led to a paradigm shift in systems design, such as automating multi-objective tasks including designing new data structures [1], integrated circuits [2, 3], and design verification [20, 21], implementing control algorithms for applications including compilers [12, 13, 19], databases [8], operating systems [37], memory management [9, 10], cloud platform orchestration [33, 34], and ML training frameworks [11]. The rise of LLMs and generative AI agents has presented new opportunities and challenges within the diverse domain of computer systems. For the 10th edition of the ML for Systems workshop, we are shaping the program around three key pillars: (1) Pushing the frontier: demanding mature, scalable, and cost-efficient research on LLMs/agents for systems problems that moves the community beyond one-off demonstrations; (2) Agents for cybersecurity and agentic security: Introducing a critical new focus on security and reliability as LLM-generated code and autonomous agents increasingly enter production systems, to have a grounded understanding besides lay reports on "Mythos solved security"; and (3) Charting the future: Hosting a 10-year retrospective and forward-looking panel to define the next generation of systems intelligence.
Show more
View full details
Workshop

On-Device Intelligence: Foundation Models under Real-World Constraints

Niao He ⋅ Bingcong Li ⋅ Shiwei Liu ⋅ Hao Ma ⋅ Michael Muehlebach ⋅ Daniela Rus ⋅ Marko Zaric ⋅ Melanie Zeilinger
Dec 12, 8:00 AM - 5:00 PM MR C3.6
Intelligence should not rely solely on the cloud. As foundation models enter real-world interactive and embodied settings, their cloud-centric capabilities must be realized on devices under constraints on latency, memory, reliability and other factors. From LLMs to embodied foundation models, the key question is no longer only how to improve capability, but how to make it usable for efficient on-device inference and reasoning, stable interaction, continual adaptation, and safe execution. On-device intelligence is therefore not merely model compression or deployment, but a tightly coupled research problem spanning theory, algorithms, systems, and evaluation, where capability must be considered jointly with efficiency, adaptation, execution, and reliability under real-world constraints. This workshop will bring together researchers working on LLMs, embodied intelligence, and edge systems to clarify this shared problem space and identify key questions through invited talks, panels, contributed posters, and open discussions, helping connect separated communities and shape a shared research agenda for on-device intelligence.
Show more
View full details
Workshop

XAI4science: Knowledge Discovery and Trust through Interpretable Foundation Models

Leonardo Pesce ⋅ Jiawen Wei ⋅ Max Welling ⋅ Gianmarco Mengaldo
Dec 12, 8:00 AM - 5:00 PM Cockle Bay Room 1
Climate change represents an existential race that humanity cannot afford to lose. Through accurate weather forecasting models and Environmental Sciences, researchers make informed choices towards a more sustainable future. While traditional models require immense computational resources, as more data is collected, Machine Learning (ML) has proven to be a viable and efficient solution for it. Weather and Climate forecasting Foundation Models (FM) have shown that comparable results can be achieved with a fraction of the resources, opening the field to smaller countries and research teams with less investment and representation. While this provided a solution to the “efficiency” crisis in climate modeling, it introduced a “transparency” crisis. In fact, these black-box models pose a significant threat to scientific trust and social equity, impeding wider adoptions, such as in extreme weather event preparation and intervention, or towards sustainable initiatives and global net-zero projects. This workshop focuses on two converging crises: the technical challenge of extracting physically consistent explanations from high-dimensional FM and the risk that non-transparent models will reinforce existing inequalities through biased resource allocation. Therefore, by integrating eXplainable Artificial Intelligence (XAI) in the modeling pipeline, we aim at reducing the interpretability gap and bias of FMs. By bringing together ML researchers, climate scientists, and policy experts, we aim to investigate the fundamental roles of interpretable architectures in the ongoing challenge of extreme weather events and towards a low-carbon society for a better future.
Show more
View full details
Workshop

ATTRIB: Workshop on Data Attribution and Provenance

Sarah Cen ⋅ Andrew Ilyas ⋅ Bálint Mucsányi ⋅ Elisa Nguyen ⋅ Sam Park ⋅ Theodora Worledge ⋅ Theodoraworledge
Dec 12, 8:00 AM - 5:00 PM MR C4.1
As generative AI systems become widely deployed, a central challenge is attribution: determining how model outputs relate to data. Existing approaches span contributive attribution, which estimates the causal influence of training data on model behavior, and corroborative attribution, which identifies sources that semantically support or correspond to model outputs. While both paradigms have advanced significantly, their adoption in real-world contexts (such as copyright disputes, regulatory audits, safety investigations, and output verification) remains limited. This gap reflects a mismatch between current technical outputs, such as influence scores or semantic matches, and the forms of attribution required in practice, which must be interpretable, auditable, and grounded in identifiable sources. The rapid growth of synthetic and model-generated data further complicates attribution by blurring distinctions between training data and outputs. This workshop brings together researchers and practitioners across machine learning, law, journalism, and related domains to examine the limitations of existing attribution methods, clarify practical requirements, and identify research directions that can enable reliable, actionable attribution in deployed AI systems.
Show more
View full details
Workshop

Real-Time Conversational Agents: Toward Natural Multimodal Interaction

Niki Foteinopoulou ⋅ Alessandro Conti ⋅ Jack Saunders ⋅ Oya Celiktutan ⋅ Cigdem Beyan ⋅ Ioannis Patras ⋅ Jack Saunders
Dec 12, 8:00 AM - 5:00 PM MR C4.11
Real-time conversational agents have rapidly moved from research demonstration to deployed product, with voice modes, embodied avatars, and full-duplex speech systems now powering applications across customer service, education, healthcare, and accessibility. To feel natural, however, such systems cannot rely on the offline generation paradigm that has dominated multimodal machine learning: they must produce speech, video, and language in a streaming fashion while continuously listening, watching, and re-planning over partial observations. The methodological constraints introduced by this regime, namely sub-second latency budgets, causal and incremental computation, full-duplex audio modelling, tight cross-modal temporal alignment, and handling of interruptions, backchannels, and overlapping speech, differ qualitatively from those addressed by the offline literature, and techniques that have driven offline progress (non-causal attention, large-beam decoding, multi-pass refinement, slow diffusion sampling) frequently fail to transfer. Despite recent advances in full-duplex audio–language modelling, real-time talking-head and avatar synthesis, low-latency speech generation, and streaming automatic speech recognition, deployed agents remain perceptually robotic: turn-taking is stilted, backchannels are absent, prosody is monotone, and gaze and gesture are routinely mis-timed with respect to the linguistic and affective content of the utterance. The community has not yet converged on a shared vocabulary, benchmarks, or methodology for evaluating interactional naturalness as distinct from per-utterance quality measured by mean opinion scores or offline win-rates. The Workshop on Real-Time Conversational Agents (RTCA) addresses this gap by convening researchers across speech, vision, language, human–computer interaction, social signal processing, and machine learning systems around three intertwined questions. First, real-time generation: how to produce high-quality speech, video, and language under hard latency budgets, in a streaming or full-duplex fashion, with attention to the architectural, training, and inference-system innovations that distinguish streaming from offline modelling. Second, naturalness in interaction: what perceptual, linguistic, and behavioural elements prosody, gaze, timing, grounding and expressivity. Third, evaluation of live systems: how to design metrics, benchmarks, and study protocols that capture naturalness, responsiveness, and conversational quality in interactive settings, where standard offline metrics and held-out test sets are demonstrably inadequate. The workshop solicits short papers, full-papers, and demo papers on streaming speech synthesis and recognition, full-duplex audio–language models, real-time talking-head and embodied avatar generation, incremental and speculative decoding, turn-taking and floor management, multimodal alignment under partial observation, prosody and paralinguistic generation, memory and grounding in live conversation, interactive evaluation protocols, efficient inference and on-device deployment, and the safety and identity considerations specific to real-time multimodal generative agents. The programme combines invited talks across the workshop's four thematic pillars with contributed talks and posters, a live Conversational Agents Showcase in which accepted demonstration systems are run on stage, and a closing panel on what naturalness in conversational AI actually means and how it should be measured. By foregrounding evaluation alongside generation and by convening communities that currently publish across separate venues, the workshop aims to consolidate shared problems, datasets, and methodological standards for a research area whose downstream user impact already substantially exceeds its academic visibility.
Show more
View full details
Workshop

AI Foundations for Power Grids: From Models to Deployment at Scale

Andrea Britto Mattos Lima ⋅ Wenqi Cui ⋅ Nicolas Christianson ⋅ Christopher Yeh ⋅ Rabab Haider ⋅ Baosen Zhang ⋅ Thomas Brunschwiler
Dec 12, 8:00 AM - 5:00 PM MR C2.5 & C2.6
The power grid presents a compelling yet underexplored domain for machine learning, combining hard physical constraints, real-time operation, and large-scale societal impact. Despite increasing interest in applying learning-based methods to problems such as optimal power flow, contingency analysis, and system control, evaluation practices have not kept pace with methodological advances. Many existing studies rely on small, static benchmarks and in-distribution metrics, which fail to capture the challenges of real-world deployment in evolving and safety-critical environments. This workshop aims to position power systems as a first-class methodological challenge for machine learning, with a central focus on evaluation under realistic operating conditions. We will bring together researchers from machine learning and power systems to discuss datasets, benchmarks, model classes, and validation protocols needed to support robust and trustworthy deployment. The program includes invited talks from academia and industry, contributed papers, panel discussions, and interactive sessions designed to foster collaboration across communities. By emphasizing rigorous evaluation, distributional robustness, and system-level behavior, the workshop seeks to shape the development of machine learning methods for infrastructure systems and to establish shared standards for future research.
Show more
View full details
Workshop

PTA: From Pretrained Representations to Acting Agents -- Bridging Pretraining, Planning, and Test-Time Decision Making

Ping-Chun Hsieh ⋅ Kuang-Huei Lee ⋅ Yen-Ling Kuo ⋅ Bo Dai ⋅ Georgia Chalvatzaki ⋅ Karen Leung ⋅ Co Yong ⋅ Claas A. Voelcker
Dec 12, 8:00 AM - 5:00 PM Parkside 2
Large-scale pretraining has become a dominant paradigm across machine learning, providing foundation models for language, vision, robotics, multimodal interaction, and decision-support systems. Yet as pretrained models are increasingly deployed in sequential decision-making settings, one central challenge remains: how can pretrained representations be learned, aligned with acting agents, and transformed into reliable, adaptive, and controllable behavior at test time? The PTA workshop—From Pretrained Representations to Acting Agents—will bring together researchers from representation learning, reinforcement learning, robotics, world modeling, planning, test-time scaling, and adaptive control to study this emerging post-pretraining decision interface. The workshop will focus on what makes representations actionable: how they encode structure useful for prediction, planning, control, adaptation, uncertainty estimation, and generalization, and how agents can reliably use such knowledge under distribution shift and changing task demands. It will examine both the learning of action-relevant abstractions and the mechanisms that convert pretrained knowledge into robust decisions during deployment. By emphasizing the shared problem of learning representations that are not merely predictive but actionable, PTA aims to clarify common principles, expose open challenges, identify benchmark needs, and accelerate progress toward agents that can reliably reuse, adapt, and deploy pretrained knowledge in dynamic environments.
Show more
View full details
Workshop

Transitioning from Pre-training to Post-training

Rachit Bansal ⋅ Clara Mohri ⋅ Tian Qin ⋅ Harman Singh ⋅ Samy Jelassi ⋅ Sham Kakade
Dec 12, 8:00 AM - 5:00 PM Cockle Bay Room 2
This workshop examines the interaction between pre-training and post-training in modern foundation models. We solicit theoretical, empirical, and methodological work on how pretraining choices determine the success, or failure, of post-training methods including, but not limited to, supervised fine-tuning, RLXF, RLVR, self-improvement, and distillation. We also solicit work studying on how these post- training procedures reshape the base model. The goal is to crystallize the science of the pretraining-to-post-training transition.
Show more
View full details
Workshop

Foundation Models for the Brain and Body

Mehdi Azabou ⋅ Nanda H Krishna ⋅ Pierre Guetschel ⋅ Jonathan McCart ⋅ Alexandre Andre ⋅ Melanie Segado
Dec 12, 8:00 AM - 5:00 PM Hall 1
Our brains and bodies speak a rich and complex biological language of neural and physiological signals, a language that AI models are increasingly capable of deciphering as large-scale datasets become available. Recent advances in neural technology, including EEG, intracortical electrophysiology, fMRI, EMG, MEG, and ECG, have enabled the broad collection of biosignals across real-world contexts and diverse populations. This growing wealth of data is driving a shift toward foundation models: large-scale, pretrained AI systems designed to consume incredibly large and diverse datasets with the goal of generalizing across diverse downstream applications, from brain-computer interfacing to health monitoring and robotics. Realizing this potential, however, requires addressing the unique challenges that come with the recordings available from these modalities: they are noisy and heterogeneous timeseries that were collected under variable conditions across subjects, devices, and environments. To get truly generalist foundation models we need to meet these challenges. To this end, this workshop brings together neuroscientists, biomedical engineers, wearable tech researchers, and machine learning experts advancing foundation model approaches. Through interdisciplinary dialogue, we aim to catalyze the next generation of AI models that can capture the complexity of the brain, body, and behavior at scale.
Show more
View full details
Workshop

AI4Mat-NeurIPS-2026: NeurIPS-2026 Workshop on AI for Accelerated Materials Design

Santiago Miret ⋅ Mara Schilling-Wilhelmi ⋅ N M Anoop Krishnan ⋅ Vijay K Narasimhan ⋅ Stefano Martiniani
Dec 12, 8:00 AM - 5:00 PM Hall 2
AI4Mat-NeurIPS-2026 explores applications of AI to materials via: (1) AI-Guided Materials Design; (2) Automated Chemical Synthesis; and (3) Automated Material Characterization. As the focused meeting point for the AI-for-materials community at NeurIPS, the workshop emphasizes structured, expert-driven dialogue on making machine learning more impactful for real-world materials discovery. As reasoning models mature and self-driving laboratories come online, a timely question arises: how does compute translate into genuine scientific discovery, and how do we close the gap between algorithmic progress and deployment in messy, real experiments? Our two main sessions address this: Scaling Laws for Materials Reasoning: From Compute to Scientific Discovery examines reasoning models and the utility of compute; Automating Discovery That Delivers: When AI Meets the Messiness of Real Experiments tackles the critical challenges of real-world experimental data collection. Beyond invited talks, a substantial portion of the program is devoted to contributed spotlights, a poster session, and a town hall. Building on previous AI4Mat workshops, we strengthen this community through in-depth peer feedback on spotlight presentations, an expanded travel grant program, and a dedicated focus collection in a high-impact journal. Our location preference is Sydney to broaden the reach of the AI4Mat community.
Show more
View full details
Workshop

Can We Trust AI Evaluation? Robustness, Causality, and Risk in Modern AI Assessment

Sarah Erfani ⋅ Paolo Giudici ⋅ Xingjun Ma ⋅ Eduard Hovy ⋅ Hanxun Huang
Dec 12, 8:00 AM - 5:00 PM MR C2.2 & C2.3
AI capabilities are advancing rapidly, yet our ability to evaluate these systems has not kept pace. Benchmarks, leaderboards, and aggregate metrics increasingly influence decisions about model selection, deployment, regulation, and investment, but growing evidence shows that evaluation conclusions can be fragile, misleading, or insufficient for real-world use. Performance may vary under small evaluation changes, benchmark reuse can induce overfitting, contamination can distort comparisons, and offline metrics often provide limited evidence about safety, reliability, and deployment outcomes. As foundation models, AI agents, and autonomous systems are increasingly deployed in high-stakes settings, evaluation is becoming a critical bottleneck for trustworthy AI adoption. This workshop is motivated by a central question: When is AI evaluation evidence strong enough to guide deployment decisions? We argue that the next decade of AI research requires a science of AI evaluation – a research agenda that studies evaluation protocols themselves, not only the models being evaluated. The goal is to develop principled foundations for determining what an evaluation measures, what assumptions it relies on, what uncertainty remains, and when its conclusions can be trusted. The workshop will bring together researchers and practitioners from machine learning, statistics, causal inference, robustness, AI safety, and high-risk application domains to address three interconnected challenges: (1) uncertainty and robustness of evaluation evidence, (2) benchmark and leaderboard auditing, and (3) deployment risk and decision relevance. Through invited talks, contributed papers, posters, and panel discussions, the workshop aims to catalyze a research community around trustworthy AI evaluation and help establish the methodological foundations needed for reliable, transparent, and accountable deployment of increasingly capable AI systems.
Show more
View full details
Workshop

AI for Science: Verification in the Age of AI Scientists

Ada Fang ⋅ Yuanqi Du ⋅ Ana Rivera Him ⋅ Anvita Bhagavathula ⋅ Emilien Dupont ⋅ Priya L Donti ⋅ Marinka Zitnik
Dec 12, 8:00 AM - 5:00 PM Grand Ballroom B1
AI Scientists are rapidly changing the pace and structure of scientific discovery. Emerging systems can generate hypotheses, design experiments, write papers, and propose scientific decisions at a scale that increasingly exceeds human capacity for manual review. Yet across domains, the ability to generate scientific outputs is advancing faster than the ability to verify them. This workshop addresses the resulting verification bottleneck: *how should the scientific community trust, judge, and act on AI-generated scientific claims when ground truth is expensive, delayed, incomplete, or unavailable?* We will bring together researchers from scientific domains spanning mathematics, physics, chemistry, biology, medicine, climate science, and energy systems, together with experts in machine learning, formal methods, simulation, experimental validation, uncertainty quantification, and safety-critical deployment, to examine verification across three settings (1) open-ended hypothesis generation, (2) imperfect simulators and surrogate verifiers, and (3) real-world constraints involving cost, uncertainty, and safety. The workshop will feature invited talks, contributed research, verifier systems, poster sessions, and a cross-scientific-domain panel on what counts as sufficient evidence for action. By centering verification as a key challenge for AI for Science, this workshop aims to build shared vocabulary, methods, and infrastructure for determining when AI-generated science is correct, trustworthy, and actionable.
Show more
View full details
Workshop

The Third Workshop on GenAI for Health: Agentic Systems, Clinical Trust, and Future Potential

Jiayuan Ding ⋅ Pranav Rajpurkar ⋅ Junyuan Hong ⋅ Joseph Lim ⋅ Ehsan Adeli ⋅ Subhabrata Mukherjee ⋅ Ying Ding ⋅ Tanveer Syeda-Mahmood ⋅ Jiawei Xu ⋅ Jinrui Fang ⋅ Tiange Xiang ⋅ Yixin Wang ⋅ Ziheng Zhang
Dec 12, 8:00 AM - 5:00 PM Parkside 1
Generative AI (GenAI) has rapidly matured from a promising research direction into an active clinical reality, yet the path to trustworthy, human-centered deployment remains incomplete. Building on two successful GenAI4Health workshops, the field has progressed from exploratory studies and early deployments to increasingly autonomous, agentic AI systems operating across diagnosis, care planning, and patient interaction — raising new and urgent questions about safety, governance, and human oversight. This third workshop convenes machine learning researchers, healthcare professionals, policy experts, and clinical practitioners to address three interconnected frontiers: the rise of agentic, reasoning, and multimodal foundation models for health; the challenge of deploying trustworthy, policy-compliant, and human-collaborative AI in real-world clinical settings; and a forward-looking vision for ambient, conversational, and embodied AI systems that reshape how care is delivered, documented, and experienced. Our goal is to advance GenAI that is not only technically capable but also safe, equitable, and ready to earn the trust of patients, clinicians, and the institutions that serve them.
Show more
View full details
Workshop

Interpreting Agent Behavior (IAB): Human-Centered Interpretation for Understanding Agents, Humans, and Interaction

Sophia Gao ⋅ Kaiser Sun ⋅ Teresa Yeo ⋅ Jen-Tse Huang ⋅ Zhuoran Lu ⋅ Daniel Khashabi ⋅ Boyuan Zheng ⋅ Katherine Van Koevering ⋅ Sijie Ji
Dec 12, 8:00 AM - 5:00 PM MR C3.4 & C3.5
Commercial autonomous agents such as Claude and Codex now run for hours or even days to complete tasks, and along the way they show complex behavior: they plan, reason, use tools, recover from errors, coordinate with subagents, and communicate with users. We use the word behavior, as in the study of human behavior, for the full range of what an agent does during runtime. This behavior spans three levels: what agents do and how they do it (decomposing tasks, deciding under uncertainty), what people do in response (instructing, verifying, stepping in), and how the two work together through instructions and corrections. All three levels are generating vast behavioral data such as execution logs and interaction traces. Yet existing approaches read this data largely for outcomes, not through a behavior lens: benchmarks tell us whether an agent succeeds or fails, but not what it did or how it did it. Understanding what and how is what people actually need. It lets agent developers and model trainers debug failures, compare architectures, and filter training data. It lets agent users and deployment engineers watch production agents to understand safety, cost, and reliability risks. Our purpose with IAB is to identify the emerging problem spaces and challenges, pushing the community towards building this missing layer between raw behavioral data and meaningful human oversight: debugging, governance, and calibrated trust all depend on first understanding what an agent did and how.
Show more
View full details
Workshop

Who Verifies the Agents? Toward Reliable Agent Development

Ahmad Beirami ⋅ Mert Cemri ⋅ Zhang-Wei Hong ⋅ Hung Le ⋅ Ninareh Mehrabi ⋅ Melissa Pan ⋅ Dilara Soylu ⋅ Ramya Ramakrishnan
Dec 12, 8:00 AM - 5:00 PM Hall 5
Recent advances in autonomous agents have demonstrated potential for complex reasoning and open-ended tasks. However, current agent development is often fragmented and dependent on ad-hoc methods, leading to reliability challenges where performance plateaus or regresses during iteration. The fundamental bottleneck preventing the transition to scalable, reliable agent systems is verification: the inability to systematically determine whether a change to an agent constitutes genuine improvements. This workshop seeks to formalize verification as a core discipline in agent development. We convene researchers and practitioners to explore three critical pillars: developing robust verifiers that resist reward hacking, leveraging environment-grounded simulation as the ground truth for evaluation, and integrating heterogeneous signals like latency, cost, and calibration into the verification loop. By centering verification in the development lifecycle, this workshop aims to establish the foundation necessary to move toward truly autonomous and reliable agentic systems.
Show more
View full details
Workshop

Workshop on the Linguistic Principles for Foundation Models

Zhiqin Yang ⋅ Yhliu ⋅ Xiuying Chen ⋅ Ekaterina Vylomova ⋅ Bo Han ⋅ Masashi Sugiyama ⋅ Jingwen Fu
Dec 12, 8:00 AM - 5:00 PM MR C3.3
How a task is expressed, including its wording, structure, and notation, is not merely the channel through which we query foundation models (FMs); it shapes what they can do. The same problem rendered in different but meaning-equivalent forms, whether a paraphrase, another natural language, or a formal notation such as code or logic, can shift reasoning accuracy dramatically and induce entirely different internal representations, even with model parameters held fixed. The linguistic and symbolic medium is therefore a design axis for FM capability, comparable to architecture and scale, rather than a neutral interface. This workshop establishes the principled study of this medium as a first-class research agenda, linking classical linguistic structures such as compositionality, ambiguity, tokenization, pragmatics, and typological variation to model reasoning, generalization, and alignment. It convenes machine learning researchers and linguists to advance a direction complementary to scaling: improving FMs by rethinking the media through which they perceive, reason, and act.
Show more
View full details
Workshop

Beyond Next Token Prediction: Diffusion and Flow Models for Next-Generation Decoding

Minhyuk Sung ⋅ Jaihoon Kim ⋅ Sophia Tang ⋅ Subham Sahoo ⋅ Nolan Dey ⋅ Pranam Chatterjee ⋅ Molei Tao
Dec 12, 8:00 AM - 5:00 PM Hall 4
While autoregressive (AR) decoding remains the dominant paradigm for sequential data generation, it is constrained by the inherent limitations of next-token prediction. To move beyond this, discrete diffusion and flow-based generative models have emerged as compelling non-causal alternatives that enable parallel generation across a wide array of discrete domains, spanning from natural language to complex scientific data. However, transitioning to these paradigms presents a broad spectrum of open challenges, ranging from establishing theoretical frameworks and efficient algorithms to building scalable systems, and developing benchmarks and datasets for novel applications. The goal of this workshop is to bring together researchers with diverse backgrounds from both academia and industry to provide the community with a deeper understanding of the opportunities and limitations of discrete diffusion and flow models, highlight recent breakthroughs, and outline the future of next-generation decoding.
Show more
View full details
Workshop

Self-Evolving Diversity-Driven Search for Robust AI Systems

Jiao Liu ⋅ Haofeng Wu ⋅ Seyed-Mohsen Moosavi-Dezfooli ⋅ Chaoqi Chen ⋅ Catherine Huang ⋅ Sergio Escalera ⋅ Yew Soon Ong
Dec 12, 8:00 AM - 5:00 PM MR C4.6 & C4.7
This workshop studies robust AI systems through the lens of self-evolving diversity-driven search. As AI systems become increasingly interactive, multimodal, and agentic, safety failures can emerge across languages, modalities, tools, users, contexts, and multi-turn interaction trajectories. Static benchmarks and single-objective safety metrics are therefore insufficient for discovering novel safety scenarios and previously unseen failure modes. The workshop will bring together researchers from AI safety, trustworthy machine learning, evolutionary computation, multi-objective optimization, quality-diversity search, adversarial machine learning, red teaming, privacy, fairness, human-AI interaction, and AI governance. It will focus on formalizing safety scenario spaces, measuring diversity in red-teaming and evaluation benchmarks, identifying behavioral descriptors for safety-relevant failures, and using multi-objective and quality-diversity methods to navigate safety trade-offs. The goal is to build a cross-community venue for developing adaptive evaluation pipelines, diverse failure discovery methods, and robust AI safety frameworks.
Show more
View full details
Workshop

Continual World Models

Laura Leal-Taixé ⋅ Xindi Wu ⋅ Jonathan Lorraine ⋅ Qianqian Wang ⋅ Peter Tong ⋅ Amir Bar
Dec 12, 8:00 AM - 5:00 PM MR C4.4
Continual World Models is a one-day NeurIPS 2026 workshop focused on the next step beyond offline generalization: world models that update, adapt, and improve after initial training. Current video, multimodal, and embodied models learn rich priors from static datasets, but they are still typically evaluated as frozen predictors. Physical intelligence requires systems that notice prediction failures, incorporate new observations and feedback, maintain memory, and revise their understanding of objects, scenes, agents, and dynamics over time. The workshop will bring together researchers from computer vision, robotics, reinforcement learning, generative modeling, multimodal learning, and cognitive science. Through invited talks, contributed papers, posters, breakout discussions, and a panel, the workshop will define key problems for continual world modeling, including adaptation mechanisms, memory and representation design, embodied data acquisition, evaluation protocols, and the distinction between genuine model revision and superficial test-time heuristics. Our goal is to build a cross-community research agenda for adaptive world models that keep learning from the worlds they observe, imagine, and act within.
Show more
View full details
Workshop

Beyond Private Training: The New Landscape of AI Privacy

Eli Chien ⋅ Ruihan Wu ⋅ Erchi Wang ⋅ Jiachen (Tianhao) Wang ⋅ Niloofar Mireshghallah ⋅ Antti Honkela ⋅ Yu-Xiang Wang ⋅ Kamalika Chaudhuri ⋅ Ruihan
Dec 12, 8:00 AM - 5:00 PM MR C3.2
The rapid deployment of Large Language Models has underscored privacy as a critical and evolving frontier in artificial intelligence. Historically, the pursuit of privacy-preserving AI has been heavily anchored to the training phase, relying predominantly on Differentially Private Stochastic Gradient Descent. However, as the landscape of AI applications rapidly shifts toward efficient, inference-time, and non-finetuning paradigms, a significant mismatch in the current research ecosystem has been exposed. While academic literature remains disproportionately focused on privacy problems tied to the training and fine-tuning stages, industry practitioners are actively confronting novel vulnerabilities that training-stage interventions simply cannot address. These emerging challenges include the generation of differentially private synthetic data strictly at inference time, cascading data leakages within multi-step autonomous agentic systems, and privacy risks in in-context learning. The primary objective of the "Beyond Private Training: The New Landscape of AI Privacy" workshop is to call for a paradigm shift, moving the privacy discourse forward beyond the training stage. By gathering privacy researchers from both industry and academia, alongside non-privacy AI domain experts, we aim to collectively define the most pressing emerging privacy problems and synchronize theoretical rigor with production-level deployments. Furthermore, this workshop will invert the traditional paradigm by exploring how advanced foundation models can actively enforce data sanitization, verify mathematical privacy guarantees, and advance the theoretical foundations of differential privacy. Ultimately, this workshop seeks to establish robust privacy frameworks for non-finetuning scenarios before they become locked-in legacy infrastructure.
Show more
View full details
Workshop

8th Robot Learning Workshop: Is Physical AI Going Zero-Shot?

Alex Bewley ⋅ Andrey Kolobov ⋅ Hamidreza Kasaei ⋅ Roberto Calandra ⋅ Johannes V Busch ⋅ Moritz Reuss ⋅ Jiaxu Xing ⋅ Jen Jen Chung
Dec 12, 8:00 AM - 5:00 PM Grand Ballroom B2 & B3
The 8th Robot Learning Workshop returns to NeurIPS 2026 to examine a timely question: Is Physical AI going zero-shot? Driven by large-scale robotics foundation models trained across diverse data, tasks, and embodiments are providing increasingly general-purpose robotic systems. We will critically examine how these paradigms shift the traditional boundaries of robot learning. Are we moving past narrow task-specific fine-tuning toward reasoning-based physical agents? What role do scaling laws, diverse datasets, and multi-modal models play in achieving robust real-world performance? Can generalization alone lead us to the strong performance required in robotics use cases? We seek diverse perspectives from the machine learning and robotics communities—both academia and industry—to map the trajectory of zero-shot physical AI. Capitalizing on our prior experience, we will solicit several robotics researchers and companies to exhibit their latest works during the workshop's poster sessions to ground these models in physical reality.
Show more
View full details
Workshop

Continual Learning in the Era of Foundation Models and Embodied Agents

Jonghyun Choi ⋅ Xiao-Ming Wu ⋅ Rahaf Aljundi ⋅ Jorge Mendez-Mendez ⋅ Liyuan Wang ⋅ Yujie Feng ⋅ Yuankai Luo
Dec 12, 8:00 AM - 5:00 PM Hall 3
This workshop focuses on continual learning as a shared challenge for foundation models and embodied agents in dynamic, open-ended, and interactive environments. We bring these areas together because they are increasingly intertwined in modern AI: foundation models are becoming central components of many embodied systems, while embodied settings place these models in changing conditions that demand continual adaptation. This creates overlapping challenges, including catastrophic forgetting, memory and knowledge consolidation, online adaptation, long-term skill acquisition, and safe model updates. The workshop is motivated by the growing recognition that static train-once paradigms are increasingly insufficient for real-world AI systems, which must adapt over time to new tasks, environments, and user needs. It aims to bring together researchers from continual learning, foundation models, robotics, and embodied AI to identify shared technical problems, discuss emerging methods and benchmarks, and encourage closer exchange across these communities.
Show more
View full details
Workshop

Responsible Communication of Machine Learning Research in Biomedicine

Siobhan Sanford ⋅ Julia A. Meister ⋅ Qiyao Wei ⋅ Nikhil C Kurian ⋅ Ben Glocker ⋅ Jessica Schrouff
Dec 12, 8:00 AM - 5:00 PM MR E5.4 & 5
A persistent gap has emerged between what machine learning (ML) systems can deliver in biomedical contexts and how their capabilities are communicated to those who use or regulate them. In early discovery, high-profile systems such as AlphaFold and agentic research platforms such as AI Scientists have increasingly accelerated the pace at which unchecked capability claims enter public and policy discourse. In clinical deployment, decision-making tools are expanding rapidly but frameworks for communicating their limitations remain underdeveloped, illustrated by cases from IBM Watson for Oncology's widely documented capability overstatements to more recent findings that LLMs achieving near-perfect medical benchmark scores fail to improve clinical decision-making with real patients. Across the pipeline, this gap drives hype, misuse, misinterpretation and poorly informed governance, with direct consequences for user trust, funding priorities and the effective adoption of advances in ML. Yet in early discovery, frameworks to address this challenge are largely absent; in clinical settings, reporting standards such as TRIPOD+AI and TRIPOD-LLM represent important steps, but adherence remains low. A deeper contributing challenge is that the technical conventions and vocabulary that make findings legible within ML do not translate cleanly across the diverse stakeholders involved: researchers, clinicians, policymakers and science communicators frequently lack a common language and are left uncertain about what ML systems can and cannot do and unable to evaluate their claims. Because the challenges arise in the translation between communities, reporting standards alone or solutions developed by a single community in isolation cannot feasibly close the gap that unintended miscommunication creates. In response, this workshop treats structured, interdisciplinary dialogue as the method. Initiated from within the ML research community, it brings those who produce ML findings into direct exchange with the clinicians, policymakers and science communicators who must interpret and act on them. Previous NeurIPS, ICML and ICLR workshops have advanced related themes primarily from the ML perspective, including explainability, interpretability and responsible AI (RAI). This workshop builds on that foundation by shifting focus from how findings are communicated within ML to how they translate across the biomedical landscape and those who shape it. Grounded in real-world case studies, the workshop is designed to surface opportunities for evidence-based communication approaches when moving from problem to practice, with interdisciplinary exchange at its core.
Show more
View full details
Workshop

2nd Embodied Spatial Reasoning (ESR) Workshop

Wufei Ma ⋅ Haoyu Chen ⋅ Hunar Batra ⋅ Tommie Kerssies ⋅ Adam Kortylewski ⋅ Yilun Du ⋅ Alan Yuille
Dec 12, 8:00 AM - 5:00 PM MR C4.8
Embodied spatial reasoning is the ability to analyze and interpret positions, orientations, physical properties, and temporal dynamics of an agent and its surrounding objects in 3D space. Rooted in core vision tasks like 3D detection and pose estimation, this capability serves as a foundation for advanced spatial understanding, reasoning, and generation in modern embodied AI and world modeling. However, prior discussions on this topic have been fundamentally fragmented. While embodied multimodal agents derive strong spatial reasoning capabilities from multimodal alignment, world models often extract powerful priors from the video domain. Yet these two lines of work lack a shared and reliable interface for spatial, physical, and temporal modeling. In particular, current systems often struggle to maintain object permanence, spatial memory, physical consistency, and long-horizon temporal coherence when agents move, interact with objects, or revisit previously observed regions. In this workshop, we invite prominent researchers across embodied AI, spatial reasoning, and world modeling to discuss the frontiers of this intersection. For example, we will discuss various spatial priors that enable agents to ``imagine'' and plan future actions, while endowing video world models with strict physical and temporal consistency. Furthermore, we encourage submissions that explore how agents actively learn from experience, leveraging reinforcement learning and/or curiosity to build action-conditioned spatial priors.
Show more
View full details